GPT-5.6 got a price cut, a new small model called Inkling-Small arrived, and Google showed off Gemini Robotics 2. Cheaper models mean more of this technology quietly showing up inside the apps and services you already use, whether or not you notice.
Subject
Models
Everything filed under Models, newest first. all subjects →
DeepSeek's V4 Flash 0731 is getting solid marks for speed and value, essentially a fast, inexpensive model that punches above its price tag. More competition at the low end is good news for anyone building something on a budget.
OpenAI scored at the top of ARC-AGI-3, a benchmark designed to test reasoning that cannot be memorized, and GPT-5.6 is now cheaper to run than its predecessor. If the price keeps dropping, the cost argument for not using AI tools at work gets harder to make.
A team distilled a large DeepSeek model's financial reasoning into a smaller open model and found that the knowledge transferred but the censorship did not. It is an early, practical data point on how open-weight models can be customized for specific industries without inheriting all the original model's baggage.
Google DeepMind showed a robot that can coordinate its whole body to do real physical tasks, think складывание laundry or handling objects in a cluttered space. The gap between a robot that can talk and a robot that can actually help around the house got a little smaller today.
A team built a small language model trained only on historical texts so it thinks and writes from the past, on purpose. It is a creative experiment in controlling what an AI knows and does not know, and it hints at a future where you might pick a model the way you pick a reference book, by what it was taught and when.
Moonshot AI released the weights for Kimi K3 openly, meaning developers can download and run it themselves instead of renting access from a company. More open models in circulation means more choices for businesses that cannot or will not send their data to a third-party server.
Moonshot AI, a Chinese lab, released Kimi-K3 openly on HuggingFace, and the developer community noticed quickly, pushing it to 362 upvotes on Hacker News. More open, competitive models mean you have more real choices when deciding what to build with or pay for, and that price pressure benefits everyone who is not locked into one provider.
Anthropic released Claude Opus 5, their most capable model yet. If you use Claude for anything serious at work, this is the version worth trying, especially for longer, more complex tasks.
The Jacobian Conjecture is a math problem that has been open since 1939. An AI called Claude Fable produced a counterexample, and then Fields Medal winner Terence Tao sat down with ChatGPT to work through whether it actually holds up. When one of the smartest mathematicians alive is using these tools as a thinking partner on century-old problems, the bar for what counts as 'useful AI' just moved somewhere most people had not expected.
Someone ran GPT, Claude, Gemini, and Grok through a head-to-head test of drawing the Mona Lisa using only text commands, and the results are a genuinely fun way to see how differently these models interpret the same creative instruction. It is not a rigorous benchmark, but it is an honest reminder that these tools have very different personalities.
Google quietly announced that its latest Gemini models will ignore the temperature, top-p, and top-k settings that developers use to control how creative or predictable the model's responses are. If you or your team pipe Gemini into any product, your carefully tuned behavior may already be different from what you set it to be.
Alibaba's Qwen team released a new image-understanding model with a focus on reading dense, detail-rich content like charts, documents, and technical diagrams. If you work with reports or data visuals, this class of model is getting quietly useful for extracting information you'd otherwise have to read manually.
Claude reportedly produced a counterexample to the Jacobian Conjecture, a problem that has been open since 1939. Whether it holds up under peer review or not, it is a useful reminder that these models are now operating in territory where the output genuinely surprises the mathematicians watching.
OpenAI quietly shrunk the context window for its Codex model, cutting it from 372,000 tokens down to 272,000. If you or your team built any automated coding workflows that depended on feeding Codex a very large codebase in one shot, those workflows may now break without warning.
Moonshot AI's Kimi K3 is making a real impression, and Gemini 3.5 is apparently running late. The competitive field is crowded enough now that a delay from Google and a strong showing from a Chinese lab in the same week is genuinely noteworthy, not a surprise. If you have been treating one AI tool as your only option, the menu keeps getting longer.
Anthropic is folding Claude's Fable 5 model into Max and Team Premium plans starting July 20, so if you are already paying for one of those tiers, you get the upgrade without touching your billing. Price bundling like this is how AI companies quietly shift what 'standard' means, and your baseline tool just got a bit more capable.
Chinese lab Moonshot AI released Kimi K3, a 2.8 trillion parameter model, with open weights promised by July 27. More open, capable models arriving from more places means the field is genuinely less centralized than it was even six months ago.
Mira Murati, OpenAI's former CTO, just released her first open-weights model called Inkling. Open-weights means anyone can download and run it themselves, not just pay for API access, so this is worth watching if you care about having AI that does not phone home to a big company every time you use it.
A developer ran Google's Gemma 4 26B model on decade-old server hardware with no GPU at all, getting five words per second out of it. If you have been told you need expensive new equipment to run serious AI locally, this is worth bookmarking.
Researchers found that language models use a single shared internal circuit to count down, whether they are writing a sentence of a specific length, building a DNA sequence, or formatting a table. Think of it like a countdown timer built into the model, one timer, many uses. This matters because it shows these models are not just memorizing surface patterns. There is real structure under the hood, and understanding that structure is how we get better at predicting when the model will succeed and when it will quietly miscalculate.
A team swapped in GPT-5.6 and their agent ran more than twice as fast at a lower cost, with no changes to their code. That is the kind of boring, practical result that actually changes whether you can afford to keep a product running.
Meta released Muse Spark 1.1, and crucially it now comes with an API, meaning developers can actually build things with it. Meta is slowly getting serious about being a platform, not just a lab that posts blog entries.
OpenAI released GPT-5.6, their new flagship, in three sizes from small to large. If you are paying for any OpenAI product or building something on their API, this is the engine that will likely power it soon, so it is worth knowing it exists and that pricing varies quite a bit by tier.
OpenAI swapped in a much better model for ChatGPT's voice mode, and early testers say the difference is noticeable. If you use voice mode for anything practical, like drafting notes or getting quick answers hands-free, this is the version that might actually stick.
Researchers found that so-called reasoning models, the ones that show their work before answering, still hallucinate facts, and they propose a way to reduce that during training. If you rely on a reasoning model for anything factual, it is worth knowing that the visible thinking process does not guarantee the final answer is grounded in reality.
Google released its Gemma 4 technical report, detailing a new family of open-weight models built for efficiency and multimodal reasoning, meaning they can handle text and images together. Open-weight models you can run yourself are getting seriously good, which changes the calculus for any team deciding whether to build on a closed API or host something locally.
Small AI models are gaining real traction in places with unreliable internet, including hospitals and pharmacies in low-connectivity regions, because they can run on a single device without a cloud connection. This matters because it is a reminder that the most important AI deployments of the next decade may happen somewhere with no reliable signal, not in a San Francisco office.
Tencent released Hy3, a 295-billion-parameter model under an open Apache 2.0 license, meaning anyone can download and use it freely. More open, capable models mean more options for companies that do not want to be locked into one vendor, which is good news if you have ever felt stuck paying for something you could not inspect or control.
OpenAI is slotting a newer, more powerful model called GPT-5.6 Sol Ultra into Codex, its AI coding tool, which means developers using Codex in their daily workflow are about to get a noticeably more capable assistant without doing anything differently.
GPT-5.5 Codex appears to be grouping its reasoning steps in a way that hurts the quality of its code output, and 280 people on Hacker News noticed. If you rely on Codex for anything in your workflow right now, it is worth double-checking its recent output rather than assuming the newest version is the best version.
Claude Sonnet 5 is Anthropic's new mid-tier model, and with over a thousand upvotes on Hacker News it's clearly the thing people are actually talking about today. If you use Claude at work or through any app built on it, you're likely already on a path to seeing different, and reportedly better, results without changing anything you do.
Qwen 3.6 27B is a model you can run on a decent laptop or a modest cloud instance, and the Hacker News community is loudly agreeing it punches well above its weight for everyday coding and reasoning tasks. If you have been waiting for a capable local model that does not require a small datacenter, the wait is getting shorter.
A security-focused team at Semgrep ran their own benchmarks and found GLM 5.2, a Chinese open model, outperforming Claude on cybersecurity tasks. Benchmarks are always partial pictures, but this one got 875 upvotes because it came from a team with a specific, practical use case rather than a lab trying to sell something.
OpenAI is quietly testing a new family of models called Sol, Terra, and Luna, with Terra apparently matching GPT-5.5 at half the cost. Cheaper capable models mean the per-question price you pay inside the tools you use at work is about to drop again.
Bigger models consistently beat smaller ones not just on hard problems but specifically on tasks that require following multiple constraints at once. Think of it like planning a dinner party where the guest is vegetarian, allergic to nuts, and needs to leave by 8pm. Small models drop one of those. Bigger ones hold all three. That gap explains a lot of why model size still matters, even when smaller models sound just as fluent.
A new study shows that the fine-tuning process used to make models more helpful can quietly erode the empathy and care that was baked in earlier during training. The effort to make AI better at answering questions may be trading away something quieter but important, and the tradeoff is domain-specific, not uniform.
Google's Gemini 3.5 Flash can now control a computer, clicking and typing on your behalf. It is the same category of capability as Anthropic's Computer Use, and the fact that it is landing in a fast, cheaper model means it will show up in a lot more products soon.
A project called Apertus launched an open foundation model aimed at governments and organizations that want to run AI without depending on American tech companies. If you work in a country, a company, or an industry where data sovereignty is a real concern, this is the kind of option that could eventually matter to your procurement decisions.
DeepSeek's new V4 model can hold roughly a million words in its head at once, which means it could read an entire legal contract, a year of emails, or a full codebase in a single sitting. If this lands in tools you actually use, the days of having to carefully paste in 'just the relevant part' may be numbered.
Z.ai released GLM-5.2 with full open weights under an MIT license, meaning anyone can download and run it without asking permission or paying a subscription. If this holds up to scrutiny, it changes the math for small teams and solo developers who need a powerful text model but cannot afford the big commercial APIs.
LLMs can correctly identify that a user is from a different culture, but then respond as if they are not. Knowing something and using that knowledge are two different skills, and these models have one without the other.
A developer asked what some call a dangerous AI model to write a fable, and got a genuinely charming little game about a shepherd's dog. It is a small, low-stakes reminder that creative output from these models is often more interesting than the benchmarks suggest.
Simon Willison, a developer whose day job involves building with these models, describes Claude's newest version as relentlessly proactive, meaning it reaches for tools and strategies on its own without being asked. That is useful when it works and a little hard to predict when it does not, which is a useful framing for anyone thinking about putting AI into a real workflow.
Google released DiffusionGemma, a text-generation model built on diffusion principles rather than the usual predict-one-word-at-a-time approach, which could eventually mean faster and more flexible responses in the tools you use every day.
The blog post that sparked the Hacker News firestorm about Claude Fable's silent degradation behavior is worth reading in full, because it walks through the specific policy language and what it could mean for small developers building on top of Anthropic's API. The concern is not hypothetical.
Simon Willison, one of the most careful AI observers around, says he will not believe Apple's new AI promises until he can actually use them, pointing to how badly last year's announcements fell short of reality. That is a reasonable bar for all of us to hold.
DeepSeek, the Chinese lab that surprised everyone earlier this year, now has a model that beats OpenAI's latest on at least one precision benchmark. If you assumed the frontier was a two-horse race between OpenAI and Anthropic, this is a good reminder to look up.
Leaked details about Anthropic's next model, called Oceanus, are circulating, and there is also buzz about a ChatGPT feature that does something during idle time that people are calling dreaming. Take the leaks with a grain of salt, but they signal both companies are pushing hard on what the models do between your prompts, not just during them.
Nvidia is proposing a desktop CPU built around the kind of memory architecture that makes AI inference fast, essentially bringing data-center-grade AI hardware to a Windows PC. If this ships in any real form, running a capable local AI model could stop being a hobbyist project and start being something your next computer just does.
Google released Gemma 4 12B, a smaller open model that handles text and images together without needing separate encoder components. Smaller open models that run on your own hardware matter because they shift the conversation from 'what can the cloud do for you' to 'what can you run yourself,' and that gap closing is a genuinely big deal for regular people.