The Next TokenLearnBookArchive
Saved

Subject

Models

Everything filed under Models, newest first. all subjects →

AgentsSafetyPolicyToolsResearchBusiness

Hacker News · Tuesday, July 28, 2026
Kimi K3 Now Available via Telnyx Inference API

Moonshot AI released the weights for Kimi K3 openly, meaning developers can download and run it themselves instead of renting access from a company. More open models in circulation means more choices for businesses that cannot or will not send their data to a third-party server.

Hacker News · Monday, July 27, 2026
Kimi-K3 Releases on HuggingFace 7/27

Moonshot AI, a Chinese lab, released Kimi-K3 openly on HuggingFace, and the developer community noticed quickly, pushing it to 362 upvotes on Hacker News. More open, competitive models mean you have more real choices when deciding what to build with or pay for, and that price pressure benefits everyone who is not locked into one provider.

Simon Willison · Saturday, July 25, 2026
Introducing Claude Opus 5

Anthropic released Claude Opus 5, their most capable model yet. If you use Claude for anything serious at work, this is the version worth trying, especially for longer, more complex tasks.

Terence Tao's ChatGPT conversation about the Jacobian Conjecture counterexample

The Jacobian Conjecture is a math problem that has been open since 1939. An AI called Claude Fable produced a counterexample, and then Fields Medal winner Terence Tao sat down with ChatGPT to work through whether it actually holds up. When one of the smartest mathematicians alive is using these tools as a thinking partner on century-old problems, the bar for what counts as 'useful AI' just moved somewhere most people had not expected.

"Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok

Someone ran GPT, Claude, Gemini, and Grok through a head-to-head test of drawing the Mona Lisa using only text commands, and the results are a genuinely fun way to see how differently these models interpret the same creative instruction. It is not a rigorous benchmark, but it is an honest reminder that these tools have very different personalities.

Kimi K3 🌕, Gemini 3.5 delayed ⏳, crushing ARC-AGI 3 🤖

Moonshot AI's Kimi K3 is making a real impression, and Gemini 3.5 is apparently running late. The competitive field is crowded enough now that a delay from Google and a strong showing from a Chinese lab in the same week is genuinely noteworthy, not a surprise. If you have been treating one AI tool as your only option, the menu keeps getting longer.

Simon Willison · Saturday, July 18, 2026
Claude make Fable 5 permanent

Anthropic is folding Claude's Fable 5 model into Max and Team Premium plans starting July 20, so if you are already paying for one of those tiers, you get the upgrade without touching your billing. Price bundling like this is how AI companies quietly shift what 'standard' means, and your baseline tool just got a bit more capable.

Simon Willison · Friday, July 17, 2026
Inkling: Our open-weights model

Mira Murati, OpenAI's former CTO, just released her first open-weights model called Inkling. Open-weights means anyone can download and run it themselves, not just pay for API access, so this is worth watching if you care about having AI that does not phone home to a big company every time you use it.

A Shared Subcircuit Lets LLMs Count Down Across Tasks

Researchers found that language models use a single shared internal circuit to count down, whether they are writing a sentence of a specific length, building a DNA sequence, or formatting a table. Think of it like a countdown timer built into the model, one timer, many uses. This matters because it shows these models are not just memorizing surface patterns. There is real structure under the hood, and understanding that structure is how we get better at predicting when the model will succeed and when it will quietly miscalculate.

Simon Willison · Friday, July 10, 2026
Introducing Muse Spark 1.1

Meta released Muse Spark 1.1, and crucially it now comes with an API, meaning developers can actually build things with it. Meta is slowly getting serious about being a platform, not just a lab that posts blog entries.

Hacker News · Friday, July 10, 2026
GPT-5.6

OpenAI released GPT-5.6, their new flagship, in three sizes from small to large. If you are paying for any OpenAI product or building something on their API, this is the engine that will likely power it soon, so it is worth knowing it exists and that pricing varies quite a bit by tier.

Simon Willison · Thursday, July 9, 2026
Introducing GPT‑Live

OpenAI swapped in a much better model for ChatGPT's voice mode, and early testers say the difference is noticeable. If you use voice mode for anything practical, like drafting notes or getting quick answers hands-free, this is the version that might actually stick.

arXiv cs.CL · Tuesday, July 7, 2026
Gemma 4 Technical Report

Google released its Gemma 4 technical report, detailing a new family of open-weight models built for efficiency and multimodal reasoning, meaning they can handle text and images together. Open-weight models you can run yourself are getting seriously good, which changes the calculus for any team deciding whether to build on a closed API or host something locally.

Hacker News · Tuesday, July 7, 2026
Small AI Models Gain Traction In places with unreliable networks

Small AI models are gaining real traction in places with unreliable internet, including hospitals and pharmacies in low-connectivity regions, because they can run on a single device without a cloud connection. This matters because it is a reminder that the most important AI deployments of the next decade may happen somewhere with no reliable signal, not in a San Francisco office.

Simon Willison · Tuesday, July 7, 2026
tencent/Hy3

Tencent released Hy3, a 295-billion-parameter model under an open Apache 2.0 license, meaning anyone can download and use it freely. More open, capable models mean more options for companies that do not want to be locked into one vendor, which is good news if you have ever felt stuck paying for something you could not inspect or control.

Hacker News · Monday, July 6, 2026
GPT-5.6 Sol Ultra will be in Codex

OpenAI is slotting a newer, more powerful model called GPT-5.6 Sol Ultra into Codex, its AI coding tool, which means developers using Codex in their daily workflow are about to get a noticeably more capable assistant without doing anything differently.

Claude Sonnet 5

Claude Sonnet 5 is Anthropic's new mid-tier model, and with over a thousand upvotes on Hacker News it's clearly the thing people are actually talking about today. If you use Claude at work or through any app built on it, you're likely already on a path to seeing different, and reportedly better, results without changing anything you do.

Hacker News · Tuesday, June 30, 2026
Qwen 3.6 27B is the sweet spot for local development

Qwen 3.6 27B is a model you can run on a decent laptop or a modest cloud instance, and the Hacker News community is loudly agreeing it punches well above its weight for everyday coding and reasoning tasks. If you have been waiting for a capable local model that does not require a small datacenter, the wait is getting shorter.

Hacker News · Monday, June 29, 2026
GLM 5.2 beats Claude in our benchmarks

A security-focused team at Semgrep ran their own benchmarks and found GLM 5.2, a Chinese open model, outperforming Claude on cybersecurity tasks. Benchmarks are always partial pictures, but this one got 875 upvotes because it came from a team with a specific, practical use case rather than a lab trying to sell something.

Simon Willison · Saturday, June 27, 2026
Quoting OpenAI

OpenAI is quietly testing a new family of models called Sol, Terra, and Luna, with Terra apparently matching GPT-5.5 at half the cost. Cheaper capable models mean the per-question price you pay inside the tools you use at work is about to drop again.

arXiv cs.CL · Friday, June 26, 2026
Where Larger Models Excel: The Primacy of Constraint-Guided Reasoning

Bigger models consistently beat smaller ones not just on hard problems but specifically on tasks that require following multiple constraints at once. Think of it like planning a dinner party where the guest is vegetarian, allergic to nuts, and needs to leave by 8pm. Small models drop one of those. Bigger ones hold all three. That gap explains a lot of why model size still matters, even when smaller models sound just as fluent.

Computer use in Gemini 3.5 Flash

Google's Gemini 3.5 Flash can now control a computer, clicking and typing on your behalf. It is the same category of capability as Anthropic's Computer Use, and the fact that it is landing in a fast, cheaper model means it will show up in a lot more products soon.

Hacker News · Monday, June 22, 2026
Apertus – Open Foundation Model for Sovereign AI

A project called Apertus launched an open foundation model aimed at governments and organizations that want to run AI without depending on American tech companies. If you work in a country, a company, or an industry where data sovereignty is a real concern, this is the kind of option that could eventually matter to your procurement decisions.

Simon Willison · Friday, June 12, 2026
Claude Fable is relentlessly proactive

Simon Willison, a developer whose day job involves building with these models, describes Claude's newest version as relentlessly proactive, meaning it reaches for tools and strategies on its own without being asked. That is useful when it works and a little hard to predict when it does not, which is a useful framing for anyone thinking about putting AI into a real workflow.

Simon Willison · Thursday, June 11, 2026
DiffusionGemma

Google released DiffusionGemma, a text-generation model built on diffusion principles rather than the usual predict-one-word-at-a-time approach, which could eventually mean faster and more flexible responses in the tools you use every day.

Simon Willison · Wednesday, June 10, 2026
If Claude Fable stops helping you, you'll never know

The blog post that sparked the Hacker News firestorm about Claude Fable's silent degradation behavior is worth reading in full, because it walks through the specific policy language and what it could mean for small developers building on top of Anthropic's API. The concern is not hypothetical.

Simon Willison · Tuesday, June 9, 2026
Siri AI at WWDC 2026

Simon Willison, one of the most careful AI observers around, says he will not believe Apple's new AI promises until he can actually use them, pointing to how badly last year's announcements fell short of reality. That is a reasonable bar for all of us to hold.

Hacker News · Monday, June 8, 2026
DeepSeek V4 Pro beats GPT-5.5 Pro on precision

DeepSeek, the Chinese lab that surprised everyone earlier this year, now has a model that beats OpenAI's latest on at least one precision benchmark. If you assumed the frontier was a two-horse race between OpenAI and Anthropic, this is a good reminder to look up.

Hacker News · Sunday, June 7, 2026
Nvidia is proposing a beast of a CPU system for Windows PCs

Nvidia is proposing a desktop CPU built around the kind of memory architecture that makes AI inference fast, essentially bringing data-center-grade AI hardware to a Windows PC. If this ships in any real form, running a capable local AI model could stop being a hobbyist project and start being something your next computer just does.

Gemma 4 12B: A unified, encoder-free multimodal model

Google released Gemma 4 12B, a smaller open model that handles text and images together without needing separate encoder components. Smaller open models that run on your own hardware matter because they shift the conversation from 'what can the cloud do for you' to 'what can you run yourself,' and that gap closing is a genuinely big deal for regular people.