The Next TokenLearnBookArchive
Saved

Friday, July 10, 2026 · about a 2 minute read

Better Models, Shakier Rulers

Today the AI world released new things and quietly admitted that some of the tools we use to judge those things might not be trustworthy. That gap, between building and measuring, is worth sitting with.

Get the calm version of AI news.

One email a day on what is actually happening in AI, in plain English. No hype, no doom.

Free. One calm email a day. No hype, no doom.

Hacker NewsModels
GPT-5.6

OpenAI released GPT-5.6, their new flagship, in three sizes from small to large. If you are paying for any OpenAI product or building something on their API, this is the engine that will likely power it soon, so it is worth knowing it exists and that pricing varies quite a bit by tier.

Read
Simon WillisonBusiness
The new GPT-5.6 family: Luna, Terra, Sol

The three GPT-5.6 sizes, Luna, Terra, and Sol, are priced very differently from each other, which means the model a company picks for their app will often come down to budget, not what is smartest. The next time an AI product feels a little dumber than you expected, cost is a reasonable guess.

Read
Simon WillisonModels
Introducing Muse Spark 1.1

Meta released Muse Spark 1.1, and crucially it now comes with an API, meaning developers can actually build things with it. Meta is slowly getting serious about being a platform, not just a lab that posts blog entries.

Read
Hacker NewsTools
Building a real-time AI tutor for 5-year-olds

A team spent a year learning that teaching a five-year-old to read is genuinely hard, even for AI, because a good tutor has to know when to push, when to back off, and how to keep a child from just guessing randomly. If you have young kids, this space is moving fast and it is worth paying to what actually works versus what just looks impressive in a demo.

Read

Get this every morning.

arXiv cs.CLResearch
When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability

Researchers found that AI evaluation scores shift when you simply replace the AI doing the judging, even if the answers being graded stay identical. Think of it like this: if you wrote the same essay and handed it to three different teachers, you might get a B, a B-plus, and a C-plus. We have been using AI to grade AI and quietly assuming the ruler is fixed. It is not, and that matters because almost every you read about, every claim that model X beats model Y, relies on someone trusting that ruler.

Read
Want the slow, plain-English version of why this matters? This is exactly the kind of idea the book was written to unpack, one light-switch analogy at a time.JPWExplained properly in the book
Hacker NewsSafety
AI-generated videos to maximally drive a target brain region

Researchers at EPFL built a system that generates videos specifically designed to activate a targeted region of the brain, essentially reverse-engineering visual stimuli from neural responses. It sounds like science fiction but it is a real tool for neuroscience, and it is also the kind of research that will eventually raise serious questions about persuasion and .

Read
arXiv cs.CLResearch
Two Axes of LLM Abstention: Answer Correctness and Question Answerability

A paper laid out a simple but important idea: saying no is not one skill, it is two. A model should refuse things it would get wrong, and separately refuse things that are genuinely unanswerable or built on a false premise. Right now most models treat both the same way, which is why they sometimes confidently answer a trick question instead of stopping to say the question itself is broken.

Read

That's today. See you tomorrow.

Get this every morning.

One email a day on what is actually happening in AI, in plain English. No hype, no doom.

Free. One calm email a day. No hype, no doom.

Just Predicting Words book cover

The book behind this newsletter

Just Predicting Words

How ChatGPT, Claude, and Modern AI Actually Work

The trick is small. The world it built is not.

PaperbackKindleAudiobook · SpotifyAudiobook · Google Play