OpenAI released GPT-5.6, their new flagship, in three sizes from small to large. If you are paying for any OpenAI product or building something on their API, this is the engine that will likely power it soon, so it is worth knowing it exists and that pricing varies quite a bit by tier.
Friday, July 10, 2026 · about a 2 minute read
Better Models, Shakier Rulers
Today the AI world released new things and quietly admitted that some of the tools we use to judge those things might not be trustworthy. That gap, between building and measuring, is worth sitting with.
Get the calm version of AI news.
One email a day on what is actually happening in AI, in plain English. No hype, no doom.
Free. One calm email a day. No hype, no doom.
The three GPT-5.6 sizes, Luna, Terra, and Sol, are priced very differently from each other, which means the model a company picks for their app will often come down to budget, not what is smartest. The next time an AI product feels a little dumber than you expected, cost is a reasonable guess.
Meta released Muse Spark 1.1, and crucially it now comes with an API, meaning developers can actually build things with it. Meta is slowly getting serious about being a platform, not just a lab that posts blog entries.
A team spent a year learning that teaching a five-year-old to read is genuinely hard, even for AI, because a good tutor has to know when to push, when to back off, and how to keep a child from just guessing randomly. If you have young kids, this space is moving fast and it is worth paying to what actually works versus what just looks impressive in a demo.
Get this every morning.
Researchers found that AI evaluation scores shift when you simply replace the AI doing the judging, even if the answers being graded stay identical. Think of it like this: if you wrote the same essay and handed it to three different teachers, you might get a B, a B-plus, and a C-plus. We have been using AI to grade AI and quietly assuming the ruler is fixed. It is not, and that matters because almost every you read about, every claim that model X beats model Y, relies on someone trusting that ruler.
Researchers at EPFL built a system that generates videos specifically designed to activate a targeted region of the brain, essentially reverse-engineering visual stimuli from neural responses. It sounds like science fiction but it is a real tool for neuroscience, and it is also the kind of research that will eventually raise serious questions about persuasion and .
A paper laid out a simple but important idea: saying no is not one skill, it is two. A model should refuse things it would get wrong, and separately refuse things that are genuinely unanswerable or built on a false premise. Right now most models treat both the same way, which is why they sometimes confidently answer a trick question instead of stopping to say the question itself is broken.
That's today. See you tomorrow.
Get this every morning.
One email a day on what is actually happening in AI, in plain English. No hype, no doom.
Free. One calm email a day. No hype, no doom.

The book behind this newsletter
Just Predicting Words
How ChatGPT, Claude, and Modern AI Actually Work
The trick is small. The world it built is not.