The Next TokenLearnBookArchive
Saved

Friday, June 12, 2026 · about a 2 minute read

AI That Listens, Learns, and Sometimes Gets Distracted

Today's research keeps bumping into the same honest problem: these models are impressive until you hand them something slightly messy, and then things get interesting.

Get the calm version of AI news.

One email a day on what is actually happening in AI, in plain English. No hype, no doom.

Free. One calm email a day. No hype, no doom.

arXiv cs.CLResearch
Multi-Turn Reasoning When Context Arrives in Pieces: Scalable Sharding and Memory-Augmented RL

When you give a chatbot important details piece by piece across a long conversation, its accuracy can drop by 65%, even though all the information is technically still there in the window. If you use an AI assistant for anything that builds up over multiple messages, like a project plan or a medical question, this is a real and current limitation you should know about.

Read
Simon WillisonModels
Claude Fable is relentlessly proactive

Simon Willison, a developer whose day job involves building with these models, describes Claude's newest version as relentlessly proactive, meaning it reaches for tools and strategies on its own without being asked. That is useful when it works and a little hard to predict when it does not, which is a useful framing for anyone thinking about putting AI into a real workflow.

Read
arXiv cs.CLSafety
One Jailbreak, Many Tongues: Learning Language-Insensitive Intention Representations for Multilingual Jailbreak Detection

Researchers found that if you ask an AI to do something harmful in a language other than English, the safety guardrails are noticeably weaker, because safety training has been concentrated almost entirely in English. If you or your organization use AI tools with a global team, the safety behavior your English-speaking colleagues see may not be what everyone else gets.

Read
arXiv cs.CLTools
Shopping Reasoning Bench: An Expert-Authored Benchmark for Multi-Turn Conversational Shopping Assistants

A new was built specifically to test AI shopping assistants on multi-turn conversations, the kind where you say 'I need something for a wedding' and then 'actually it is outdoor' and then 'my budget changed.' Current models handle these real shopping conversations worse than the clean single-question tests suggest, so the gap between demo and reality is still wide.

Read

Get this every morning.

arXiv cs.CLResearch
The Structural Attention Tax: How Retrieval Format Hijacks In-Context Learning Independent of Content

Researchers found that the format you use to feed information to a model, how you structure and present the text, changes the model's output independently of what the information actually says. Think of it like reading a memo versus reading a legal contract. Same words, different shape on the page, and your brain pays differently. Models do the same thing, and that is worth knowing if you ever craft prompts or build anything on top of a retrieval system.

Read
Want the slow, plain-English version of why this matters? This is exactly the kind of idea the book was written to unpack, one light-switch analogy at a time.JPWExplained properly in the book
arXiv cs.CLPolicy
Polar: A Benchmark for Evaluating Political Bias in LLMs

A new called Polar tests AI models for political bias across multiple countries and languages, not just American English political categories. As these models get used in news tools and civic applications around the world, knowing whether they lean one way in one country and another way elsewhere is a genuinely important question that nobody has had a clean way to measure until now.

Read
arXiv cs.CLResearch
Observable Patterns Are Not Explanations: A Causal-Geometric Analysis of Latent Reasoning Models

A paper argues that when researchers see interesting patterns inside a model's internal states, like evidence it is doing something that looks like reasoning, those patterns are not the same thing as proof that reasoning is actually happening. It is a small distinction that sounds philosophical until you realize it affects how much you should trust a model when it confidently shows its work.

Read

That's today. See you tomorrow.

Get this every morning.

One email a day on what is actually happening in AI, in plain English. No hype, no doom.

Free. One calm email a day. No hype, no doom.

Just Predicting Words book cover

The book behind this newsletter

Just Predicting Words

How ChatGPT, Claude, and Modern AI Actually Work

The trick is small. The world it built is not.

PaperbackKindleAudiobook · SpotifyAudiobook · Google Play