The Next TokenLearnBookArchive
Saved

Monday, June 29, 2026 · about a 2 minute read

Who's Watching the Watchers

Today's stories keep circling the same quiet question: when AI is grading your exam, reading your MRI, or writing your code, who is actually in charge of checking its work?

Get the calm version of AI news.

One email a day on what is actually happening in AI, in plain English. No hype, no doom.

Free. One calm email a day. No hype, no doom.

Hacker NewsTools
I used Claude Code to get a second opinion on my MRI

A person ran their MRI results through Claude and found it surfaced details their radiologist had mentioned but not fully explained. This is not a replacement for a doctor, but it is a preview of what it looks like when AI becomes the second reader you can actually afford.

Read
Hacker NewsPolicy
Professor denounces mass AI fraud on an exam at Brown

A Brown University professor flagged what appears to be widespread AI-written submissions on a single exam, enough to make it a public story. If you work in education, or manage people whose work you review, this is the moment where 'assume good faith' gets complicated.

Read
Hacker NewsModels
GLM 5.2 beats Claude in our benchmarks

A security-focused team at Semgrep ran their own benchmarks and found GLM 5.2, a Chinese open model, outperforming Claude on cybersecurity tasks. Benchmarks are always partial pictures, but this one got 875 upvotes because it came from a team with a specific, practical use case rather than a lab trying to sell something.

Read

Get this every morning.

arXiv cs.CLResearch
Can LLMs Judge Better Than They Generate? Evaluating Task Asymmetry, Mechanistic Interpretability and Transferability for In-Context QA

Researchers tested whether AI models are better at judging answers than generating them, and the result is genuinely interesting: not always. Think of it like a student who can spot a bad essay but still writes bad essays. The AI tools that grade output, including their own output, are not some neutral referee sitting above the model. They are the same model wearing a different hat, and that matters every time a product uses 'AI as judge' to tell you whether another AI did a good job.

Read
Want the slow, plain-English version of why this matters? This is exactly the kind of idea the book was written to unpack, one light-switch analogy at a time.JPWExplained properly in the book
Simon WillisonAgents
Quoting Jon Udell

Jon Udell flipped the phrase 'human in the loop' to ' in the loop,' arguing the framing matters because it changes who we think of as being in charge. Small language shift, but worth sitting with if you are building workflows where AI takes actions on your behalf.

Read

That's today. See you tomorrow.

Get this every morning.

One email a day on what is actually happening in AI, in plain English. No hype, no doom.

Free. One calm email a day. No hype, no doom.

Just Predicting Words book cover

The book behind this newsletter

Just Predicting Words

How ChatGPT, Claude, and Modern AI Actually Work

The trick is small. The world it built is not.

PaperbackKindleAudiobook · SpotifyAudiobook · Google Play