← Back to archive

Signals for 2026-07-27

Published 2026-07-27T08:15+02:00

4 geselecteerde signalen uit de lokale hybride Daily Signal Brief pipeline.

Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence

The Decoder

Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, nearly quadrupling GPT-5.6 Sol's previous record of 7.8 percent. The benchmark's developers say the model independently formulated reflection equations, a behavior they had never seen from another model, and attribute to stronger logical reasoning.

#builder #evals #research-evals

An Inside Look at the Relay Market Powering Token Resellers and Fraud

Simon Willison

Fascinating investigation by Matt Lenhard into the market that has grown up around reselling LLM tokens at a discount by pooling API keys from various sources. This looks to be mostly a thing in China.

#builder #tooling-runtime

Cursor's agent swarm suggests cheaper models can handle most coding when frontier models plan the work

The Decoder

Cursor asked its upgraded agent swarm and its predecessor to rebuild SQLite in Rust using only the documentation, with no source code or internet access. Every configuration of the new system, which separates planners from workers, eventually scored 100 percent on the test suite.

#agent #builder #models-architecture

Hundreds asked ChatGPT for poison and bioweapon recipes and some got step-by-step high school level guides

The Decoder

In summer 2025, OpenAI internally flagged GPT-5 as high-risk because it helped users create biological hazards, but downgraded the model's risk rating that fall. According to the Wall Street Journal, some users got step-by-step instructions for making poisons and biological weapons.

#implementation-adoption