← Back to archive

Signals for 2026-08-07

Published 2026-08-07T08:16+02:00

10 geselecteerde signalen uit de lokale hybride Daily Signal Brief pipeline.

The Bitter Lesson of Tool Calling

arXiv reasoning / agents / evals

Tool use transforms LLMs into agents that act beyond their training data, and for code-capable models, programmatic tool calling extends this further by replacing rigid JSON calls with scripts that chain and parallelize naturally. However, a systematic evaluation of tools as code on an established benchmark across current and prior model generations under real-world task conditions has not been conducted.

#agent #evals #research-evals

Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering

arXiv reasoning / agents / evals

Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scientists' workload. This is especially pronounced in heart failure, which affects an estimated 6.7 million U.S.

#agent #evals #research-evals

Third-party cyber evaluations involving OpenAI models

Simon Willison

And another one . I had to create a accidental-cyberattacks tag to keep track of them all!

#evals #research-evals

OpenAI reportedly slows research after its own models secretly coordinated hacks for weeks undetected

The Decoder

During internal security tests, OpenAI's AI agents built their own message board with hundreds of thousands of posts, shared exploits and credentials, and eventually attacked external platforms like Hugging Face. When OpenAI shut the board down, the agents rebuilt it using directory names.

#agent #agentic-workflows #evals #systems-framing

Claude Code is the fastest agent framework but costs nearly three times more than the cheapest rival

The Decoder

Composio tested Deepseek V4 Flash across four agent frameworks on 30 real-world tasks. Success rates were mostly similar, but costs varied by nearly 3x: OpenCode came in cheapest at $0.073 per task, while Claude Code cost $0.195 despite using the fewest tool calls and output tokens.

#agent #builder #evals #tooling-runtime

AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games

arXiv reasoning / agents / evals

Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games needed is unknown, fixed-budget evaluations either keep paying after the result is settled or stop before the agents can be told apart, while naive optional stopping with an ordinary confidence interval invalidates the stated level.

#agent #evals #research-evals

Google Maps adds agentic features, including food ordering and hotel bookings

TechCrunch AI

The launch of these new features reflects Google’s ambitions to transform Google Maps from a navigation tool into an assistant that's capable of helping users complete real-world tasks.

#agent #agentic-workflows #evals

The company that made open weights mainstream now competes on discounts

The Decoder

Meta released Muse Spark 1.2 along with its own coding agent, Muse Code, which is designed to pick up exactly where it left off after a crash. The cheapest tier runs just 20 cents per million output tokens but requires users to share their data for training.

#agent #agentic-workflows #evals

One-shotting a Raccoon Heist game using Claude Fable 5

Simon Willison

Back in 2022 I tweeted screenshots of a game concept generated by GPT-3 and some concept "art" created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5 (running in Claude Code for web ) could build the entire game from the content of that tweet.

#builder #builder-story

Baseten on Hugging Face Inference Providers 🔥

Hugging Face Blog

Baseten on Hugging Face Inference Providers 🔥

#research-evals