← Back to archive

Signals for 2026-08-03

Published 2026-08-03T08:16+02:00

8 geselecteerde signalen uit de lokale hybride Daily Signal Brief pipeline.

OpenAI Presence wants to make AI agents production-ready for businesses

The Decoder

OpenAI's new enterprise offering, Presence, is designed to get AI agents into production for customer service and internal workflows. Unlike the existing Workspace Agents, Presence targets external deployments.

#agent #implementation #implementation-adoption

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

arXiv reasoning / agents / evals

Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata.

#agent #evals #implementation #research-evals

After Hugging Face incident, METR urges independent root-cause investigations into AI agent misbehavior

The Decoder

Research organization METR is calling for systematic, independently led investigations whenever AI agents act autonomously against their developers' intentions. The push comes partly in response to the Hugging Face hack carried out by OpenAI models.

#agent #agentic-workflows #builder #evals

Meta AI uses a second AI agent as a memory coach to keep long tasks on track

The Decoder

Meta AI wants to stop AI agents from forgetting errors they've already diagnosed and repeating failed steps during complex tasks. A separate memory agent maintains a structured memory bank and decides when to remind the main agent and when to stay silent.

#agent #agentic-workflows #evals

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers

arXiv reasoning / agents / evals

As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly important. Existing benchmarks typically focus on static code generation, paper replication, or final answer correctness, but do not directly assess whether agents can interpret experimental evidence and use it to guide subsequent hyperparameter decisions.

#agent #evals #research-evals

From Code Review to Code Critique: Intent, Drift, and Spotlight for AI-Generated Diffs at Scale

arXiv reasoning / agents / evals

AI coding agents are generating code at volumes that exceed the capacity of traditional peer review. At the same time, existing AI code review tools over-index on low-value suggestions such as style and best practices while under-indexing on the concerns human reviewers prioritize most: correctness, security, and performance.

#agent #builder #evals #research-evals

July 2026 newsletter

Simon Willison

The June edition of my sponsors-only monthly newsletter is out. If you are a sponsor (or if you start a sponsorship now) you can access it here .

#builder #models-architecture

Latest open artifacts (#23): Laguna S2.1, Inkling, & Kimi K3 show the utility of open models on the Pareto frontier

Interconnects

Capacity to train strong models is proliferating.

#models-architecture