← Back to archive
Signals for 2026-08-03
Published 2026-08-03T08:16+02:00
8 geselecteerde signalen uit de lokale hybride Daily Signal Brief pipeline.
The Decoder
OpenAI's new enterprise offering, Presence, is designed to get AI agents into production for customer service and internal workflows. Unlike the existing Workspace Agents, Presence targets external deployments.
#agent #implementation #implementation-adoption
arXiv reasoning / agents / evals
Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata.
#agent #evals #implementation #research-evals
The Decoder
Research organization METR is calling for systematic, independently led investigations whenever AI agents act autonomously against their developers' intentions. The push comes partly in response to the Hugging Face hack carried out by OpenAI models.
#agent #agentic-workflows #builder #evals
The Decoder
Meta AI wants to stop AI agents from forgetting errors they've already diagnosed and repeating failed steps during complex tasks. A separate memory agent maintains a structured memory bank and decides when to remind the main agent and when to stay silent.
#agent #agentic-workflows #evals
arXiv reasoning / agents / evals
As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly important. Existing benchmarks typically focus on static code generation, paper replication, or final answer correctness, but do not directly assess whether agents can interpret experimental evidence and use it to guide subsequent hyperparameter decisions.
#agent #evals #research-evals
arXiv reasoning / agents / evals
AI coding agents are generating code at volumes that exceed the capacity of traditional peer review. At the same time, existing AI code review tools over-index on low-value suggestions such as style and best practices while under-indexing on the concerns human reviewers prioritize most: correctness, security, and performance.
#agent #builder #evals #research-evals
Simon Willison
The June edition of my sponsors-only monthly newsletter is out. If you are a sponsor (or if you start a sponsorship now) you can access it here .
#builder #models-architecture
Interconnects
Capacity to train strong models is proliferating.
#models-architecture