← Back to archive
Signals for 2026-07-28
Published 2026-07-28T08:17+02:00
10 geselecteerde signalen uit de lokale hybride Daily Signal Brief pipeline.
MIT Technology Review AI
For the enterprise, the promise of agentic AI is much more than just a better chatbot. It is software agents that execute business tasks end-to-end across people, business workflows, data, and systems.
#agent #agentic-workflows #implementation #systems-framing
Simon Willison
It's interesting watching the evolution of Ethan Mollick's guide over time. A year ago it was still all about chat - ChatGPT, Claude, Gemini - with o3, Claude 4 Opus, and Gemini 2.5 Pro as the models and Deep Research as a useful alternative mode.
#agent #agentic-workflows #builder #evals
arXiv reasoning / agents / evals
Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events. However, expert-centered warning workflows are costly, labor-intensive, and difficult to scale throughout the warning-to-action process.
#agent #evals #research-evals #systems-framing
The Decoder
Microsoft introduces MAI-Cyber-1-Flash, a compact security model that scores 96 percent on the CyberGym benchmark when embedded in its MDASH multi-agent system. Microsoft says costs should drop by 50 percent compared to pure frontier models, since only tough cases get passed to GPT-5.4.
#agent #agentic-workflows #evals
arXiv reasoning / agents / evals
AI-driven autonomous research (AR) systems are becoming increasingly effective across a broad range of tasks. Their performance, however, is still evaluated primarily by the quality of the final outcome.
#evals #research-evals #systems-framing
MIT Technology Review AI
Imagine a healthcare system made up of multiple AI agents: one that manages symptom assessment, another scheduling, a third insurance, and a fourth pharmacy. Each is an expert in its domain.
#agent #agentic-workflows
The Decoder
METR's new metric, the "expenditure horizon," puts a dollar figure on how cost-effective AI agents are at solving problems. Early results on the NanoGPT speedrun are underwhelming, the metric has blind spots, and the newest generation of models could change the picture.
#agent #agentic-workflows #evals
arXiv reasoning / agents / evals
Background: Large language models (LLMs) are increasingly used to automate code review, but the reasoning behind their decisions remains hard to understand. Developers struggle to assess the validity of LLM-generated reviews, making it difficult to gauge how much trust to place in them.
#builder #evals #research-evals
TechCrunch AI
Microsoft bolstered its AI cybersecurity offerings this week with the launch of its first AI security model and a new security platform.
#agent #agentic-workflows #systems-framing
The Decoder
Moonshot AI has released Kimi K3's model weights and made parts of its infrastructure open source. The Chinese model nearly matches Western frontier models such as Fable 5 and GPT-5.6 Sol on popular benchmarks, but independent tests found major gaps in cyber and math performance, possibly pointing to distillation.
#evals #models-architecture