← Back to archive

Signals for 2026-07-28

Published 2026-07-28T08:17+02:00

10 geselecteerde signalen uit de lokale hybride Daily Signal Brief pipeline.

Building the enterprise environment for agentic AI

MIT Technology Review AI

For the enterprise, the promise of agentic AI is much more than just a better chatbot. It is software agents that execute business tasks end-to-end across people, business workflows, data, and systems.

#agent #agentic-workflows #implementation #systems-framing

An opinionated guide to which AI to use to do stuff

Simon Willison

It's interesting watching the evolution of Ethan Mollick's guide over time. A year ago it was still all about chat - ChatGPT, Claude, Gemini - with o3, Claude 4 Opus, and Gemini 2.5 Pro as the models and Deep Research as a useful alternative mode.

#agent #agentic-workflows #builder #evals

SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents

arXiv reasoning / agents / evals

Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events. However, expert-centered warning workflows are costly, labor-intensive, and difficult to scale throughout the warning-to-action process.

#agent #evals #research-evals #systems-framing

Microsoft launches its own cybersecurity model MAI-Cyber-1-Flash but still depends on OpenAI for the toughest tasks

The Decoder

Microsoft introduces MAI-Cyber-1-Flash, a compact security model that scores 96 percent on the CyberGym benchmark when embedded in its MDASH multi-agent system. Microsoft says costs should drop by 50 percent compared to pure frontier models, since only tough cases get passed to GPT-5.4.

#agent #agentic-workflows #evals

Efficiency Matters in Autonomous Research

arXiv reasoning / agents / evals

AI-driven autonomous research (AR) systems are becoming increasingly effective across a broad range of tasks. Their performance, however, is still evaluated primarily by the quality of the final outcome.

#evals #research-evals #systems-framing

The path to artificial superintelligence

MIT Technology Review AI

Imagine a healthcare system made up of multiple AI agents: one that manages symptom assessment, another scheduling, a third insurance, and a fourth pharmacy. Each is an expert in its domain.

#agent #agentic-workflows

METR introduces a new metric to calculate exactly when AI agents become more expensive than humans

The Decoder

METR's new metric, the "expenditure horizon," puts a dollar figure on how cost-effective AI agents are at solving problems. Early results on the NanoGPT speedrun are underwhelming, the metric has blind spots, and the newest generation of models could change the picture.

#agent #agentic-workflows #evals

Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code Review

arXiv reasoning / agents / evals

Background: Large language models (LLMs) are increasingly used to automate code review, but the reasoning behind their decisions remains hard to understand. Developers struggle to assess the validity of LLM-generated reviews, making it difficult to gauge how much trust to place in them.

#builder #evals #research-evals

Microsoft launches its first cybersecurity model, plus a new agentic cybersecurity system

TechCrunch AI

Microsoft bolstered its AI cybersecurity offerings this week with the launch of its first AI security model and a new security platform.

#agent #agentic-workflows #systems-framing

Moonshot AI releases Kimi K3 open weights and infrastructure after shaking up the frontier model race

The Decoder

Moonshot AI has released Kimi K3's model weights and made parts of its infrastructure open source. The Chinese model nearly matches Western frontier models such as Fable 5 and GPT-5.6 Sol on popular benchmarks, but independent tests found major gaps in cyber and math performance, possibly pointing to distillation.

#evals #models-architecture