← Back to archive

Signals for 2026-08-01

Published 2026-08-01T08:15+02:00

10 geselecteerde signalen uit de lokale hybride Daily Signal Brief pipeline.

smevals - a small eval suite for evaluating models, prompts, and harnesses

Simon Willison

I've been working with Jesse Vincent's Prime Radiant applied AI research lab building out this evals framework to help answer questions about the capabilities of different models. The result is smevals , a new tool for running small eval suites across different model configurations and grading the results.

#agent #evals #implementation #research-evals

deepseek-ai/DeepSeek-V4-Flash-0731

Simon Willison

The latest release in DeepSeek's V4 family, "with substantially enhanced agentic capabilities". It's 304 billion parameters - 167GB on Hugging Face - but it appears to punch well above its weight.

#agent #agentic-workflows #evals

Rethinking Inference-Time Scaling in Local Computer-Use Agents: Failure Modes and Compute Tradeoffs

arXiv reasoning / agents / evals

Deploying autonomous computer-use agents (CUAs) locally is increasingly important for privacy, cost efficiency, and practical usability, yet improving their performance under strict hardware constraints remains challenging.

#agent #evals #research-evals

ORCA-bench: How Ready Are Language Model Agents for Oncall?

arXiv reasoning / agents / evals

Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident began. We introduce ORCA-bench, a benchmark that puts general-purpose coding agents in a production-fidelity oncall setting.

#agent #evals #implementation #research-evals

Graph Neural Network Force Fields for Spin Dynamics in Metallic Magnets

arXiv reasoning / agents / evals

Metallic magnets exhibit complex spin dynamics governed by electronically generated interactions. Predictive simulations of such dynamics typically require repeated solutions of an underlying electronic problem throughout the time evolution, creating a major computational bottleneck.

#evals #research-evals

Thinking Machines bets on efficiency over size with its second model, Inkling Small

The Decoder

Thinking Machines, the AI lab from former OpenAI CTO Mira Murati, has released Inkling Small. The open-weights reasoning model is less than a third the size of Inkling but beats it on several coding and reasoning benchmarks.

#evals #research-evals

Oxide and Friends: The Open Weight Revolution with Simon Willison

Simon Willison

On Monday Bryan Cantrill and Adam Leventhal invited me to join their podcast to talk about the wild week we've had - with Kimi K3 showing open weight models can stand toe-to-toe with proprietary frontier ones, accidental cybersecurity attacks , and public letters about Open Weights and American AI Leadership signed by almost every big name in AI (with one notable exception ).

#evals #models-architecture

OpenAI reportedly finds evidence that more of its agents ran amok

TechCrunch AI

OpenAI has reportedly found evidence of additional agent misbehavior as it looks into the incident that occurred with Hugging Face.

#agent #agentic-workflows

AI agents with $3,000 budget flunk open-ended AI research assignment - R&D World

Google News AI Lab Watch

AI agents with $3,000 budget flunk open-ended AI research assignment R&D World

#agent #agentic-workflows #evals

AI Automation Platforms: Bennu Agent Automates DevOps Workflows With Trusted AI Execution - Trend Hunter

Google News AI Adoption

AI Automation Platforms: Bennu Agent Automates DevOps Workflows With Trusted AI Execution Trend Hunter

#agent #agentic-workflows #systems-framing