← Back to archive
Signals for 2026-08-01
Published 2026-08-01T08:15+02:00
10 geselecteerde signalen uit de lokale hybride Daily Signal Brief pipeline.
Simon Willison
I've been working with Jesse Vincent's Prime Radiant applied AI research lab building out this evals framework to help answer questions about the capabilities of different models. The result is smevals , a new tool for running small eval suites across different model configurations and grading the results.
#agent #evals #implementation #research-evals
Simon Willison
The latest release in DeepSeek's V4 family, "with substantially enhanced agentic capabilities". It's 304 billion parameters - 167GB on Hugging Face - but it appears to punch well above its weight.
#agent #agentic-workflows #evals
arXiv reasoning / agents / evals
Deploying autonomous computer-use agents (CUAs) locally is increasingly important for privacy, cost efficiency, and practical usability, yet improving their performance under strict hardware constraints remains challenging.
#agent #evals #research-evals
arXiv reasoning / agents / evals
Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident began. We introduce ORCA-bench, a benchmark that puts general-purpose coding agents in a production-fidelity oncall setting.
#agent #evals #implementation #research-evals
arXiv reasoning / agents / evals
Metallic magnets exhibit complex spin dynamics governed by electronically generated interactions. Predictive simulations of such dynamics typically require repeated solutions of an underlying electronic problem throughout the time evolution, creating a major computational bottleneck.
#evals #research-evals
The Decoder
Thinking Machines, the AI lab from former OpenAI CTO Mira Murati, has released Inkling Small. The open-weights reasoning model is less than a third the size of Inkling but beats it on several coding and reasoning benchmarks.
#evals #research-evals
Simon Willison
On Monday Bryan Cantrill and Adam Leventhal invited me to join their podcast to talk about the wild week we've had - with Kimi K3 showing open weight models can stand toe-to-toe with proprietary frontier ones, accidental cybersecurity attacks , and public letters about Open Weights and American AI Leadership signed by almost every big name in AI (with one notable exception ).
#evals #models-architecture
TechCrunch AI
OpenAI has reportedly found evidence of additional agent misbehavior as it looks into the incident that occurred with Hugging Face.
#agent #agentic-workflows
Google News AI Lab Watch
AI agents with $3,000 budget flunk open-ended AI research assignment R&D World
#agent #agentic-workflows #evals
Google News AI Adoption
AI Automation Platforms: Bennu Agent Automates DevOps Workflows With Trusted AI Execution Trend Hunter
#agent #agentic-workflows #systems-framing