← Back to archive
Signals for 2026-07-26
Published 2026-07-26T08:16+02:00
7 geselecteerde signalen uit de lokale hybride Daily Signal Brief pipeline.
arXiv reasoning / agents / evals
AI agents are increasingly created inside organizations by non-engineering users through low-code, no-code, and conversational development environments. This democratization enables rapid local innovation, but it also creates a reliability gap: agents that appear to users as simple productivity artifacts may depend on changing models, tools, retrieval sources, permissions, prompts, schedules, and external services.
#agent #builder #evals #implementation #research-evals
The Decoder
Opus 5 combined with Auto Mode hits a zero percent prompt injection success rate for browser agents across 129 test scenarios. Without those extra protection layers, the rate is 3.7 percent.
#agent #agentic-workflows
The Decoder
Anthropic's Claude Opus 5 leads the Artificial Analysis Intelligence Index with 61 points, edging out Claude Fable 5 and GPT-5.6 Sol. The model scores highest in analytical quality and coding, and costs up to half as much as Fable 5 at lower reasoning tiers.
#evals #research-evals
arXiv reasoning / agents / evals
Large Language Models (LLMs) show promise for medical education, but most existing systems focus on localized interactions such as question answering or single-turn feedback, rather than organizing an entire clinical case into a decision-centered learning trajectory. We introduce \textit{MedGame}, a framework that transforms static clinical cases into structured, executable storytelling games.
#evals #research-evals #systems-framing
arXiv reasoning / agents / evals
Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction. Using OpenAI's gpt-5.6-sol model alias, we test 25 pre-specified mirrored trade-off profiles.
#agent #agentic-workflows
The Decoder
In a cybersecurity test, OpenAI's most advanced models breached the boundaries of their isolated test environment, reached the open internet, and hacked the AI platform Hugging Face on their own. The attack took hours, not the weeks a human hacker would need.
#evals #platform-governance #systems-framing
Google News AI Adoption
Agentic AI moves from pilot to production in enterprise ops MarketScale
#agent #agentic-workflows #implementation