← Back to archive

Signals for 2026-07-29

Published 2026-07-29T08:16+02:00

10 geselecteerde signalen uit de lokale hybride Daily Signal Brief pipeline.

Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident

Simon Willison

Hugging Face just released this extremely detailed technical description of OpenAI's recent accidental cyberattack against their infrastructure . This attack was very sophisticated, and the resulting document doubles as a crash-course in modern adversarial security approaches.

#agent #agentic-workflows #implementation

MemLens: A Value-Aware Memory Management System with Interactive Analytics for LLM-based Agents

arXiv reasoning / agents / evals

Recently, memory management has become a key infrastructure for LLM-based agents, as it directly affects long-horizon reasoning, personalized responses, and knowledge reuse. However, existing LLM memory systems typically adopt a coarse-grained (utility-agnostic) manner that treats heterogeneous user-LLM interaction records uniformly, leading to redundant and low-impact records persisting in the memory repository.

#agent #evals #research-evals

Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?

arXiv reasoning / agents / evals

Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success or single-frame grounding.

#agent #evals #research-evals

Empirical Evaluation of Out-Of-Distribution Performance of Tabular Foundation Models

arXiv reasoning / agents / evals

Tabular Foundation Models (TFMs) have emerged as novel approaches for tabular predictive tasks, demonstrating competitive predictive performance to ensemble tree-based models. Most TFMs are trained and evaluated on independent and identically distributed data, but this assumption changes in real-world scenarios due to distribution shifts, which compromise the robustness of models.

#evals #research-evals

The OlmoEarth Platform: Geospatial inference at planetary scale

Hugging Face Blog

The OlmoEarth Platform: Geospatial inference at planetary scale

#research-evals #systems-framing

Discovering cryptographic weaknesses with Claude

Simon Willison

The best part of this article (here's the repo ) about how Anthropic researchers used Claude Mythos to find mathematical flaws in both HAWK and a weaker version of AES ("neither of these results has a practical impact on today’s computer systems") is the prompts that they shared, spelling mistakes included: the models tend to think it is impossible to solve so they don't try they need a good amount of prompting.

#evals #models-architecture

Cyera agrees to acquire Oasis Security for $1B to safeguard proliferating AI agents

TechCrunch AI

The deal is Cyera's third acquisition this year.

#agent #agentic-workflows

Anthropic says its Mythos model found vulnerabilities in cryptographic algorithms that secure the internet

The Decoder

Anthropic's Claude Mythos Preview found weaknesses in key cryptographic algorithms, including a better attack on HAWK, a post-quantum signature scheme that human experts had reviewed for more than two years. The model found it in just 60 hours at an API cost of about $100,000.

#builder #tooling-runtime

Cognizant launches EMEA AI Unit to help enterprises scale agentic AI adoption - Cognizant Technology Solutions

Google News AI Adoption

Cognizant launches EMEA AI Unit to help enterprises scale agentic AI adoption Cognizant Technology Solutions

#agent #agentic-workflows #implementation

LFM2.5-Encoders for Fast Long-Context Inference on CPU

Hugging Face Blog

LFM2.5-Encoders for Fast Long-Context Inference on CPU

#models-architecture