← Back to archive
Signals for 2026-07-29
Published 2026-07-29T08:16+02:00
10 geselecteerde signalen uit de lokale hybride Daily Signal Brief pipeline.
Simon Willison
Hugging Face just released this extremely detailed technical description of OpenAI's recent accidental cyberattack against their infrastructure . This attack was very sophisticated, and the resulting document doubles as a crash-course in modern adversarial security approaches.
#agent #agentic-workflows #implementation
arXiv reasoning / agents / evals
Recently, memory management has become a key infrastructure for LLM-based agents, as it directly affects long-horizon reasoning, personalized responses, and knowledge reuse. However, existing LLM memory systems typically adopt a coarse-grained (utility-agnostic) manner that treats heterogeneous user-LLM interaction records uniformly, leading to redundant and low-impact records persisting in the memory repository.
#agent #evals #research-evals
arXiv reasoning / agents / evals
Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success or single-frame grounding.
#agent #evals #research-evals
arXiv reasoning / agents / evals
Tabular Foundation Models (TFMs) have emerged as novel approaches for tabular predictive tasks, demonstrating competitive predictive performance to ensemble tree-based models. Most TFMs are trained and evaluated on independent and identically distributed data, but this assumption changes in real-world scenarios due to distribution shifts, which compromise the robustness of models.
#evals #research-evals
Hugging Face Blog
The OlmoEarth Platform: Geospatial inference at planetary scale
#research-evals #systems-framing
Simon Willison
The best part of this article (here's the repo ) about how Anthropic researchers used Claude Mythos to find mathematical flaws in both HAWK and a weaker version of AES ("neither of these results has a practical impact on today’s computer systems") is the prompts that they shared, spelling mistakes included: the models tend to think it is impossible to solve so they don't try they need a good amount of prompting.
#evals #models-architecture
TechCrunch AI
The deal is Cyera's third acquisition this year.
#agent #agentic-workflows
The Decoder
Anthropic's Claude Mythos Preview found weaknesses in key cryptographic algorithms, including a better attack on HAWK, a post-quantum signature scheme that human experts had reviewed for more than two years. The model found it in just 60 hours at an API cost of about $100,000.
#builder #tooling-runtime
Google News AI Adoption
Cognizant launches EMEA AI Unit to help enterprises scale agentic AI adoption Cognizant Technology Solutions
#agent #agentic-workflows #implementation
Hugging Face Blog
LFM2.5-Encoders for Fast Long-Context Inference on CPU
#models-architecture