Investigating three real-world incidents in our cybersecurity evaluations
Simon Willison
It happened again! This is turning into something of a pattern.
10 geselecteerde signalen uit de lokale hybride Daily Signal Brief pipeline.
Simon Willison
It happened again! This is turning into something of a pattern.
arXiv reasoning / agents / evals
Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability.
Simon Willison
Huge price drop from OpenAI today: GPT-5.6 Terra got a 20% reduction, and GPT-5.6 Luna got a massive 80% drop.
arXiv reasoning / agents / evals
The rapid growth of large language models (LLMs) has resurrected age-old questions in sociolinguistics and world Englishes, such as who decides what counts as legitimate English, whose English is suspect etc.
arXiv reasoning / agents / evals
SWE-bench-like benchmarks are widely used for evaluating LLM's issue resolution capability. They typically follow a common construction pipeline: each PR (Pull Request) is paired with its linked issue by extracting issue references from the PR description; the issue description is used as the problem statement, and the PR patch serves as the test oracle.
TechCrunch AI
The deal gives Okta identity threat detection capabilities as enterprises seek to secure AI agents and other non-human identities across cloud environments.
MIT Technology Review AI
It is impossible to make large language models fully secure against hacks because of a fundamental flaw in how they work, a team of researchers argue in a paper presented at the International Conference on Machine Learning, a top AI conference, this month. The claim has huge implications for the safety of this technology, which…
The Decoder
Can language models spark a scientific revolution? In a position paper titled "LLMs can't jump," Google Deepmind's Tom Zahavy argues they can't.
The Decoder
Pangram 4 detects 99.66 percent of AI-generated text with just one false positive per 24,000 documents, the company claims. The model also resists "humanizer" tools that disguise AI writing as human.
TechCrunch AI
A new study estimates only 2,000 U.S. engineers have the expertise to deliver meaningful AI ROI, as enterprises race to hire forward-deployed engineers to implement AI at scale.