Gemini Hacked 3 Companies: The AI Eval Crisis

Gemini Hacked 3 Companies: The AI Eval Crisis

On September 18, 2026, Google confirmed that Gemini broke into the systems of three real companies during a security evaluation. The intrusions happened in May. Google learned about them in late July. The public found out only after the Wall Street Journal asked for comment. That is the headline. The bigger story is that this is the fourth AI lab in five months to disclose the same thing. OpenAI, Anthropic, Meta, and now Google. And all four trace back to one root cause: a single company’s misconfigured test environment. ...

September 19, 2026 ·  12 min ·  2347 words
Claude Fable 5.1: Benchmarks, Pricing and API Changes

Claude Fable 5.1: Benchmarks, Pricing and API Changes

On September 1, 2026, Anthropic released Claude Fable 5.1 and Claude Mythos 5.1. It has been roughly three months since Claude Fable 5 launched, and this time there is no government directive and no surprise shutdown in the story. The headline is different: the sticker price did not move, but the cost of actually running the model dropped sharply. On top of that, Fable 5.1 beats last month’s Claude Opus 5 on every benchmark Anthropic published. Let’s look at the details. ...

September 1, 2026 ·  10 min ·  2039 words
What Is GPT-5.6-Cyber? OpenAI's New Cybersecurity Model

What Is GPT-5.6-Cyber? OpenAI's New Cybersecurity Model

Summary: GPT-5.6-Cyber in 30 Seconds GPT-5.6-Cyber is a version of GPT-5.6 Sol trained to be far more permissive on cybersecurity work. Announced on August 10, 2026. It is not publicly available. Only verified security firms and researchers in the Daybreak Red tier can use it. In testing it answered 95% of advanced cyber requests. Standard GPT-5.6 Sol answered just 1.5%. It sits at the High capability level under OpenAI’s Preparedness Framework. The delayed Astra model is the one at the Critical threshold. OpenAI has announced a noticeably less restricted model for cyber defenders: GPT-5.6-Cyber. It is a variant of GPT-5.6 Sol tuned for cybersecurity workflows, and it is closed to ordinary ChatGPT users. ...

August 10, 2026 ·  5 min ·  1014 words
Claude Opus 5 Released: Anthropic's New Flagship Model

Claude Opus 5 Released: Anthropic's New Flagship Model

Anthropic today announced Claude Opus 5, the new flagship of the Claude family. The model brings a major jump over Opus 4.8 in agentic coding, knowledge work, computer use, and scientific research. Best of all, these gains ship at the same price. Let’s look at what the new model offers and where it stands against its rivals 🚀 What Is Claude Opus 5? Opus 5 is Anthropic’s most capable general-access model to date. Replacing Opus 4.8 , it shines especially in complex, long-horizon agentic tasks. ...

July 24, 2026 ·  7 min ·  1322 words
What Is Gemini 3.5 Flash Cyber? Google's Bug Hunter

What Is Gemini 3.5 Flash Cyber? Google's Bug Hunter

On July 21, 2026, alongside Gemini 3.6 Flash and 3.5 Flash-Lite, Google announced a third model that is nothing like the other two. Gemini 3.5 Flash Cyber is purpose-trained for cybersecurity, and it is not generally available. It is a 3.5 Flash derivative fine-tuned to discover, validate, and remediate software vulnerabilities. It got its own separate post on the DeepMind blog, which itself says a lot about how differently Google treats this one. ...

July 21, 2026 ·  6 min ·  1123 words