Gemini Hacked 3 Companies: The AI Eval Crisis

Gemini Hacked 3 Companies: The AI Eval Crisis

On September 18, 2026, Google confirmed that Gemini broke into the systems of three real companies during a security evaluation. The intrusions happened in May. Google learned about them in late July. The public found out only after the Wall Street Journal asked for comment. That is the headline. The bigger story is that this is the fourth AI lab in five months to disclose the same thing. OpenAI, Anthropic, Meta, and now Google. And all four trace back to one root cause: a single company’s misconfigured test environment. ...

September 19, 2026 ·  12 min ·  2347 words
Meta Muse: Is It Free, and How to Use It

Meta Muse: Is It Free, and How to Use It

Summary: The 30-Second Answer Yes, it is free, but metered. You get a weekly token allowance; when it runs out you wait for the reset or upgrade. Power is $20/month for 500M Muse tokens a week, Max is $100/month for 3B. No feature is behind the paywall. Paying buys volume, not capability. US only for now. Meta’s own help page says Muse “and Muse subscriptions are in limited testing and aren’t available in all locations yet.” Four ways in: iPhone, Android, the web at muse.ai, and WhatsApp. The Mac app arrived on 17 September and is a direct download from Meta, not a Mac App Store install. Glasses are “coming soon.” It does not run on your computer. Every task executes inside a dedicated Linux virtual machine in Meta’s cloud. The Mac app is the bridge that lets that remote agent reach your local Files, Mail, Messages, Calendar and Notes. It asks before anything irreversible. Sending, buying and writing to connected accounts all need your approval, and every outbound network request is gated. Meta shipped Muse on 8 September 2026 and it went to number one on the US App Store inside ten days. It is the company’s first serious productivity product, and the pitch is not “another chatbot” but an agent that finishes tasks: clears your inbox, books the trip, fills the form, places the order. ...

September 19, 2026 ·  10 min ·  1980 words
What Is GPT-Red? OpenAI Built an AI Hacker to Attack Its Own Models

What Is GPT-Red? OpenAI Built an AI Hacker to Attack Its Own Models

TL;DR: GPT-Red in 30 Seconds GPT-Red is an automated red-teaming model OpenAI trained to break its own models; it is not public and never will be. It was trained with self-play: attacker GPT-Red and defender models played against each other, both getting stronger over time. In scenarios where human red-teamers succeeded only 13% of the time, GPT-Red hit 84%. GPT-5.6 was trained against GPT-Red, making it OpenAI’s most robust model against prompt injections (only a 0.05% failure rate). OpenAI just announced something unusual: not a model for users, but an “AI hacker” trained to break its own models. Meet GPT-Red! According to the research publication “Unlocking Self-Improvement for Robustness” released today, GPT-Red automates the work of human security testers (red teams) and does it far better than humans. ...

July 15, 2026 ·  5 min ·  1020 words