Meta Muse Glimmer 30B Review: Benchmarks, VRAM

Meta Muse Glimmer 30B Review: Benchmarks, VRAM

Summary: Muse Glimmer in 30 Seconds Muse Glimmer is a 30 billion parameter agentic model released by Meta Superintelligence Labs on August 10, 2026. The weights are open under Apache 2.0. No cloud needed. Four-bit quantization shrinks the language model to under 20 GB, so it runs on a single consumer GPU with 24 GB or 32 GB of memory. It reads text and images, and returns text. Default context window is 128K tokens, and the model supports more. It clears its size class on agentic tests: 75.5 on MCP Atlas (Gemma4-31B 54.2, Qwen3.6-27B 62.5). On math, 94.7 on AIME 2026. The bundled DFlash drafter speeds generation up 3.1x on an RTX 5090 and 1.8x on an M5 Max. It is distilled from Muse Spark, Meta’s cloud model, so the big model’s agentic behavior is transferred into a small one. On the same day Zuckerberg announced that the weights for Muse Spark 1.2, Meta’s flagship model, will be opened too. No date, just “soon.” An AI agent normally means an API key, an internet connection and a bill that ticks up with every request. Meta just tried the opposite. ...

August 10, 2026 ·  13 min ·  2664 words
What Is Qwen3.8-Max? The AI That Ran Alone for 125 Hours

What Is Qwen3.8-Max? The AI That Ran Alone for 125 Hours

Summary: Qwen3.8-Max in 30 Seconds Qwen3.8-Max is Alibaba’s new flagship model, made generally available on August 2, 2026. 2.4 trillion parameters, 95 billion active (MoE architecture). It reads text, images and video, and returns text. Context window is in the 1 million token class. On one task it ran 125 hours (about 5 days) with no human input, rebuilding an experiment from a machine learning paper from scratch, confirming its six findings, then inventing a method that beats the paper. It beats Claude Opus 4.8 on most coding and agent tests, trades blows with Claude Fable 5 and GPT-5.6 Sol, and falls behind on some. API pricing is $2 input / $6 output per million tokens. Repeated input costs $0.25. This is the first time Alibaba has open-weighted a Max-class model. The weights landed on Hugging Face on August 12, 2026, though under Alibaba’s own Qwen3.8-Max license rather than Apache 2.0. Two days later, on August 14, Qwen3.8-27B followed: a dense 27B model under Apache 2.0 that fits on a single GPU. Ask an AI to “rebuild the experiment in this paper, then improve on it” and it normally stalls after a few turns, waiting for you to step in and steer. ...

August 3, 2026 ·  Updated: August 16, 2026 ·  19 min ·  3926 words
Gemini 3.6 Flash and 3.5 Flash-Lite Introduced

Gemini 3.6 Flash and 3.5 Flash-Lite Introduced

On July 21, 2026, Google announced three new Gemini models at once: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and the defense-focused Gemini 3.5 Flash Cyber. The theme of the announcement is clear: not higher scores, but getting the same job done with fewer tokens. As agentic workflows move into production, the bill is driven less by how smart a model is and more by how many tokens that intelligence burns. Google aimed straight at that. Here are the details. ...

July 21, 2026 ·  9 min ·  1754 words
Claude Opus 4.8 Released: More Honest and Capable Than Ever!

Claude Opus 4.8 Released: More Honest and Capable Than Ever!

Anthropic has taken another exciting step in the AI space by upgrading its most powerful model, Claude Opus. Meet Claude Opus 4.8! Built on the foundations of Opus 4.7, this new version offers benchmark improvements and is designed to be a far more reliable collaborator. Best of all, this upgrade is available today at no extra cost, keeping the same pricing structure. Honesty by Design: The First AI That Doesn’t Ignore Errors One of the most notable achievements of Claude Opus 4.8 is its progress on AI hallucinations and overconfidence. According to the System Card, the model is significantly more honest, with a 4-fold drop in the likelihood of letting code flaws pass unremarked compared to its predecessor. ...

May 28, 2026 ·  4 min ·  842 words