DeepSeek V4.1 Flash: Price and Benchmarks

DeepSeek V4.1 Flash: Price and Benchmarks

Summary: DeepSeek V4.1 Flash in 30 Seconds DeepSeek-V4.1-Flash shipped on September 10, 2026. License is MIT, weights are on Hugging Face. 552B backbone parameters, but only 8B active while reading input and 16B while generating output. That asymmetry is the whole point of the model. New architecture: Causal Encoder-Decoder (CED). 40 layers, 20 causal encoder plus 20 decoder. The decoder’s global KV cache is projected from the encoder’s final hidden states instead of from each decoder layer’s own. Result: a global KV cache of 890 bytes per token. One quarter of V4-Flash, and 1/437 of DeepSeek-V1. Natively multimodal: images go through DeepSeek-ViT, a vision encoder trained from scratch, from the very start of language-model pre-training. Pre-training corpus is 45T tokens. Beats Claude Opus 5 on several agentic benchmarks: 90.6 vs 89.1 on Terminal-Bench 2.1, 74.2 vs 74.0 on DeepSWE v1.1, 54.8 vs 50.3 on AutomationBench. Reasoning effort is a continuous 1-100 dial, not an on/off thinking toggle. API pricing is $0.15 in / $0.60 out per million tokens off-peak, double that at peak. Roughly one twentieth of Claude Opus 5’s output price. DeepSeek is retiring V4 Pro: from September 14, 2026, deepseek-v4-pro requests get routed to V4.1 Flash and billed at Flash rates. Everyone in the open-weight race this year is chasing the same two numbers: active parameters and KV cache. The first decides what each token costs, the second decides how much memory a long context eats. GLM-5.3-Flash answered with hybrid attention, Qwen3.8-Flash-Next with 6B active parameters. ...

September 10, 2026 ·  14 min ·  2923 words
What People Built With GPT-6 Astra: 12 Real Runs

What People Built With GPT-6 Astra: 12 Real Runs

Summary: Five Days of Receipts Finished Portal on its own. 3,336 tool calls, roughly 21 hours, a $571.18 token bill. Launched a rocket in Factorio Space Age 2.1. No model had ever pushed past blue science before. Beat Pokémon in 18h 12m. GPT-5.6 Sol needed 96h 35m for the same run. Scored 19 out of 20 on a robot arm, at $0.94 per attempt. Produced 3,295 editable objects in Blender from a single prompt. What ties them together: every one of these has a price tag, and it is not small. The launch write-up covered the benchmark table, the pricing and the access rules. Five days on, the picture has changed: instead of OpenAI’s slides, we can now look at what people actually got the model to do. ...

September 8, 2026 ·  10 min ·  1994 words
How to Use GPT-6 Astra (and Is It Free?)

How to Use GPT-6 Astra (and Is It Free?)

Summary: Where to Find Astra Not free. Free and Go plans do not include Astra; their default model stays GPT-5.6 Luna. On Plus, Astra does not appear in regular Chat. Access arrives through ChatGPT Work and Codex. On Pro, Business and Enterprise, it shows up in the model picker under the Pro option, labelled GPT-6 Pro. The model underneath is Astra. Enterprise workspaces get it off by default at launch. An admin has to enable it. Codex CLI 0.153.0 or newer is required, and the desktop app needs updating too. In the API the model id is gpt-6-astra: 1.05M token context, 128K output, $10 in / $50 out per million tokens. Our GPT-6 Astra launch write-up covered what the model is, the benchmark table and the price tag. Two days later the question landing in the inbox is a different one: “I pay for Plus, there is no Astra anywhere in my model picker, where is it?” ...

September 5, 2026 ·  12 min ·  2465 words
GPT-6 Astra: Price, Benchmarks, Access

GPT-6 Astra: Price, Benchmarks, Access

GPT-6 Astra in 30 Seconds GPT-6 Astra is OpenAI’s new flagship, announced on September 3, 2026. The company calls it “the world’s most intelligent and aligned model”. There is no Sol/Terra/Luna split this time. The lineup is Astra and Astra Pro. API pricing is $10 per million input tokens and $50 per million output tokens: 2.5x GPT-5.6 Sol’s promotional price, and identical to Claude Fable 5.1. The scores are high but footnoted. The headline 98.6% on ARC-AGI-3 came from a custom harness; the same model scores 62.7% on the standard one. Astra is the first OpenAI model to cross the Critical cybersecurity threshold in the Preparedness Framework. Standard access refuses parts of that work outright. Rollout is staged: Daybreak enterprise customers first, then Plus, Pro, Business, Enterprise, the API and AWS. Pro, Business and Enterprise also get Astra Pro. OpenAI launched GPT-6 Astra today, September 3, 2026. At the press briefing, president Greg Brockman first conceded that AGI remains a “gray, fuzzy thing”, then went ahead anyway: “I think it’s not unreasonable to feel that we are now in the AGI era.” He closed with the same line: “Welcome to the AGI era.” ...

September 3, 2026 ·  13 min ·  2731 words
Gemini 3.8 Flash: Benchmarks, Price, Cyber

Gemini 3.8 Flash: Benchmarks, Price, Cyber

TL;DR: Gemini 3.8 Flash in 30 seconds Gemini 3.8 Flash landed on September 2, 2026, three weeks after 3.7 Flash. That is the third Flash release in three months. Price did not move: $0.75 input, $3.75 output per million tokens. The promo ends December 31, 2026, then it doubles. A second model shipped alongside it: Gemini 3.8 Flash Cyber, tuned for vulnerability discovery and locked behind the new Fairwind Program. The benchmark table is split. It tops the chart on finance, legal, long video and chart reasoning, and it trails Claude Opus 5 badly on long horizon terminal and computer use work. The model “works harder”: more reasoning steps, more iterative tool calls. Same sticker price, potentially a bigger bill. The Flash release cadence has stopped being funny. 3.6 Flash shipped on July 21, 3.7 Flash on August 13, and Gemini 3.8 Flash today, September 2, 2026. Three releases in three months. ...

September 2, 2026 ·  10 min ·  2046 words