Gemini 3.8 Flash: Benchmarks, Price, Cyber

Gemini 3.8 Flash: Benchmarks, Price, Cyber

TL;DR: Gemini 3.8 Flash in 30 seconds Gemini 3.8 Flash landed on September 2, 2026, three weeks after 3.7 Flash. That is the third Flash release in three months. Price did not move: $0.75 input, $3.75 output per million tokens. The promo ends December 31, 2026, then it doubles. A second model shipped alongside it: Gemini 3.8 Flash Cyber, tuned for vulnerability discovery and locked behind the new Fairwind Program. The benchmark table is split. It tops the chart on finance, legal, long video and chart reasoning, and it trails Claude Opus 5 badly on long horizon terminal and computer use work. The model “works harder”: more reasoning steps, more iterative tool calls. Same sticker price, potentially a bigger bill. The Flash release cadence has stopped being funny. 3.6 Flash shipped on July 21, 3.7 Flash on August 13, and Gemini 3.8 Flash today, September 2, 2026. Three releases in three months. ...

September 2, 2026 ·  10 min ·  2046 words
Claude Fable 5.1: Benchmarks, Pricing and API Changes

Claude Fable 5.1: Benchmarks, Pricing and API Changes

On September 1, 2026, Anthropic released Claude Fable 5.1 and Claude Mythos 5.1. It has been roughly three months since Claude Fable 5 launched, and this time there is no government directive and no surprise shutdown in the story. The headline is different: the sticker price did not move, but the cost of actually running the model dropped sharply. On top of that, Fable 5.1 beats last month’s Claude Opus 5 on every benchmark Anthropic published. Let’s look at the details. ...

September 1, 2026 ·  10 min ·  2039 words
Gemini Omni 1.1 Flash: 40s Video, 4K Upscale

Gemini Omni 1.1 Flash: 40s Video, 4K Upscale

Summary: Gemini Omni 1.1 Flash in 30 Seconds Gemini Omni 1.1 Flash is Google’s video generation model, released on 27 August 2026. It replaces Veo 3.1 on the video side. A single generation is still 10 seconds. The headline “40 seconds” is a cumulative length you reach by extending the same video in 10-second steps. When extending, the model now reads up to 10 seconds of prior footage. Previous models only referenced the final second, which is where the consistency gain comes from. You can pin the first and last frame and let the model generate the motion between them. Built for orbit shots, transitions and looping clips. The 360p draft tier runs up to 60% faster at one third of the cost of 720p. Google’s recommended loop: draft at 360p, render the pick at 720p, upscale to 1080p or 4K at the end. 4K is an upscale, not native generation. The model generates 720p; anything above that is an enlargement. Per second: 360p $0.03 · 720p $0.10 · 1080p $0.15 · 4K $0.30. In tokens, $1.50 per million in and $17.50 per million for video out. The API has real gaps: no system instructions, no function calling, no structured output, no context caching. The problem with Google’s video models for the past year was length. An 8-10 second clip is impressive but it isn’t a scene. Gemini Omni 1.1 Flash doesn’t fix that by making the clip longer, it fixes it by making clips stackable. ...

August 29, 2026 ·  10 min ·  2031 words
GLM-5.3-Flash: 320B MoE, 18B Active, MIT

GLM-5.3-Flash: 320B MoE, 18B Active, MIT

Summary: GLM-5.3-Flash in 30 Seconds GLM-5.3-Flash is Z.ai’s (formerly Zhipu AI) new model, released 26 August 2026. It is the first natively multimodal member of the GLM-5 series: text and images go through the same model. It is 320 billion parameters, but only 18 billion run per token. Layer count is roughly half of GLM-4.5’s: 45 against 92. The licence is MIT. Most strong Chinese models this summer shipped under bespoke community licences; here there is no fine print to read before shipping a product. It beats GLM-5.2 by a wide margin on coding and agentic tests (63.4 against 46.2 on DeepSWE) and approaches Claude Opus 4.8 overall, at roughly one tenth of GLM-5.2’s price. API pricing is $0.15 in / $0.50 out per million tokens, or $0.075 and $0.25 with the 50% discount running until 9 September 2026. It is the first model in the series to use hybrid attention: linear attention carries local dependencies, sparse attention retrieves distant context. Against GLM-5.3 that is 3x less attention compute and a 4.4x smaller KV cache. Before launch it was tested anonymously as ox-alpha on OpenCode and OpenRouter, where it became the most used model of the week. All of that traffic was served on Chinese AI chips. There is one race in open-weight models this summer: producing the same intelligence with less compute. Alibaba’s Qwen3.8-Flash-Next beat its own 397B sibling with 6 billion active parameters. Z.ai’s answer is GLM-5.3-Flash: 320 billion total parameters with only 18 billion running per token, leaving GLM-5.2 behind at a tenth of the cost. ...

August 26, 2026 ·  14 min ·  2875 words
Qwen3.8-Flash-Next: 125B MoE, 6B Active Params

Qwen3.8-Flash-Next: 125B MoE, 6B Active Params

Summary: Qwen3.8-Flash-Next in 30 Seconds Qwen3.8-Flash-Next is Alibaba’s new AI model, out in August 2026. It reads text and images and writes text back, tuned for writing code and running multi-step work on your behalf. Anyone can download the model files, live since 24 August. The license is not fully permissive though, so read it before shipping it in a product. The real story: Alibaba shipped this as a dry run for the next big release, Qwen4. The identifier inside the model files literally says qwen4_exp, as in “Qwen4 experimental”. It is 125B parameters in size, but only 6B of them run for any given word. Big-model knowledge, small-model bill. It reads about 750,000 words in one go (1M tokens), and at that length it is 8x faster than its much larger sibling. Training it cost roughly 1/9 of the 397B Qwen3.7-Plus, and it still beats that model on coding and office work. It is cheap to run: $0.16 in / $0.47 out per million tokens. The flagship in the same family costs 12x that. When Alibaba shipped Qwen3-Next, the pitch was: this is not a finished product, it is next generation’s architecture released early so the community can poke at it. That architecture then carried the whole Qwen3.5 through Qwen3.8 line. ...

August 26, 2026 ·  10 min ·  2130 words