Qwen3.8-Flash-Next: 125B MoE, 6B Active Params

Qwen3.8-Flash-Next: 125B MoE, 6B Active Params

Summary: Qwen3.8-Flash-Next in 30 Seconds Qwen3.8-Flash-Next is Alibaba’s new AI model, out in August 2026. It reads text and images and writes text back, tuned for writing code and running multi-step work on your behalf. Anyone can download the model files, live since 24 August. The license is not fully permissive though, so read it before shipping it in a product. The real story: Alibaba shipped this as a dry run for the next big release, Qwen4. The identifier inside the model files literally says qwen4_exp, as in “Qwen4 experimental”. It is 125B parameters in size, but only 6B of them run for any given word. Big-model knowledge, small-model bill. It reads about 750,000 words in one go (1M tokens), and at that length it is 8x faster than its much larger sibling. Training it cost roughly 1/9 of the 397B Qwen3.7-Plus, and it still beats that model on coding and office work. It is cheap to run: $0.16 in / $0.47 out per million tokens. The flagship in the same family costs 12x that. When Alibaba shipped Qwen3-Next, the pitch was: this is not a finished product, it is next generation’s architecture released early so the community can poke at it. That architecture then carried the whole Qwen3.5 through Qwen3.8 line. ...

August 26, 2026 ·  10 min ·  2130 words
What Is Qwen3.8-Max? The AI That Ran Alone for 125 Hours

What Is Qwen3.8-Max? The AI That Ran Alone for 125 Hours

Summary: Qwen3.8-Max in 30 Seconds Qwen3.8-Max is Alibaba’s new flagship model, made generally available on August 2, 2026. 2.4 trillion parameters, 95 billion active (MoE architecture). It reads text, images and video, and returns text. Context window is in the 1 million token class. On one task it ran 125 hours (about 5 days) with no human input, rebuilding an experiment from a machine learning paper from scratch, confirming its six findings, then inventing a method that beats the paper. It beats Claude Opus 4.8 on most coding and agent tests, trades blows with Claude Fable 5 and GPT-5.6 Sol, and falls behind on some. API pricing is $2 input / $6 output per million tokens. Repeated input costs $0.25. This is the first time Alibaba has open-weighted a Max-class model. The weights landed on Hugging Face on August 12, 2026, though under Alibaba’s own Qwen3.8-Max license rather than Apache 2.0. Two days later, on August 14, Qwen3.8-27B followed: a dense 27B model under Apache 2.0 that fits on a single GPU. Ask an AI to “rebuild the experiment in this paper, then improve on it” and it normally stalls after a few turns, waiting for you to step in and steer. ...

August 3, 2026 ·  Updated: August 16, 2026 ·  19 min ·  3926 words
Wan Streamer: Real-Time AI Video Interaction

Wan Streamer: Real-Time AI Video Interaction

Are you ready to meet the video assistants of the future? Until today, when we talked about AI “video calls,” clunky, cascaded systems came to mind. First, the audio was listened to, then transcribed to text, a response was generated, and finally, a video animation was rendered. This delayed architecture is now history. Wan-Streamer is the world’s first native-streaming, end-to-end AI model. By processing language, audio, and video simultaneously within a single model, it offers a truly full-duplex video call experience. ...

June 26, 2026 ·  2 min ·  380 words
Qwen3.5 Released! Native Multimodal AI Performance

Qwen3.5 Released! Native Multimodal AI Performance

Taking a closer look at the Qwen3.5 model, which is reshuffling the deck in the artificial intelligence world. Focusing heavily on increasing the capacities of foundation models in recent months, Alibaba Cloud officially released Qwen3.5 on February 16, 2026. They have genuinely showcased an ambitious stride in the race of large language models. Garnering attention especially with its native multimodal agent capabilities and efficiency-focused architecture, this version goes head-to-head with tech giants like GPT-5.2 and Claude 4.5 Opus. So, what exactly does Qwen3.5 promise, when did it come out, and why is it so vital for developers? Let’s dive into the details together. 👇🏻 ...

February 21, 2026 ·  6 min ·  1097 words