GLM-5.3-Flash: 320B MoE, 18B Active, MIT

GLM-5.3-Flash: 320B MoE, 18B Active, MIT

Summary: GLM-5.3-Flash in 30 Seconds GLM-5.3-Flash is Z.ai’s (formerly Zhipu AI) new model, released 26 August 2026. It is the first natively multimodal member of the GLM-5 series: text and images go through the same model. It is 320 billion parameters, but only 18 billion run per token. Layer count is roughly half of GLM-4.5’s: 45 against 92. The licence is MIT. Most strong Chinese models this summer shipped under bespoke community licences; here there is no fine print to read before shipping a product. It beats GLM-5.2 by a wide margin on coding and agentic tests (63.4 against 46.2 on DeepSWE) and approaches Claude Opus 4.8 overall, at roughly one tenth of GLM-5.2’s price. API pricing is $0.15 in / $0.50 out per million tokens, or $0.075 and $0.25 with the 50% discount running until 9 September 2026. It is the first model in the series to use hybrid attention: linear attention carries local dependencies, sparse attention retrieves distant context. Against GLM-5.3 that is 3x less attention compute and a 4.4x smaller KV cache. Before launch it was tested anonymously as ox-alpha on OpenCode and OpenRouter, where it became the most used model of the week. All of that traffic was served on Chinese AI chips. There is one race in open-weight models this summer: producing the same intelligence with less compute. Alibaba’s Qwen3.8-Flash-Next beat its own 397B sibling with 6 billion active parameters. Z.ai’s answer is GLM-5.3-Flash: 320 billion total parameters with only 18 billion running per token, leaving GLM-5.2 behind at a tenth of the cost. ...

August 26, 2026 ·  14 min ·  2876 words
What Is Qwen3.8-Max? The AI That Ran Alone for 125 Hours

What Is Qwen3.8-Max? The AI That Ran Alone for 125 Hours

Summary: Qwen3.8-Max in 30 Seconds Qwen3.8-Max is Alibaba’s new flagship model, made generally available on August 2, 2026. 2.4 trillion parameters, 95 billion active (MoE architecture). It reads text, images and video, and returns text. Context window is in the 1 million token class. On one task it ran 125 hours (about 5 days) with no human input, rebuilding an experiment from a machine learning paper from scratch, confirming its six findings, then inventing a method that beats the paper. It beats Claude Opus 4.8 on most coding and agent tests, trades blows with Claude Fable 5 and GPT-5.6 Sol, and falls behind on some. API pricing is $2 input / $6 output per million tokens. Repeated input costs $0.25. This is the first time Alibaba has open-weighted a Max-class model. The weights landed on Hugging Face on August 12, 2026, though under Alibaba’s own Qwen3.8-Max license rather than Apache 2.0. Two days later, on August 14, Qwen3.8-27B followed: a dense 27B model under Apache 2.0 that fits on a single GPU. Ask an AI to “rebuild the experiment in this paper, then improve on it” and it normally stalls after a few turns, waiting for you to step in and steer. ...

August 3, 2026 ·  Updated: August 16, 2026 ·  19 min ·  3928 words
Kimi K3: The AI That Designed a Chip in 48 Hours

Kimi K3: The AI That Designed a Chip in 48 Hours

TL;DR: Kimi K3 in 30 Seconds Kimi K3 is the new AI model Chinese company Moonshot AI announced on July 16, 2026; at 2.8 trillion parameters, it is the largest open source model released so far. In a single 48-hour run it designed a working chip, finished an astrophysics study in 2 hours instead of 2 weeks, and cut a teaser video from 56 raw clips. It reads 1 million tokens at once: roughly like reading the entire Lord of the Rings trilogy in one sitting and remembering all of it. On coding it plays in the same league as the strongest Claude and ChatGPT models, and beats them on some tests. It falls behind on hard knowledge questions. You can try it in the Kimi app and on kimi.com. The model files go public on July 27, 2026. Ask an AI to design a chip for you and what happens? It probably writes you a nice article about how chips are designed. ...

July 17, 2026 ·  10 min ·  2082 words
Grok 4.5 Is Here! Cheaper and 4.2x More Efficient AI

Grok 4.5 Is Here! Cheaper and 4.2x More Efficient AI

A powerful move has arrived on the price and efficiency front of the AI race! On July 8, 2026, xAI, Elon Musk’s AI company now operating under the SpaceX umbrella, announced Grok 4.5, its first major model release since going public. The model is also the fruit of a collaboration with the Cursor team, and its pitch is clear: not the highest benchmark score, but the most efficient way to get work done. ...

July 8, 2026 ·  4 min ·  719 words
Kimi K2.5: China's Native Multimodal and Agentic AI Revolution

Kimi K2.5: China's Native Multimodal and Agentic AI Revolution

I’m back with a groundbreaking development that is shaking up the tech world! Yes, as you guessed from the title, we are talking about Kimi K2.5. Developed by the Chinese company Moonshot AI, this model is currently taking the world by storm with its 1.04 Trillion parameters and technical specifications. 🚀 In this post, we will take a close look at the technical details, features, and popularity of Kimi K2.5, which is challenging giants like GPT-4.1 and Claude. 👇🏻 ...

February 1, 2026 ·  4 min ·  797 words