Skip to content

Qwen3.8-Max Hits 2.4 Trillion Parameters and Undercuts Claude Opus 5 on API Price

Alibaba's Qwen3.8-Max brings 2.4 trillion parameters, a one-million-token context window and API pricing far under Claude Opus 5, with open weights next week.

A
Argal
Argal
4 min read
Alibaba Qwen logo artwork
Branding for Alibaba's Qwen family of large language models. Image: AI Business

Alibaba released Qwen3.8-Max on August 3, 2026, and called it its largest and most capable model so far. The headline numbers: 2.4 trillion total parameters, a context window of up to one million tokens, and API pricing that sits far below Anthropic's Claude Opus 5. Weights are due to be published in the following week, which would make it the largest open-weight model any company has shipped outside Moonshot AI's Kimi K3.

What is inside Qwen3.8-Max

The 2.4 trillion figure is the total parameter count, not what runs on every request. Qwen3.8-Max uses a sparse mixture-of-experts design, meaning the model is split into many specialist sub-networks and only a few are switched on per query. CGTN reported that about 95 billion parameters are active per query. That is the trick that keeps a model this size affordable to serve.

It is multimodal, handling text, images and video, and the one-million-token context window means it can hold a very large document set, or a whole codebase, in a single conversation.

Alibaba is pitching it hardest at long-running work. The company says the model can "code autonomously for weeks with little human input," and lists application design, legal document review, sports analytics and financial research among the workloads it targets, according to AI Business.

Where it lands on the leaderboards

Alibaba's own benchmark claims put it near, but not at, the top:

LeaderboardRankNote
Text Arena5thbehind the Claude family
Vision Arena2ndits strongest showing
Frontend Code Arena4thscore of 1,668

Alibaba describes the results as comparable to, and sometimes better than, Anthropic's Fable 5. Treat these as vendor-selected comparisons until independent evaluations land, which is the standard caution for every model launch, not a knock specific to this one.

The price is the point

The more interesting number is not the parameter count. In international markets, Qwen3.8-Max input tokens are priced at roughly 40 percent of Claude Opus 5, and output tokens at roughly 24 percent. Output is where most production agent workloads actually spend money, so a roughly four-times reduction there changes what a small team can afford to run.

For comparison on capability, Claude Opus 5 launched with the same one-million-token context window. Alibaba is not claiming to beat it outright. It is claiming to get close enough at a fraction of the cost, which is the same play Chinese labs have been running all year.

The open-weight race it is joining

Qwen3.8-Max lands about a month after Kimi K3, Moonshot's 2.8-trillion-parameter open model, which remains larger on paper. DeepSeek also shipped a V4-Flash model days earlier, with early figures suggesting it is among the cheapest models to run anywhere.

The practical catch with open weights at this scale is hardware. Running Kimi K3's weights locally needs roughly 1.4TB of memory, which puts it out of reach of almost everyone outside a data centre. Expect the same to be true here. "Open weights" mostly means other cloud providers and research labs can host it, not that you will run it on a workstation.

What it means for developers in the Philippines

Qwen3.8-Max is reachable from here today through Alibaba Cloud Model Studio, which serves developers globally, and that is the honest extent of the local angle: there is no Philippine-specific pricing, no local distribution partner, and no announcement tied to the Philippine market. What does travel is the cost. Filipino startups and agencies buying model capacity pay in dollars against a peso budget, so a model that costs roughly a quarter as much per output token for broadly similar work is a real change to what an AI feature costs to ship.

The caution is equally practical. Data residency, service terms, and the political risk attached to Chinese AI providers are live questions for anyone handling regulated Philippine data, and none of them are answered by a benchmark score. Independent testing on your own workload remains the only reliable way to know whether the cheaper token is actually cheaper in total.

Explore topics related to this article

A
Argal

Argal

@argal

Clurky is a Philippine tech news site owned and run by Argal, a Philippines-born software developer based in Singapore with a Computer Science background. He covers Philippine tech, fintech, and digital services - from gadgets and AI to software and security - along with evergreen guides and explainers, all with a builder's eye for how these systems actually work. Every article is fact-checked against primary sources.

170 posts

Comments

Join the conversation

Sign in to leave a comment and reply to others.

Sign in
Loading comments...