Alibaba released Qwen3.8-Max on Monday and said it will publish the model's weights next week, opening the internals of its most capable AI system to anyone who wants to download them. Qwen3.8-Max runs on 2.4 trillion parameters, the values a model adjusts during training and the industry's crude measure of size, although only about 95 billion of them are active for any given word it produces. It accepts up to 1 million tokens in one query.
Investors treated it as a win. Alibaba's Hong Kong-listed shares closed 7% higher at HK$125.20, and its US-listed shares rose 4.2% in premarket trading. Citi analysts linked part of the move to the new model's benchmark results.
That reaction is less about the model than about what Alibaba is doing with it. The company spent much of this year keeping its strongest models proprietary, and Monday's return to open-sourcing its top-tier releases reverses that.
Alibaba's own framing is blunt. This is "the first time we will open-source the weights of a Qwen-Max-class model," the team wrote, with files due on Hugging Face and ModelScope next week. Until now, a business that wanted Alibaba's strongest model had to send its data to Alibaba's servers and pay per use. From next week it can download the same model and run it on hardware it controls. None of the leading US labs publish weights for their frontier models at all.
Alibaba is also arriving into a pattern already set by its neighbours. Moonshot AI published weights for the 2.8 trillion parameter Kimi K3 on July 27, and independent trackers placed it behind only Claude Fable 5 and GPT-5.6 Sol Max. DeepSeek retrained its cheap V4-Flash model on July 31, trailing Claude Opus 4.8 on all nine of the coding benchmarks it published at roughly one fifty-fourth of the token cost. Near-frontier, downloadable and cheap is now the standard Chinese offer, and it prices against US models that are none of those three.
Days-long tasks, not single answers
Alibaba built the launch around long jobs rather than clever answers. It set the model loose on a software project for more than 10 days; as of July 30, the repository had accumulated 265 commits, 127 pull requests and 151 issues after roughly 16 days with no human involvement.
In another test it handed the model a research paper on choosing training data, plus a set of GPUs and no starter code. Over about 125 hours it wrote roughly 7,600 lines of code, ran 33 rounds of training, reproduced the paper's six findings, then tested 18 ideas of its own and landed on a method that beat the paper by 2.7 points on a competition-level math benchmark.
It also entered a live machine-learning contest on Alibaba's Tianchi platform against 526 human teams. Working alone under a 24-hour limit, it made 45 submissions and finished ahead of 458 of them. In a separate chip-design run, it cut a cryptographic circuit from 8,298 logic gates to 678, and shrank the physical layout by 81%.
A year-long simulated e-commerce business is the case retail investors may find most legible. Starting with ¥100,000, the model finished with ¥416,252, 38% ahead of the second-place GLM 5.2 and 152% above Alibaba's previous flagship, Qwen3.7-Max.
Ahead on some benchmarks, behind on others
The published comparison table is mixed rather than triumphant. Qwen3.8-Max scores 86.6 on Terminal Bench 2.1, ahead of Claude Opus 4.8 and Claude Fable 5 at 84.6 but behind GPT-5.6 Sol at 88.8. It leads the table on PaperBench, a test of reproducing research papers, at 93.0.
On other tests it trails clearly. Claude Fable 5 scores 80.0 on SWE-bench Pro against 67.7 for Qwen3.8-Max, and 53.3 on Humanity's Last Exam, a broad reasoning test, against 43.6. Bloomberg reported that the model ranks above Kimi K3 on several benchmarks, which is the comparison that carries most weight inside China.
What the numbers do not settle
Much of the evidence is Alibaba's own. Several benchmarks in the table were built in-house, the autonomous coding and chip-design results come from Alibaba's own unaudited runs, and a footnote warns that the Claude Fable 5 scores may involve fallbacks. Independent trackers have not yet rated the released model.
Open weights also do not mean usable weights. A 2.4 trillion parameter file needs a rack of datacenter GPUs, the same wall Kimi K3 ran into last week. Alibaba has not said which license will govern the release.
What next week buys Alibaba is distribution rather than revenue. Every company that downloads Qwen3.8-Max is one that did not sign a contract with OpenAI or Anthropic, and every company that builds on it is harder to move later. The weights ship next week. The license will show how much of that bet Alibaba is actually making.
