Mistral and Reflection Announced Open-Weight Models a Day Apart. Neither Has Released the Weights Yet.
On Monday, October 5, Reflection AI introduced Beam, which it calls its first open-weight model. On Tuesday, October 6, Mistral launched a public preview of Mistral Large 4, which it nicknamed "le Chonk." Two well-funded Western open model labs, two announcements in about a day, and both framed against the same rivals: the Chinese open models whose weights anyone can already download.
There is one thing you cannot do with either model today. You cannot download it.
Reflection says it will release the weights, technical report and model card "later this month." Mistral says it will release the weights "by the end of the month," and The Next Web reported a target of October 27. When we checked Hugging Face at about 14:00 UTC on Tuesday, Mistral's organization showed no Large 4 repository and we found no Beam repository at all.
What Each Lab Actually Put Out
Beam is a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active. Reflection says it was pretrained on 23.8 trillion tokens in under four weeks on 6,144 Nvidia GB300 GPUs, then put through a reinforcement learning run on 10.5K GB300s for four weeks that generated more than 100 million rollouts. It is text-only, with an effective context length of 1 million tokens. Access today is a waitlist for "a select group of users."
Mistral Large 4 is bigger: 1 trillion total parameters, 49 billion active, natively multimodal across text and images. Mistral says it was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in its own European datacenters, and that the preview is served from the same machines. Unlike Beam, you can pay for it now, through Mistral Studio, at $1.36 per million input tokens and $4.18 per million output tokens.
| Reflection Beam | Mistral Large 4 | |
|---|---|---|
| Announced | Monday, October 5, 2026 | Tuesday, October 6, 2026 |
| Parameters | 501B total, 23B active | 1T total, 49B active |
| Inputs | Text only, 1M token context | Text and images |
| Training hardware (company stated) | 6,144 GB300 GPUs for pretraining; 10.5K GB300s for 4 weeks of RL | 3,800 Grace Blackwell GPUs |
| Access today | Early access waitlist | Paid API preview, $1.36 in / $4.18 out per million tokens |
| Weights promised | "Later this month" | "By the end of the month" (October 27 per The Next Web) |
| License stated | Apache 2.0 | Not stated in the announcement |
Both Pitches Are Measured Against China
Read the two posts side by side and the reference point is the same. Reflection says Beam "advances the Western open-weight frontier" and is competitive with GLM 5.2 while approaching Qwen 3.8-Max on coding and agentic tasks. It concedes that "frontier open models like Kimi K3 remain ahead on raw capability" and stakes its claim on efficiency instead: scores comparable to GLM-5.2 on advanced reasoning benchmarks while using 3 to 4 times less inference compute.
Mistral is more aggressive. It says Large 4 is "significantly outperforming any open-weight model developed in the US or Europe" and leads open-weight models developed outside China on the Artificial Analysis Cyber Index "by a wide margin." Both qualifiers do the same work. Neither lab is claiming the open frontier. Each is claiming the best seat outside China.
That is an honest framing, and we think it is the right one. The numbers in Reflection's own table show why. On DeepSWE v1.1, Reflection lists Beam at 44.4, GLM 5.3 at 61.0, Kimi K3 at 68.0 and DeepSeek V4.1 Flash at 74.2. Mistral reports Large 4 at 61.7 on the same benchmark, from a private run by Artificial Analysis ahead of the harness's public launch.
| DeepSWE v1.1 | Score | Reported by | Weights on Hugging Face today |
|---|---|---|---|
| DeepSeek V4.1 Flash | 74.2 | Reflection's table | Yes |
| Kimi K3 | 68.0 | Reflection's table | Yes |
| Mistral Large 4 | 61.7 | Mistral (Artificial Analysis private run) | No |
| GLM 5.3 | 61.0 | Reflection's table | Yes |
| Qwen 3.8 Max | 51.0 | Reflection's table | Not checked |
| Reflection Beam | 44.4 | Reflection's table | No |
| GLM 5.2 | 44.0 | Reflection's table | Yes |
Treat this table as a sketch, not a leaderboard. The two labs ran their numbers separately, and Reflection says it took scores for other models from Artificial Analysis and DataCurve. Mistral's figure has not been posted publicly by Artificial Analysis yet. What it does show is the shape of the field on two vendors' own data: the two Western entrants each land next to a GLM release, and the top of this chart is Chinese.
The weights column matters most. Kimi K3, GLM 5.2 and GLM 5.3 sit in official repositories on Hugging Face in repos dating to June and August, and DeepSeek V4.1 Flash has been there since September. Anyone can download them, quantize them and run their own evals. Beam and Large 4 can only be judged on what their makers chose to publish.
The Gap Between Announcement and Download
Both labs say they are using the window for safety work. Reflection says Beam "is undergoing final red-teaming and evaluations." Mistral goes further, and this is the part of its post we would read twice. Until the weights ship, it says, it is red-teaming with "cybersecurity leaders, vetted partners, and state authorities, who will access the same model with reduced moderation and expanded cyber capabilities."
That is the same pattern we described with Google's Gemini 4 Argon last week: defenders first, guardrails loosened for them, everyone else later. Mistral makes the cyber case explicitly. It says Large 4 scores 82 percent on an index test that asks a model to reproduce a real vulnerability and patch it, "the highest of any model," and that Claude Opus 5.5 and GPT-6 Astra score near zero on that test because they refuse. It also reports 93 percent on Cybench, a set of 40 security competition exercises.
Here is the tension. Mistral's argument for open weights in cyber is that "provider-level refusals can block legitimate vulnerability research." Once the weights are public, nobody controls refusals, for defenders or anyone else. Mistral also says Large 4 refuses malicious cyber prompts at a higher average rate than all open models it compared, which is the right thing to measure, but a refusal rate is a property of the released checkpoint, not of what people do with it after fine-tuning.
Caveats
Every capability number in this piece is self-reported. Reflection estimates its 3 to 4 times compute advantage from active parameters and generated token counts, and says the estimate excludes prompt prefill, attention costs and serving overhead, so it is "an approximate compute comparison rather than measured inference cost." Mistral's blind human coding evaluation, run with Surge AI, placed Large 4 second of five models at 3.74, behind Claude Opus 5 at 4.22, and ahead of Kimi K3, GLM-5.3 and GLM-5.2. That is a respectable result, and a vendor-commissioned one.
Mistral also says its reinforcement learning run "is still in flight," so the checkpoint that ships at the end of the month may not be the one behind today's numbers. That cuts both ways: it could be better, and it will not be the model anyone actually benchmarked.
Our Take
We do not object to a red-team window. After the year this industry has had, a few weeks of outside testing before weights go public is the responsible order of operations, and we would rather see it announced than skipped.
What we object to is the label. An open-weight model is one you can download. Until then, Large 4 is a closed API preview with a promise attached, and Beam is a waitlist with a blog post. The practical difference is not semantic: an open checkpoint lets anyone check the benchmarks, and these two launches are asking for trust on precisely the numbers that open weights exist to verify.
Of the two, Reflection made the more credible pitch. It named a license, published a full comparison table that includes models beating it by wide margins, and said plainly that Kimi K3 is ahead. Mistral has the stronger product today, a priced and usable API with images and a serious cyber story, but its post does not name a license, and a weights release without Apache 2.0 or something close would make "open" a much weaker word.
The bigger picture is that both companies now define success as leading the West. That can be a real market, for any government or regulated buyer that wants a model it can host without Chinese provenance. It is not the frontier, and both posts say so if you read the qualifiers.
Three signposts for the next 60 days: whether both sets of weights actually land in October, and under what license; whether independent runs on the released checkpoints, from Artificial Analysis or anyone else, reproduce the DeepSWE and cyber numbers within a few points; and whether Mistral publishes who got the reduced-moderation version during the red-team window and what they found.
For background on how both labs got here, see our June look at Reflection's Colossus compute deal and our analysis of Mistral's Series D.
Sources: Reflection: Introducing Beam, Mistral: Introducing Mistral Large 4, TechCrunch, SiliconANGLE, The Next Web, WIRED, and Hugging Face model pages for Kimi K3, GLM-5.3, GLM-5.2 and DeepSeek V4.1 Flash.
