Mistral has unveiled Large 4, its largest and most capable model to date, according to the French AI company. Nicknamed Le Chonk, it can read text and images and respond in text. Early independent comparisons put it close to GPT-6 Luna and DeepSeek V4.1 Flash on broad capability, but practical tests have produced a more uneven picture. The preview opened on 6 October; downloadable model weights are promised by the end of the month.
Based on 14 primary sources · See sources
What has Mistral released?
Mistral’s launch announcement describes a model trained and served during the preview on its own European infrastructure. Developers can access it through an API: a connection that lets their software send requests to the model. It is not yet a downloadable release. Open weights would let organisations run the model on their own hardware rather than rely entirely on Mistral’s service.
Mistral announces the API preview and a later open-weights release.
View the original post on X ↗Large 4 accepts images as well as text, so it can examine a document or picture and write an answer about it. It produces text, not generated images. Mistral also reports training data in more than 160 languages, including every official EU language. That breadth is worth investigating, although the announcement alone does not establish how well it writes in each language.
The promised weights release will make the self-hosting story easier to assess: the licence, hardware requirements and final release details still need checking. Open weights do not remove the expense of running a large model. Our coverage of Kolibri’s deployment requirements shows what those costs can involve for another European model; it is not a hardware guide for Large 4.
How does it compare with other models?
Artificial Analysis’s independent launch evaluation gives Large 4 a score of 38 on its Intelligence Index, which combines a collection of tests into one measure. That places it alongside GPT-6 Luna at maximum reasoning effort and just behind DeepSeek V4.1 Flash at maximum effort, on 39. It is a useful indication of broad capability, rather than a verdict on every task.
The picture changes when the tests become more specialised. Large 4 matches GLM-5.3-Flash on Artificial Analysis’s cybersecurity index, while trailing Kimi K3 on its document-and-image test. Vals AI’s results show a similar split: sixth of 75 models on a legal-agent test, but 32nd of 44 on its aggregate index. The table below keeps those different measures separate; a strong result in one field does not imply the same standing elsewhere.
A separate coding evaluation reported by Surge AI offers a more favourable result. Expert software engineers reviewed outputs from five models without knowing which model produced them. Large 4 came second behind Claude Opus 5 and first among the open-weight candidates. Mistral arranged that evaluation for its release, with external annotators, so it should be read alongside the independent results rather than treated as another general ranking.
Broad intelligence: close to Luna and Flash
Artificial Analysis Intelligence Index · selected configurations · higher is better
Specialist results tell different stories
Each row is a separate evaluation. Scores cannot be compared across rows.
| Evaluation | Large 4 result | Comparison / scope |
|---|---|---|
| Cybersecurity · AA Cyber Index | 50 | Level with GLM-5.3-Flash (50) |
| Documents and images · AA GDP.pdf | 19% | Kimi K3: 22%; document and image tasks |
| Overall · Vals Index | 48.05% | 32nd of 44; aggregate Vals evaluation |
| Legal tasks · Harvey Legal Agent | 15.83% | 6th of 75; specialist legal tasks |
| Surge human coding review | 2nd of 5 | Behind Opus 5; blind evaluation arranged for Mistral’s release |
Early users see progress—and uneven results
Developer Simon Willison tried a familiar test: asking the model to write the code for a drawing of a pelican riding a bicycle. The format, SVG, describes shapes that a browser can display. His published comparison shows outputs from both available reasoning settings, none and high. He preferred the high-effort drawing and saw a substantial improvement over Large 3, while still judging Large 4 below the leading models.
“Mistral are back in the game.”
Wayne Lowry tested something more demanding: seven game builds through Hermes Agent, a coding tool working with the model. Each game got one attempt, without retries or manual fixes, before a bot tried playing it. His original game-building post reports a mixture of successful and weaker results.
A one-attempt game-building test reports a mixture of successful and weaker outputs.
View the original post on X ↗Lowry’s follow-up the next morning was more critical of his live experience. It illustrates the gap between a promising example and a tool that feels dependable in use. Both posts describe one person’s experience, rather than a measured failure rate across users.
“My experience so far? Less than stellar.”
TradingToni’s early impression was more positive, describing Large 4 as fast and a useful all-rounder. Unlike the drawing and game tests, that post does not supply prompts or a testing method. It adds a user’s perspective, but gives us less evidence to assess the result.
Michael D. Olmos’s walkthrough, which includes a Rubik’s-cube build, highlights another issue: getting a model to work inside a coding tool can introduce its own difficulties. That matters when interpreting the mixed reports. They test the model together with particular tools and settings, not one identical setup. The table records those differences, including the more detailed failure report examined next.
What was tested, and what remains unproven
A small early sample of firsthand reports, not a community survey.
| Reviewer | Test / observation | Reported result | Limit |
|---|---|---|---|
| Simon Willison | Code for a pelican drawing; none vs high reasoning | Preferred high; improved over Large 3. Output tokens: high 2,717; none 3,275 | One drawing prompt |
| Wayne Lowry | Seven games through Hermes; one attempt each | Mixed builds; later post more critical | One developer and one coding tool |
| TradingToni | Early general impression | Fast, useful all-rounder | No published test protocol |
| Michael D. Olmos | Walkthrough with Rubik’s-cube build | Highlights coding-tool integration friction | Workflow observation; no controlled ranking |
| Adam Holter | Eight preserved test variants | Refusals, build and rendering defects | 4,096-token output cap; different access routes; no reasoning tokens reported |
What the failed builds tell us
Adam Holter’s eight-variant report goes beyond impressions by saving the responses, screenshots and checks of the generated work. Among the results were refusals, a learning app with unreadable dark-on-dark text and content extending beyond a phone screen, another app with code that would not run, and controller drawings with incorrect shapes. Those are visible failures a polished initial screenshot could conceal.
There is an important limitation to that evidence. Holter capped each response at 4,096 output tokens—the pieces of text used to measure generated output—and requested continuations when responses were cut short. He used both a console and an API, with a sampling setting of 0.7. His medium reasoning request through OpenRouter reported no reasoning tokens. These conditions do not establish how Mistral’s high-reasoning setting performs, but they do show why a generated build needs checking beyond its appearance.
Speed, price and the details still in question
Mistral’s model card lists standard API prices of $1.36 per million input tokens and $4.18 per million output tokens. Input is the material sent to the model; output is what it generates. A two-week launch offer halves those rates to $0.68 and $2.09. The chart below asks a different question: what did it cost to complete Artificial Analysis’s benchmark tasks?
That distinction matters because the amount of text generated affects the bill. Artificial Analysis’s measurements put Large 4 at about 116 output tokens per second, but across its full Intelligence Index evaluation the model generated 200 million output tokens, compared with a median of 81 million. Those totals cover the evaluation, not one request. A model can generate text quickly and still cost more per completed task if it produces substantially more of it.
Some published specifications also disagree. The announcement gives roughly one trillion parameters—the numerical values learned during training—with 49 billion used at a time. The model card lists 1.05 trillion and 52 billion active, plus a component for processing images. The card also advertises a one-million-token context window, meaning the amount of material the model can hold at once, while Artificial Analysis lists about 512,000. We have not resolved those differences; the actual endpoint limits matter more than assuming the larger figure applies.
Similar intelligence, different task costs
USD per Artificial Analysis Intelligence Index task · lower is better
What to watch next
The preview is enough to show a substantial step forward for Mistral, with strengths that vary by task and practical results that are still mixed. The next milestone is the promised weights release. It will let readers examine the final model, licence and requirements for running it themselves. Until then, comparisons should name the preview version, settings and tools used, so that a result can be understood—and, where possible, reproduced.
Sources
- Large 4 announcement and vendor evaluations
Mistral · Announcement ·
- Original Large 4 announcement on X
Mistral AI · Announcement ·
- Public preview model card and current pricing
Mistral · Documentation ·
- Preview release and two-week launch discount
Mistral · Documentation ·
- Independent launch evaluation and peer comparisons
Artificial Analysis · Research ·
- Intelligence, throughput and output-token measurements
Artificial Analysis · Research ·
- Specialist and aggregate benchmark results
Vals AI · Research ·
- Blind human coding evaluation for the Mistral release
Surge AI · Research ·
- Reasoning-level comparison with saved SVG outputs
Simon Willison · Research ·
- Seven one-shot game builds using Hermes Agent
Wayne Lowry · Research ·
- Follow-up criticism of the preview experience
Wayne Lowry · Research ·
- Positive early impression without a published test protocol
TradingToni · Research ·
- Walkthrough and coding-harness observations
Michael D. Olmos · Research ·
- Eight test variants with preserved outputs and verification notes
Adam Holter · Research ·
Written with AI assistance from the sources above. Analysis and practical implications are our interpretation. Our editorial approach.



