Mistral Large 4: Global Strength, With Reservations About Running Agents
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4: Global Strength, With Reservations About Running Agents on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, a major improvement over its predecessor but below leading US and Chinese models in the cited rankings. The supplied analysis also flags higher task costs than two Chinese models that score better, unusually high output volume and an author’s unverified observations of confident hallucinations. The model remains in research preview, with weights and licensing details still pending.

Mistral released Large 4 in research public preview, and the model scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, a substantial improvement on its predecessor but still below every listed US and Chinese flagship in the supplied comparison. The release gives France a stronger contender in advanced AI, while the ranking and cost figures raise questions about whether buyers should use it for long-running agent tasks.

Artificial Analysis’ current index places Large 4 below the listed leaders, including Claude Opus 5.5 at 57.6, GPT-6 Astra at 52.7 and China’s GLM-5.3 at 44.8. Large 4 is also below GLM-5.3-Flash at 41.8 and DeepSeek V4.1 Flash at 39.5. The source material describes it as the highest-scoring model outside the United States and China in the comparison; that distinction does not put it level with the leading models overall.

The index score marks a sharp step up for Mistral: the source says Large 3 scored 9 and Medium 3.5 scored 14 on the same index version. Large 4 has one trillion total parameters, with 49 billion active, accepts text and images, produces text, and has a 512,000-token context window. Mistral says reinforcement learning is ongoing, so its performance scores may change.

Large 4 is available through Mistral’s API as a research public preview. The listed price is $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14; the source says Mistral is offering a 50% discount for the first two weeks. Mistral has promised to release the weights at the end of October, but the licence has not been published. Until then, users cannot treat the preview as an openly licensed weights release.

At a glance
reportWhen: Released yesterday, according to the so…
The developmentMistral’s newly released Large 4 made a large leap on an independent model ranking, but the same data places it below leading US and Chinese rivals and raises concerns about its cost and use in agent workflows.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

Agent Work Faces Cost and Reliability Tests

The comparison matters most for companies evaluating models to carry out multi-step work, rather than answer a single prompt. The source notes that Artificial Analysis’ index includes agent-focused tasks such as AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. A lower score on a benchmark spanning those tasks may signal a wider capability gap, although an index result alone does not predict whether a particular company’s workflow will succeed.

The source reports that Large 4 used 200 million output tokens to complete the index, compared with a median of 81 million for comparable models. That is a benchmark observation, not a general measure of every user’s token consumption. Still, higher output volume can add cost and time in workflows that repeatedly call a model. The source calculates a cost of $1.13 per Intelligence Index task for Large 4, against $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash, both of which scored higher in the cited results.

For buyers, the headline is not simply that a European model improved. It is that the performance, price and reliability trade-offs differ by task. A company may value access to a French provider or a specific capability, but the supplied figures do not establish that Large 4 is the best-value choice for agent workloads. Procurement decisions require testing against the company’s own tasks, token use and tolerance for errors.

Amazon

AI language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Mistral’s Previous Index Scores

The launch analysis frames Large 4 as both a major gain for Mistral and a reminder of the distance between its model and current leaders. Its reported score rose from 9 for Large 3 to 38.4 for Large 4 on the same version of Artificial Analysis’ index. Those results make the release a significant improvement within Mistral’s own lineup, but they do not erase the gap to the highest-scoring models listed.

The source describes Large 4 as an alternative to models from the US and China, but its ranking table shows Chinese open-weight models ahead of it. Its characterization as the strongest model outside those two countries applies to the set in the source’s comparison, not to a broad demonstration that it matches the global frontier. The promised end-of-October weights release could also change how developers assess access, deployment and licensing, once the actual terms are available.

Artificial Analysis is the source of the index scores and task-cost comparisons cited here. The source material separately reports hands-on observations of hallucinations; those are the author’s experience, not a published Artificial Analysis metric. Keeping those evidence types distinct is important when interpreting claims about reliability.

Amazon

large language model for developers

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights, Licence and Reliability Remain Open

Several points are still unsettled. Mistral has promised weights for the end of October, but the source does not give a specific date or explain what licence will apply. Until the release and licence are published, developers cannot determine the precise reuse and deployment terms. Mistral also says reinforcement learning is ongoing, leaving room for benchmark results to change.

The source’s concerns about hallucination require care: they come from the author’s hands-on experience, and no sample size, test protocol or measured Large 4 hallucination rate is provided. The benchmark scores and reported token use offer a comparison on Artificial Analysis’ tasks, but they cannot settle performance on every real-world agent workflow. It is also unclear how the first two weeks of discounted pricing affect the source’s task-cost comparison, which is reported at standard pricing.

Amazon

AI model token cost calculator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch the Weights and Updated Scores

The next concrete milestone is Mistral’s promised end-of-October weights release. Developers and businesses will be able to assess the model’s deployment options more fully once Mistral publishes the weights and licence. Buyers should also check whether Artificial Analysis updates Large 4’s score as reinforcement learning continues.

For organizations considering the API preview now, the useful next step is a controlled comparison on their own tasks. That should track not only task completion, but also output-token use, latency, cost and factual errors across multi-step runs. The supplied index offers a common reference point; it does not replace testing under the conditions in which a system will be used.

Amazon

AI model performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How did Mistral Large 4 score?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the supplied source. The source reports scores of 9 for Mistral Large 3 and 14 for Medium 3.5 on the same index version.

Does Large 4 lead the global model rankings?

No. In the comparison supplied, it trails all listed US and Chinese flagship models, including GLM-5.3 at 44.8 and DeepSeek V4.1 Flash at 39.5. Its distinction is that it is the highest-scoring model outside the US and China in that comparison.

Can developers download Large 4’s weights now?

The source says Large 4 is currently available as a research public preview through Mistral’s API. Mistral has promised weights for the end of October, but the licence is unpublished in the supplied material.

Is Large 4 suitable for AI agents?

The cited index includes agent-focused tasks, but its score does not determine performance on every workflow. The source raises concerns about its relative benchmark score, token use, cost and reported hallucinations. Teams should test it on their own tasks before relying on it for multi-step work.

How does its task cost compare with the cited alternatives?

The source estimates $1.13 per Intelligence Index task for Large 4, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Those two models scored higher in the cited index results. These are source-reported benchmark costs, not guaranteed costs for other workloads.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Explore AI Innovation By Building A Grok Bot With Grok Bot

xAI announced a project titled ‘Designing Grok Bot with Grok Bot,’ indicating Grok AI’s involvement in creating a new system, but details remain limited.

How Does 512GB Storage Benefit AI Work In The M5 Ultra Mac Studio?

Exploring how the 512GB storage option in the M5 Ultra Mac Studio benefits AI workloads, including large model handling and performance implications.

Valve And Gaming Trends: The Curious Case Of The Barebones Steam Machine

Valve considered a minimalistic Steam Machine, but it never materialized. This analysis explores why and what it means for gaming trends.

The Core Engines Of AI: A Look Inside Twelve Machines

A detailed look into twelve fundamental AI models, explaining how they work, why they matter, and what remains uncertain about their capabilities.