CompaniesInvestorsPeople
Home
Loading

aVenture is in Beta: research coverage is expanding as we build, so please independently verify key details before making investment decisions.

aVenture is in Beta: research coverage is expanding as we build, so please independently verify key details before making investment decisions.

Get in Touch

  • Contact

  • Request a Demo

  • Request Data Updates

  • Add a Company

Research

  • Companies

  • Investors

  • People

aVenture

  • Download App

  • Pricing

Download the aVenture Research beta for iOS and iPadOSDownload aVenture Research on the Mac App Store

Resources

  • Documentation

  • Use Cases

  • CLI

  • MCP

  • Feature Requests

  • Sitemap

Member

Backed by

© aVenture Investment Company, 2026. All rights reserved.

San Francisco, CA, USA

Privacy Policy · Terms of Service · Privacy FAQ

aVenture Investment Company ("aVenture") is an independent research platform providing detailed analysis and data on startups, venture capital investments, and key industry individuals. It is not a registered investment adviser, broker-dealer, or investment advisor and does not provide investment advice or recommendations. The data provided by aVenture does not constitute recommendations or advice, whether by methodology, analysis, AI-generated content, or a statement written by a staff member of aVenture.

aVenture is not affiliated with any of the people, companies, organizations, government agencies, regulatory bodies, or investment funds we provide coverage for on this site unless explicitly stated otherwise. Users assume full responsibility for decisions made based on information obtained from this platform. Links to external websites do not imply endorsement or affiliation with aVenture. Any links that provide the ability to invest in a primary or secondary transaction in a company are for convenience only and do not constitute solicitations or offers to buy or sell an investment. Investors should exercise heightened precaution and due diligence when investing in private companies, especially those not independently audited.

While we strive to provide valuable insights with objectivity and professional diligence, we cannot guarantee the accuracy of the information provided on our platform. Before making any investment decisions, you should verify the accuracy of all pertinent details for your decision. To the fullest extent permitted by law, aVenture shall not be liable for any direct, indirect, incidental, consequential, or financial damages arising from use of this site, whether by consumers of its contents directly or by persons or organizations covered by our research, even if we are advised of the possibility. Our best-efforts processes and correction request forms do not create a warranty or duty of care.

Profiles on this platform may include content generated in part by large language models (LLMs, artificial intelligence) that aggregate publicly available sources (e.g., SEC EDGAR, public filings, press releases). Source attribution is provided where known; always verify statements and claims here against original sources before relying on any data. Content on our site may contain inaccuracies, omissions, or what are commonly called 'hallucinations' if generated in part or in full by AI / LLMs. The risk can also exist even when content is written by a human, as internal and third-party sources may also have inaccuracies for the same or different reasons. While we randomly audit a proportion of content, this is not exhaustive.

We recommend that an independent auditor be hired to verify the accuracy of the information before relying on it for any sensitive decisions. By accessing this platform, you agree not to rely solely on any information generated by AI, aggregated, or sourced or written otherwise on this site, for investment, financial, or other decisions. aVenture assumes no responsibility for inaccuracies, omissions, or hallucinations. You must independently verify all data from primary sources. Use of this platform constitutes your waiver of claims for reliance-based damages, including negligent misrepresentation. To report an error, request a correction, or dispute information about a company or individual, contact us via our request data updates form.

Loading
Loading
Home
News
Introducing Mistral Large 4

From Mistral Blog

October 6, 2026

Introducing Mistral Large 4

Introducing Mistral Large 4

Le chonk

IntroducingMistral Large 4

Back to Blog

14 min read

October 6, 2026

By Mistral

Le Chonk 

Today, we’re launching a public preview of Mistral Large 4. Unofficially ML4, very officially: le Chonk. ML4 pushes the frontier of open-weight performance. You can try the preview API today on Mistral Studio. Weights drop end of this month.

Frontier performance

ML4 is a 1 trillion-parameter natively multimodal model with 49 billion active parameters. It is our largest and most capable model to date, and it continues to improve rapidly as we refine it.

The model demonstrates exceptional performance across coding, agentic workflows, and multimodal understanding. It already achieves performance competitive with the strongest open-source models globally, while significantly outperforming any open-weight model developed in the US or Europe. On critical enterprise workloads, including cybersecurity, finance and law, we find it to be state-of-the-art among open models. In some domains such as visual grounding, it goes further still, surpassing even frontier closed models.

We will release the weights by the end of the month. Until then, we are red-teaming the model in real-world settings with cybersecurity leaders, vetted partners, and state authorities, who will access the same model with reduced moderation and expanded cyber capabilities.

Forged in Europe. Built for AI sovereignty.

ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe. The public preview is served on that same infrastructure. It is a significant milestone in our long-term investment across infrastructure, research, and product development: state-of-the-art performance in critical verticals, delivered through open weights, designed to give customers control over their AI.

This is particularly important in cybersecurity, where provider-level refusals can block legitimate vulnerability research and incident response, and where losing access to a capability mid-incident can itself become a critical security risk. ML4 pairs top-tier cyber performance with open weights and self-deployment, giving organizations both the capability and the autonomy to run advanced security work under their own policies.

The model will be available across multiple regions worldwide, including a European deployment that Mistral operates end-to-end, independently of other digital service providers and under European law. Fun fact: a significant share of ML4’s training data was multilingual, spanning more than 160 languages, including every official language of the European Union. 

We’ve been working closely with leading enterprises across the world in finance, engineering, manufacturing, logistics, pharmaceuticals, science, shipping, public sector, and other mission-critical industries to train ML4. In fact, the model uses the same training, customization, and RL environment we offer our customers through Mistral Forge.

Try it today

There is still more to come. As we work toward releasing the weights, we will share further details on the model architecture, additional benchmarks, and our post-training methodology.

This model will also serve as the foundation for a new generation of specialized and optimized Mistral models. In the meantime, we invite you to try the preview API and share your feedback with us on social media.

Capabilities deep-dive

Cybersecurity

ML4 is one of the world's strongest AI models for cybersecurity. On the Artificial Analysis Cyber Index, an independent evaluation of how well AI models find and fix security flaws in real software, it ranks among the top five models globally and leads open-weight models developed outside China by a wide margin. On one of the index's tests, which asks a model to reproduce a real vulnerability in open-source software and then patch it, ML4 scores 82%, the highest of any model. It also solves 93% of the challenges in Cybench, a set of 40 exercises drawn from security competitions, one of the highest scores reported for an open-weight model.

That top score reflects a practical advantage. Several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test because they refuse to perform the task. Yet defending software often starts with proving that a flaw is real, exactly the kind of work safety filters in closed models can block. This matters even more as threat actors increasingly jailbreak those same models to support offensive cyber activity: defenders need systems that can match those capabilities without being constrained by the same refusals. ML4 can do that work, and its capabilities extend beyond what it was explicitly trained for: in internal testing, it proved useful for analysing malware, prioritising vulnerabilities, and writing detection rules. For organisations that need sovereign, auditable AI for security operations, it will be able to run on private cloud or on-premise.

ML4 against the field : efficiently reasoning over diverse complex challenges
Malware reverse-engineering: solving an out-of-distribution investigation task

Agentic coding

ML4 excels across software engineering, repository understanding, and complex terminal workflows, scoring 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4. Its combined Coding Agent Index score of 49.8% places it ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max.

We also ran a blind human evaluation with Surge AI on coding quality: professional annotators rated model outputs on a 1–5 scale, with model identities hidden. ML4 Preview ranked second of five models (3.74), ahead of Kimi K3 (3.59), GLM-5.3 (3.60) and GLM-5.2 (3.40), and behind only Claude Opus 5 (4.22).

Agentic Workflows

ML4 runs general-purpose agents that gather information, use tools, and produce finished deliverables across complex workflows. On AutomationBench — 657 business workflows across apps like Gmail, Google Sheets, Slack, and Salesforce — it scores 59.9%, ahead of Kimi K3, MiMo-V2.6-Pro, and DeepSeek V4 Pro.

It's just as strong on the professional deliverables that knowledge work actually produces: spreadsheets, slides, and PDFs. On AA-Briefcase, which evaluates long-horizon knowledge work, it reaches 1,393 Elo, ahead of DeepSeek V4 Pro.

Multimodal

ML4 is a step change in the ability of our models to understand images. It reasons powerfully across complex documents, charts, and natural images, and brings vision to the industries where perception is critical such as engineering, manufacturing, and earth observation.

The model can further combine visual grounding with agentic capabilities: from inspecting gigapixel satellite imagery — helping disaster-response teams act when time counts — to analyzing engineering-drawings — zooming in, inspecting, and verifying until the answer is exact. In our demos above, ML4 grounds dense natural scenes, verifies mechanical parts in technical drawings, retrieves evidence from PDFs, and scans massive geospatial images for the hardest-to-find objects.

On visual grounding particularly, we find ML4 to be one of the most capable models we tested, for instance surpassing GPT-6-Astra on Dense 200 (42% vs 41%).

Science and Math

ML4 brings strong scientific capabilities, built by combining AI-driven methods with our researchers' expertise in mathematics, physics, and chemistry.

It's highly proficient at agentic coding for scientific tasks such as data analysis, modeling, and simulating physical reality, which lets researchers focus on the questions rather than the plumbing. In benchmarks, ML4 is state of the art on SciCode-Verified among open-weight models. In practice, it can generate a full Hartree–Fock simulation in one shot — a complex, multi-step chemistry task built from a series of advanced routines.

ML4's math is stronger too, in both formal reasoning and applied mathematics. In our human evaluations it reasons more precisely and with more structure than GLM-5.3, and it can sustain long, domain-specific applied-mathematics tasks, including work relevant to frontier theoretical physics.

Together, these capabilities make ML4 a strong research assistant across the full technical workflow — from the first question to the final result.

SciCode-Verified tests the capabilities of models to implement complex scientific workflows in code for domains such as physics, mathematics, material science and biology. 

Internal eval on STEM tasks (math and physics) of ML4 against GLM5.3

Knowledge Work

ML4 is our most capable model for the real-world tasks which professionals handle every day. It can create, edit and fix complex spreadsheets and documents, showing exemplary performance on both legal and financial benchmarks.

Notably, we evaluated ML4 through third party evaluators (vals.ai) on representative tasks for both legal and financial tasks, finding the model exceeds GPT-6-Astra in both cases. On HarveyAI’s Legal Agent benchmark, ML4 outperforms all open-source models.

FinWorkBench tests model capabilities at creating/editing spreadsheets on real life Finance and Accounting use cases.

Financial analysis demands precision and the ability to synthesize information from multiple sources, a process that remains time-consuming at many financial institutions today. In this demo, ML4 compared to other top OSS models take on the same multistep corporate finance challenge, searching through public company filings and financial reports, such as those available via EDGAR and equivalent European databases. An animated semantic map traces each model's journey toward a solution, highlighting every document retrieved along the way. Each track's position reflects the evidence gathered, the results of calculations, and the questions that remain unresolved. Viewers can follow how the investigations unfold and compare the distinct paths each model takes before arriving at its final answer.

Model Safety

ML4 has saturated our benchmarks on robustness to indirect prompt injections, putting it at the frontier of OSS models (compared to GLM-5.2, GLM-5.3, Kimi-K2.6, Kimi-K3, DS-V4-Pro-0813). On Lakera’s public B3 AI Security Benchmark, ML4 resists 93.3% of attacks – we see no higher scores among competitors.

ML4 also engages more responsibly with users than any of our previous models. We highlight our results on the KORA Benchmark, where ML4 again sits at our highest measured score among OSS models (1.691, with 2 being the maximum denoted as “Exemplary”). 

Of particular relevance is the model’s propensity to refuse malicious requests regarding cybersecurity. Despite strong performance on Cyber benchmarks, the average refusal rate of the model on cyber prompts from JailbreakBench, StrongREJECT, and AgentHarm is higher than all OSS models.

Human Evaluation

We ran an internal evaluation in which expert annotators across coding, computer-aided design (CAD), finance, mathematics and physics compared Mistral Large 4 with GLM-5.3. ML4 was preferred in CAD and STEM, while performing on par or close to GLM-5.3 in finance and coding.

Reinforcement learning at scale

Base models are improving fast, and our post-training has to keep pace. A recipe tuned for yesterday's model leaves capability on the table with today's frontier, because ground truth samples that once pushed a model to its limits won’t anymore. We use Reinforcement Learning (RL) because it adapts as the model does: we train on the outcomes of the model's own attempts, and we can raise the difficulty and the breadth of the tasks as it gets stronger.

Our RL library was designed to make new environments easy to add and train at scale. A shared, composable interface allows a single training run to combine tasks ranging from single-turn chat and complex scientific problem solving to safety alignment, factuality, and long-horizon tool use. These environments share scaffolds and resources such as code sandboxes, web search, and external APIs. The same composability extends to verification, with reward models, unit tests, LLM judges, and static checks combined as needed for each task.

At runtime, an autoscaling fleet of actors generates tens of thousands of rollouts in parallel while model training proceeds asynchronously. The generation and training pipeline is optimized for long trajectories, supporting rollout budgets of millions of tokens across multiple compactions while keeping staleness low. Novel methods and optimizations across both stages minimize off-policy drift and enable stable RL over long horizons.

At our current scale (3k GPUs), a single training run produces roughly 33 billion tokens per day, of which around 16 billion are trainable completion tokens after filtering and masking. We can see the run progress directly in the training rollouts: training rewards rise across several representative environments as the policy learns to solve increasingly complex tasks. Below are a few examples.

The improvements are not specific to the environments we train on; they transfer to downstream evals, and the final model owes them to both post-training stages (supervised fine-tuning and RL), as shown in the charts.

What comes next

This is only the beginning. ML4 is the first milestone on the roadmap funded by our €3 billion Series D — the largest equity round ever raised by a European technology company. That capital is already being put to work: we are significantly scaling up our compute capacity in our own European datacenters, and much more is coming online in the months ahead.

More compute means more training. The reinforcement learning run behind this preview is still in flight, and the model is showing no signs of saturation — there is substantial headroom ahead. As we scale up training on our expanded infrastructure, we expect large and rapid improvements in the weeks and months to come.

We will release the weights by the end of the month, along with more details on the architecture, additional benchmarks, and our post-training methodology. And ML4 is only the foundation: it will serve as the base for a new generation of specialized and optimized Mistral models, built for the industries and workloads our customers care about most.

The pace of progress from here will be fast. Stay tuned.

Mistral Large 4

New

Open-weight hybrid instruct-and-reasoning MoE with multimodal input; unifies instruction, reasoning, and agentic capabilities in a single model, state-of-the-art among open weights on cybersecurity, finance, and manufacturing, natively fluent in 160+ languages.

Multimodal

Reasoning

Coding

Agentic

Cybersecurity

Input (/M tokens)

$1.36

Output (/M tokens)

$4.18

Read more

0%

View original article on mistral.ai

Most Recent

SoftBank-Backed DayOne Data Centers Plans to Raise Up to $5.0 Bln in U.S. IPO

The data-center operator is also backed by Coatue Management, Hillhouse and Citadel — A SoftBank-backed data-center company …

Oct 6, 2026

Report: Apple’s smart home push includes LG-made doorbell, lock, thermostat, more

Bloomberg reports that Apple’s plans for smart home products include devices, sensors, and cameras codeveloped in partnership with LG.

Oct 6, 2026

Anthropic says it found 5,500 verified vulnerabilities in April-October and Glasswing partners found 129K+ in April-July; 33K+ were critical or high severity

Anthropic is expanding a program that allows vetted cybersecurity professionals to test its most powerful AI models with fewer safeguards …

Oct 6, 2026

Waymo Taps Pimco, Blackstone for $5 Billion Loan to Accelerate Growth

Waymo increased the size of its inaugural debt raise to $5 billion, tapping lenders to help fuel the robotaxi company's growth as it expands globally.

Oct 6, 2026

Similar Posts

Introducing Claude Sonnet 5.5

Claude Sonnet 5.5 is a clear upgrade over Claude Sonnet 5, runs 30%+ faster, and costs up to 30% less for most work.

Sep 28, 2026

Anthropic Subscriptions Offer 5x+ More Value Than OpenAI

Limit testing every AI subscription plan from Anthropic, OpenAI, Meta, SpaceXSI, MiniMax, Moonshot, Z.ai, Cursor, and Cognition

Oct 5, 2026

Announcing our partnership with OpenAI

Baseten is partnering with OpenAI to serve open models natively via Codex and the Responses API.

Sep 29, 2026

GPT-6 Astra might be too powerful to understand or control

OpenAI is hailing its new model as “the world’s most intelligent and aligned”, but the details reveal an awareness of being evaluated and an ability to manipulate its visible reasoning

Sep 4, 2026