• Cloud

Inference to build your frontier in Europe.

Access and deploy frontier models on one API, powered by Mistral Cloud.

Get API key

Medium

Read our docs

from mistralai import Mistral

model = "mistral-medium-latest"

client = Mistral(api_key="")

chat_response = client.chat.complete(
    model=model,
    messages=[
        {
            "role": "user",
            "content": "Explain a merkle tree in two sentences.",
        },
    ],
)

print(chat_response.choices[0].message.content)

Add your API key

Create a key in Mistral Studio, then paste it below to complete the snippet.

Open Studio
Snowflake (black)
Harvey (black)
MongoDB (black)

Data sovereignty for every request.

Choose where it runs.

Process requests in the EU or, for supported models, US for data residency and lower latency. Each region is a fully independent deployment with its own capacity and data boundary.

Know who handles the data.

Mistral serves every model itself, so third-party open models run safely, under the same regional controls and service commitments as ours.

Limit what’s retained.

Turn on zero data retention or deploy privately for the strictest regulatory requirements.

"We use Mistral Small models in Firefox to bring fast, accurate, multilingual AI in the browser to millions of users."

Alexander Osipenko, Staff Machine Learning Engineer, Firefox.

"Mistral allows us to run open models under strict regional controls and service commitments, making it easy for us to maintain data residency and compliance requirements while furthering our commitment to open source."

Matan Grinberg, CEO and cofounder, Factory.

Choose from our most
powerful models and APIs.

Mistral models

Mistral Medium 3.5

Merged dense model unifying instruction, reasoning, and coding; configurable reasoning effort with multimodal input.

Input (/M tokens)

$1.5

Output (/M tokens)

$7.5

OCR 4.1

New

The world’s best document extraction and understanding model.

OCR

$4

/ 1000 pages

Document AI

$5

/ 1000 pages

Voxtral Mini Transcribe 2

State-of-the-art batch audio transcription.

Audio Input/min

$0.003

Available on

/v1/audio/transcriptions

Third party models

GLM 5.3

New

Third-party model specializing in long-context agentic workflows and coding.

Input (/M tokens)

$1.4

Output (/M tokens)

$4.4

Explore

See models side-by-side to find the right one for your use case.

Compare

Enterprise APIs

Enterprise APIs

Regional data processing controls, system-level SLAs, increased rate limits, and premium support.

Get in touch

Explore developer resources.

API reference

The complete REST reference, with schemas and examples.

SDKs

Official Mistral clients for Python and TypeScript.

Cookbooks

Working examples for agents, retrieval, OCR, and more.

Join our community

Get help and share what you build.

Frequently asked questions.

Yes, you can call our regional inference endpoints in the EU or, for supported models, US (api.eu.mistral.ai, api.us.mistral.ai) to process requests for supported models in your preferred region for data residency or lower latency, at a 10% upcharge. Without a regional endpoint, requests use the global endpoint. For production traffic, you can request Priority Tier, which puts your requests ahead of standard traffic, backed by an uptime SLA.

We serve third-party models ourselves, so they run safely on GPU infrastructure managed by Mistral, without modifications and under the same regional controls and service commitments as our own. Your requests are processed by Mistral. Third-party models follow a shorter deprecation schedule; see the model lifecycle policy.

Data sent through our API isn’t used for model training, and zero data retention is available on request.

Yes, in Studio, you can test-prompt general models in the agent playground, or try documents and speech in the dedicated OCR and audio playgrounds.

For general tasks, productivity, and coding, Mistral Medium and GLM by Z.ai (hosted by Mistral) are the strongest models. For structured document extraction—Mistral OCR, and Voxtral for transcription, voice cloning, and voice agents. For cost-sensitive work, Mistral Small as well as Ministral are the lightest models. Compare them in the models overview.

Most models are priced per million tokens, with input (your prompts) and output (responses) counted separately. For example, Mistral Large costs $0.5 /M tokens in and $1.5 /M tokens out. Batch processing, for high-volume work, cuts the price by 50%, and cached input tokens reduce input cost by up to 90% for repeated prompts. A few APIs are priced differently: OCR is per 1,000 pages, speech models are per minute, and tool APIs are priced per call.

Yes, our open-weight models can run anywhere—download them from Hugging Face and host them on your own hardware. Licenses vary by model, commonly between Apache 2.0 or a modified MIT, and can change. Enterprise customers can deploy our models on-premises or in their own cloud, including hyperscalers like AWS, Azure, and Google Cloud.

Start building

Prototype → Ship → Scale
with Mistral.

gradient background
cat sitting