Cloud
Inference to build your frontier in Europe.
Medium
from mistralai import Mistral
model = "mistral-medium-latest"
client = Mistral(api_key="")
chat_response = client.chat.complete(
model=model,
messages=[
{
"role": "user",
"content": "Explain a merkle tree in two sentences.",
},
],
)
print(chat_response.choices[0].message.content)from mistralai import Mistral
client = Mistral(api_key="")
agent = client.beta.agents.create(
model="mistral-medium-latest",
name="Websearch Agent",
description="Answers with up-to-date information from the web.",
tools=[{"type": "web_search"}],
)
print(agent.id)from mistralai import Mistral
with Mistral(
api_key='',
) as mistral:
ocr_response = mistral.ocr.process(
model='mistral-ocr-latest',
document={
'type': 'document_url',
'document_url': 'https://arxiv.org/pdf/2201.04234',
},
)
print(ocr_response.pages[0].markdown)from mistralai import Mistral
with Mistral(
api_key='',
) as mistral:
response = mistral.audio.transcriptions.complete(
model='voxtral-mini-latest',
)
print(response.text)from mistralai import Mistral
with Mistral(
api_key='',
) as mistral:
response = mistral.audio.speech.complete(
model='voxtral-mini-latest',
input='Today is going to be a great day.',
)
print(response)Data sovereignty for every request.
Choose from our most
powerful models and APIs.
Mistral models
Mistral Medium 3.5
Merged dense model unifying instruction, reasoning, and coding; configurable reasoning effort with multimodal input.
Input (/M tokens)
$1.5
Output (/M tokens)
$7.5
OCR 4.1
New
The world’s best document extraction and understanding model.
OCR
$4
/ 1000 pages
Document AI
$5
/ 1000 pages

Voxtral Mini Transcribe 2
State-of-the-art batch audio transcription.
Audio Input/min
$0.003
Available on
/v1/audio/transcriptions
Enterprise APIs
Enterprise APIs
Regional data processing controls, system-level SLAs, increased rate limits, and premium support.
Explore developer resources.
Frequently asked questions.
Yes, you can call our regional inference endpoints in the EU or, for supported models, US (api.eu.mistral.ai, api.us.mistral.ai) to process requests for supported models in your preferred region for data residency or lower latency, at a 10% upcharge. Without a regional endpoint, requests use the global endpoint. For production traffic, you can request Priority Tier, which puts your requests ahead of standard traffic, backed by an uptime SLA.
We serve third-party models ourselves, so they run safely on GPU infrastructure managed by Mistral, without modifications and under the same regional controls and service commitments as our own. Your requests are processed by Mistral. Third-party models follow a shorter deprecation schedule; see the model lifecycle policy.
Data sent through our API isn’t used for model training, and zero data retention is available on request.
Yes, in Studio, you can test-prompt general models in the agent playground, or try documents and speech in the dedicated OCR and audio playgrounds.
For general tasks, productivity, and coding, Mistral Medium and GLM by Z.ai (hosted by Mistral) are the strongest models. For structured document extraction—Mistral OCR, and Voxtral for transcription, voice cloning, and voice agents. For cost-sensitive work, Mistral Small as well as Ministral are the lightest models. Compare them in the models overview.
Most models are priced per million tokens, with input (your prompts) and output (responses) counted separately. For example, Mistral Large costs $0.5 /M tokens in and $1.5 /M tokens out. Batch processing, for high-volume work, cuts the price by 50%, and cached input tokens reduce input cost by up to 90% for repeated prompts. A few APIs are priced differently: OCR is per 1,000 pages, speech models are per minute, and tool APIs are priced per call.
Yes, our open-weight models can run anywhere—download them from Hugging Face and host them on your own hardware. Licenses vary by model, commonly between Apache 2.0 or a modified MIT, and can change. Enterprise customers can deploy our models on-premises or in their own cloud, including hyperscalers like AWS, Azure, and Google Cloud.


