# Mistral Large 4 hosted preview vs weights: choose by your constraints

> Choose Mistral Large 4’s hosted preview only if data, latency, and service requirements fit; wait for the planned weights when self-run control is mandatory.
> By Dave · 2026-10-07
> Source: https://otf-kit.dev/blog/mistral-large-4-preview-vs-weights

Mistral Large 4 creates two different decisions for a team evaluating it: whether to use Mistral’s hosted preview now, and whether to plan around weights that Mistral says are coming by the end of October. The first option is available today through the preview API. The second is still a future release. An announcement about planned weights is not yet a deployment artifact, license, or operating plan.

For most teams, the practical choice depends on where prompts may be processed, what latency the product needs, and who will own inference operations. If a hosted model is allowed and the current API meets your service requirements, the preview gives you a route to try the model now. If policy requires your own infrastructure, or your product cannot depend on a hosted endpoint, wait until the weights and their deployment terms are actually published.

![Hosted preview is available now while the self-run path waits for the planned weights](https://cdn.otf-kit.dev/blog/mistral-large-4-preview-vs-weights/inbody1-hosted-vs-weights-20261007.png)

## What the announcement makes available

Mistral announced Large 4 on October 6, 2026. Its announcement describes a public preview API available through Mistral Studio and says the weights are planned for release by the end of the month. Those statements describe two different delivery paths, at two different stages. You can request access to the hosted preview now; you cannot download and operate the announced weights yet.

That distinction matters when a product decision has a deadline. A team can evaluate the hosted route against its existing data policy and service requirements. It cannot responsibly promise a self-run Large 4 deployment based only on the announcement. The actual files, license, serving instructions, hardware demands, and production support requirements still need to be examined when they are released.

Mistral’s announcement positions the model for coding, agentic workflows, and multimodal use. Those claims can help decide whether the preview is relevant to your product. They do not settle where your requests will be processed, what terms apply to your account, or whether your own team can run the later weights at the latency and concurrency your users expect. Treat model capability and deployment fit as separate questions. The launch notice is the source for the preview status and stated weights timeline.

## Use the hosted preview when the provider boundary fits

A hosted preview is the shortest path to access. Mistral operates the inference service; your application sends requests to that service and receives responses. This can be a sensible choice when your organization permits that data flow, the preview terms meet your use case, and the endpoint’s measured behavior fits the product.

Before sending real product traffic, make a short checklist for the specific account and workload:

- **Data approval:** Identify the exact prompt and response data the feature sends. Confirm that your current policy permits those fields to go to a hosted model provider. Redact secrets and exclude customer content until the data owner approves the path.
- **Retention and training terms:** Read the terms that apply to the account and endpoint. Verify what happens to request and response content, operational logs, and abuse-monitoring data. Do not infer these terms from the phrase “public preview.”
- **Processing location:** If location is a requirement, confirm the actual endpoint and model support the geography you need. Check whether the requirement applies only to inference or also to account metadata and logs.
- **Latency and availability:** Measure request time from the regions where your users and backend run. Record p50 and p95 latency during the workload you intend to serve, and check rate limits, capacity, support terms, and failure behavior.
- **Cost at your traffic shape:** Estimate input and output usage from real request sizes, including retries and long context. Use the live price and any regional or preview-specific terms shown for your account; launch pricing can change.
- **Feature fit:** Confirm the preview endpoint supports the request modes your product depends on, such as streaming, structured output, or tool calls. A model name appearing in a model catalog does not prove every API feature is available in every deployment.

If any required answer is unclear, use synthetic or already approved data while you resolve it. The preview is useful for checking service fit, but trying it does not require routing customer prompts through it.

A data boundary should be explicit in the application. Keep provider selection in configuration, limit which tasks can call the preview, and log the provider and model for each request. This makes it easier to shut the route off if terms, availability, or product requirements change. The same boundary also helps if you later compare a hosted endpoint with your own deployment; see the [AI provider portability guide](/blog/ai-provider-portability) for an implementation pattern.

## Wait for weights when control is a hard requirement

Waiting is the safer decision when a hosted endpoint conflicts with a non-negotiable requirement. Examples include workloads that must run inside a controlled network, data that cannot be sent to a third-party inference service, or products that need to keep operating when an external provider is unavailable. A preference for self-hosting is not enough by itself; identify the specific control the hosted service cannot meet.

Weights would change where your team can run inference, but they do not make deployment automatic. Once files are actually available, verify the license and permitted use, supported serving software, memory and accelerator requirements, throughput at expected concurrency, and how model updates will be delivered. Estimate the people and time needed for deployment, monitoring, incident response, and capacity planning.

That operational work belongs in the cost comparison. Hosted inference has a usage bill and provider terms to review. Self-run inference adds hardware or compute capacity, utilization risk, deployment work, monitoring, and on-call ownership. Compare both paths over the same expected workload and support window. Avoid a cost-per-token comparison that excludes idle capacity or the engineering time needed to keep a service available.

Latency also needs a product-specific target. Hosted latency includes the network path to the provider and the provider’s queue and processing time. Running weights closer to your application may reduce some of that distance, but it does not guarantee a faster response: hardware choice, batching, concurrency, context length, and serving configuration all matter. Until the weights and serving guidance exist, a self-run latency estimate is a hypothesis, not a measured result.

![Dex and Nova compare hosted inference with a self-run server across data, latency, hardware, and operations](https://cdn.otf-kit.dev/blog/mistral-large-4-preview-vs-weights/inbody2-hosted-self-run-20261007-b.png)

## Choose by the constraint that can block launch

Use the hosted preview now if all of these statements are true: your data policy allows the request path, the preview terms are acceptable, the needed features are available, and measured latency and cost fit your product. Start with the narrowest approved use case and keep the route easy to disable.

Wait for the weights if one of those requirements is a hard blocker and the hosted service cannot satisfy it. While waiting, document the exact deployment conditions you need to prove: license, network boundary, hardware budget, throughput, recovery behavior, and an owner for ongoing operations. Revisit the decision when the actual release materials are available, not when a release date is announced.

If you have no hard hosted-service restriction, waiting still has a cost: you defer access to a preview that is available now. If self-hosting is mandatory, trying the API will not answer the central question, and it may create a data review the project does not need. Make the decision from the requirement that could stop launch, rather than from a general preference for open weights or a benchmark headline.

The practical verdict today is conditional. Use the hosted preview only for workloads whose data and service requirements you have cleared. Keep sensitive or location-restricted workloads out until the relevant terms and endpoint behavior are verified. Plan a self-run path only after Mistral publishes the weights, license, and serving details. The date in an announcement is a target to check again, not an artifact to build a deployment promise around.

## Sources

- [Mistral: Introducing Mistral Large 4](https://mistral.ai/news/mistral-large-4/) — public preview availability and the stated plan to release weights by the end of October 2026.
