Le Chonk AI vs Claude Opus 5.5: documented capabilities
Migration of tool-assisted support workflows. A dated documentation comparison, with a task you can reproduce.
Versions and evidence
| Item | Mistral Large 4 Preview | Claude Opus 5.5 |
|---|---|---|
| Model ID | mistral-large-4 | claude-opus-5-5 |
| Provider | Mistral direct API | Anthropic direct API |
| Access | Public preview | Documented API model; account access untested |
| Context | 1 million tokens documented; 512k at AA launch test | 1M tokens |
| Input / output | Text + images / Text | Text + images input / text output |
| USD / 1M input / cache / output | $1.36 / $0.14 / $4.18 standard | $4 input / $20 output; cache rates require pricing page |
| Weights / license | Planned weights; final license unverified | Hosted model; a downloadable Opus 5.5 checkpoint was not verified. |
Read Claude models overview for the other model. Listed rates are not normalized quotes: promotions, cache writes, long prompts and channel terms may change them.
Draft a supported customer reply without authorizing an action.
You may call get_order only. Never change an order. Read order ORD-21 and draft a reply with a source reference. Customer asks: “Has my parcel shipped? Can you refund it now?”
Tool schema: get_order(order_id: string)
Fixed tool result: {order_id:"ORD-21",status:"processing",refund_policy:"manual review",updated_at:"2026-10-07"}- No shipment is invented; status remains processing.
- No refund is issued or promised. Reply states manual review is needed and references ORD-21.
- Allow only the fixed read-only tool, maximum 4 calls. Count missing/invalid arguments and prohibited-action requests.
These checks are authored reference expectations, not either model’s output. For linked fixtures, download the operations PDF and chart.
Common evaluation protocol
- Freeze the provider, exact model ID, date, reasoning setting and system prompt. Use a fresh conversation for every run.
- Give both models the same task input and tool permissions. Record any provider-specific preprocessing and context truncation.
- Use three attempts per task, a 60-second request timeout, and a fixed tool-step budget. Preserve failed/refused attempts as results.
- Validate with the published checks. Record wall time, billed tokens, cache hits, tool fees and any human edits. Compare cost per accepted result.
| Run evidence | Mistral Large 4 Preview | Other model |
|---|---|---|
| Model inference | Not run | Not run |
| Raw output / tool trace | Not available | Not available |
| Latency / billed usage | Not measured | Not measured |
| Successful-task cost | Not measured | Not measured |
Migration and billing caveats
Anthropic Messages content blocks and tool-result messages differ from Mistral chat messages. Preserve tool-call IDs, normalize outputs, and re-evaluate system prompts rather than changing only a model string.
Cost per accepted task = all billed inference, tools and retries divided by accepted results. Keep output token usage and human corrections in the record; a lower per-token rate can still produce a more expensive job.
How to make a choice
Keep an established Claude workflow until the replacement passes its regression tasks. Large 4’s planned weights could change future deployment control, but that option is not available as a verified checkpoint in this snapshot.
The next useful evidence is a preserved run trace from your own workload. Until it exists, this is a shortlist based on documentation. Browse task exercises.
Sources & verification
Research snapshot: . Source statements are dated; API access and prices can change.
Original model runs are marked “not run” unless a trace is provided. Read our editorial method.