Le Chonk AI vs DeepSeek V4.1 Flash: documented capabilities
Cost, cache hits and extraction accuracy. A dated documentation comparison, with a task you can reproduce.
Versions and evidence
| Item | Mistral Large 4 Preview | DeepSeek V4.1 Flash |
|---|---|---|
| Model ID | mistral-large-4 | deepseek-flash |
| Provider | Mistral direct API | DeepSeek direct API |
| Access | Public preview | Documented API model; account access untested |
| Context | 1 million tokens documented; 512k at AA launch test | 1M tokens |
| Input / output | Text + images / Text | Vision supported; text output |
| USD / 1M input / cache / output | $1.36 / $0.14 / $4.18 standard | Peak $0.30 / $0.006 / $1.20; off-peak half |
| Weights / license | Planned weights; final license unverified | API terms apply to hosted use. Exact V4.1 Flash weight/license not verified here. |
Read DeepSeek models & pricing, DeepSeek change log for the other model. Listed rates are not normalized quotes: promotions, cache writes, long prompts and channel terms may change them.
Extract a ticket decision into a small, verifiable JSON object.
Return only JSON with ticket_id, deadline, approved, evidence. Cite paragraph IDs. Do not infer approval.
[P1] Ticket T-104 requests a data export by October 12, 2026.
[P2] The reviewer requested a scope clarification.
[P3] No approval was issued as of October 7, 2026.- ticket_id is T-104, deadline is 2026-10-12, approved is false.
- Evidence for approval status references P3; response is valid JSON.
- Repeat the same prefix separately to measure actual cache usage. Compare peak and off-peak billing, not assumed discounts.
These checks are authored reference expectations, not either model’s output. For linked fixtures, download the operations PDF and chart.
Common evaluation protocol
- Freeze the provider, exact model ID, date, reasoning setting and system prompt. Use a fresh conversation for every run.
- Give both models the same task input and tool permissions. Record any provider-specific preprocessing and context truncation.
- Use three attempts per task, a 60-second request timeout, and a fixed tool-step budget. Preserve failed/refused attempts as results.
- Validate with the published checks. Record wall time, billed tokens, cache hits, tool fees and any human edits. Compare cost per accepted result.
| Run evidence | Mistral Large 4 Preview | Other model |
|---|---|---|
| Model inference | Not run | Not run |
| Raw output / tool trace | Not available | Not available |
| Latency / billed usage | Not measured | Not measured |
| Successful-task cost | Not measured | Not measured |
Migration and billing caveats
The legacy deepseek-v4-flash alias is routed to V4.1 Flash in the reviewed docs. Record the returned model, date and usage. Peak/off-peak schedules and cache-hit discounts are provider-specific.
Cost per accepted task = all billed inference, tools and retries divided by accepted results. Keep output token usage and human corrections in the record; a lower per-token rate can still produce a more expensive job.
How to make a choice
DeepSeek’s published Flash token rates are lower in this snapshot. That does not prove lower cost per correct extraction. Measure accuracy, output volume, retries and the actual cache-hit fields before routing a workload.
The next useful evidence is a preserved run trace from your own workload. Until it exists, this is a shortlist based on documentation. Browse task exercises.
Sources & verification
Research snapshot: . Source statements are dated; API access and prices can change.
- Mistral Large 4 model documentation
- Artificial Analysis launch evaluation
- DeepSeek models & pricing
- DeepSeek change log
Original model runs are marked “not run” unless a trace is provided. Read our editorial method.