Le Chonk AI vs Kimi K3: documented capabilities
Document evidence and chart arithmetic. A dated documentation comparison, with a task you can reproduce.
Versions and evidence
| Item | Mistral Large 4 Preview | Kimi K3 |
|---|---|---|
| Model ID | mistral-large-4 | kimi-k3 |
| Provider | Mistral direct API | Kimi / Moonshot direct API |
| Access | Public preview | Documented API model; account access untested |
| Context | 1 million tokens documented; 512k at AA launch test | 1M tokens |
| Input / output | Text + images / Text | Text, image and video input |
| USD / 1M input / cache / output | $1.36 / $0.14 / $4.18 standard | Numeric K3 rates not reliably extracted; verify provider table |
| Weights / license | Planned weights; final license unverified | Exact K3 weight repository and license not verified here. |
Read Kimi model list, Kimi K3 guide, Kimi pricing for the other model. Listed rates are not normalized quotes: promotions, cache writes, long prompts and channel terms may change them.
Read the same operations note and chart, then separate evidence from calculations.
Use the operations note and quarterly chart. Identify the approved delivery date, compute Q1-to-Q4 unit growth, and quote the supporting paragraph and chart labels. If a claim is absent, say not stated. Is a budget approved?
Document: /fixtures/operations-note.pdf
Chart: /fixtures/quarterly-chart.svg- Approved delivery date is October 12, 2026 from P2; budget approval is not stated.
- Chart values: Q1 120, Q2 160, Q3 140, Q4 200; Q1-to-Q4 growth is approximately 66.7%.
- Use the same rasterized chart resolution and document extraction method. Record preprocessing and citation errors.
These checks are authored reference expectations, not either model’s output. For linked fixtures, download the operations PDF and chart.
Common evaluation protocol
- Freeze the provider, exact model ID, date, reasoning setting and system prompt. Use a fresh conversation for every run.
- Give both models the same task input and tool permissions. Record any provider-specific preprocessing and context truncation.
- Use three attempts per task, a 60-second request timeout, and a fixed tool-step budget. Preserve failed/refused attempts as results.
- Validate with the published checks. Record wall time, billed tokens, cache hits, tool fees and any human edits. Compare cost per accepted result.
| Run evidence | Mistral Large 4 Preview | Other model |
|---|---|---|
| Model inference | Not run | Not run |
| Raw output / tool trace | Not available | Not available |
| Latency / billed usage | Not measured | Not measured |
| Successful-task cost | Not measured | Not measured |
Migration and billing caveats
Kimi’s current model list marks K2.5 discontinued. This comparison uses K3, not an old K2.5 alias. K3 cache writes and cache hits have distinct billing semantics; extracted document content still contributes to input tokens.
Cost per accepted task = all billed inference, tools and retries divided by accepted results. Keep output token usage and human corrections in the record; a lower per-token rate can still produce a more expensive job.
How to make a choice
Both have documented long-context and image input. Start with evidence accuracy, not the nominal context maximum. Video-input needs point toward evaluating Kimi’s documented endpoint; Large 4 video support is not established here.
The next useful evidence is a preserved run trace from your own workload. Until it exists, this is a shortlist based on documentation. Browse task exercises.
Sources & verification
Research snapshot: . Source statements are dated; API access and prices can change.
- Mistral Large 4 model documentation
- Artificial Analysis launch evaluation
- Kimi model list
- Kimi K3 guide
- Kimi pricing
Original model runs are marked “not run” unless a trace is provided. Read our editorial method.