Le Chonk AI benchmarks: claims, independent tests & limits
A score is evidence about a test, not a promise about your workflow.
Independent launch evaluation
| Metric | Large 4 Preview | Source / date |
|---|---|---|
| AA Intelligence Index | 38 | Artificial Analysis launch evaluation · Oct 6, 2026 |
| AA Cyber Index | 50 | Artificial Analysis launch evaluation · Oct 6 |
| CyberGym-E2E-AA | 82% | Artificial Analysis launch evaluation · Oct 6 |
| GDP.pdf | 19% | Artificial Analysis launch evaluation · Oct 6 |
| Intelligence task cost | $1.13 standard; $0.57 launch | Artificial Analysis launch evaluation · its task distribution, not every request |
Artificial Analysis tested the preview at 512k context. These scores use different scales; do not average them or infer percentages from index values. The launch report places GDP.pdf below Kimi K3’s 22% in that test.
Official claims are a separate evidence layer
| Domain | Announcement reports | Qualification |
|---|---|---|
| Coding | DeepSWE v1.1: 61.7%; Terminal-Bench 4: 28.3%; SWE-Atlas-QnA: 59.4% | Vendor-reported tests; benchmark versions and harness matter. |
| Workflow | AutomationBench: 59.9% | Not a guaranteed completion rate for your business. |
| Vision | Dense 200: 42% | Grounding score, not general image accuracy. |
| Finance / legal / manufacturing | Enterprise and partner evaluation examples | Not reproduced by this site; demo claims do not establish task accuracy. |
Source: Mistral release announcement. We do not normalize unlike tests into one ranking.
What can change a result
- Harness, tools, time/step budgets, reasoning settings and model revision.
- Refusals and moderation: a blocked task can lower a score even when the task is technically within a model’s capability.
- Document rasterization, image count, input resolution and OCR preprocessing.
- Cache usage and long outputs: reported task cost is distinct from token price.
- Preview changes and evaluation dates: an old score may not describe a later checkpoint.
Our test status and method
- Freeze the provider, exact model ID, date, reasoning setting and system prompt. Use a fresh conversation for every run.
- Give both models the same task input and tool permissions. Record any provider-specific preprocessing and context truncation.
- Use three attempts per task, a 60-second request timeout, and a fixed tool-step budget. Preserve failed/refused attempts as results.
- Validate with the published checks. Record wall time, billed tokens, cache hits, tool fees and any human edits. Compare cost per accepted result.
Exercise inputs and acceptance checks are published under Use Cases and each comparison. No original model outputs, accuracy percentages or measured latency are available yet. A completed record must include the full response/trace, validation command, usage and any manual edits.
Read benchmark names precisely
CyberGym-E2E-AA in the independent report is not automatically identical to a vendor’s CyberGym setup. Terminal-Bench 4 results should not be compared directly to Terminal-Bench 2 or 3. A provider-specific reduced-moderation evaluation is not proof that every public API request behaves the same way.
Sources & verification
Research snapshot: . Source statements are dated; API access and prices can change.
Original model runs are marked “not run” unless a trace is provided. Read our editorial method.