Guides
Self-hosting Le Chonk AI: hardware preparation
Plan the workload before choosing deployment hardware.
Deployment is not validated
Follow release status first; executable deployment commands need a real checkpoint and test.
Storage is not active compute
Documentation lists 1.05 trillion total and 52 billion active parameters. The active slice does not remove full-weight storage. The arithmetic below assumes all total parameters use the stated precision.
| Hypothetical precision | Raw weight arithmetic | Excluded |
|---|---|---|
| 16-bit | About 2.1 TB decimal | KV cache, buffers, metadata and extra components. |
| 8-bit | About 1.05 TB decimal | Quantization scales, format overhead and quality impact. |
| 4-bit | About 525 GB decimal | Support, overhead and accuracy loss. |
Workload variables
| Variable | Effect | Record |
|---|---|---|
| Precision | Storage and quality depend on kernels and format. | Revision and quantization provenance. |
| Context | Prefill and KV-cache requirements. | Prompt length and peak memory. |
| Concurrency | Cache and scheduling pressure. | Sessions and p50/p95 latency. |
| Interconnect | Data movement can dominate. | Topology, network, host RAM and disk. |
An advertised context maximum is not an economical serving guarantee.
Validation sequence
- Verify checkpoint, license and revision.
- Choose a framework with explicit architecture support; pin driver, framework and container versions.
- Run a short single-user request. Record memory, first-token latency and output speed.
- Repeat expected context lengths and image tasks; increase concurrency gradually.
- Compare outputs with the hosted model on identical inputs; record failures.
- Estimate cost per successful task including infrastructure and engineering.
Run record
checkpoint / revision / license:
framework / container / drivers:
GPU / count / interconnect / host RAM:
precision / prompt tokens / output tokens:
concurrency / peak memory:
first-token latency / output speed:
failures / quality checks / hourly cost:Troubleshooting
- Out of memory: isolate context/concurrency and the failed memory pool.
- Unsupported architecture: check explicit framework support, not another model’s configuration.
- Slow output: separate loading, prefill and decoding; inspect offload/bandwidth.
- Image failures: verify the processor and vision component match the revision.
- Tool-format differences: validate templates and API compatibility. No Large 4 deployment fixes are claimed tested.
Sources & verification
Research snapshot: . Source statements are dated; API access and prices can change.
- Mistral Large 4 model documentation
- Mistral release announcement
- Official Mistral Hugging Face organization
Original model runs are marked “not run” unless a trace is provided. Read our editorial method.