LLM / MODEL FILE
Codestral 22B
Mistral model for code completion and generation. Memory figures are estimates, not measured benchmarks.
Parameters22.0B
Q4 estimate16.9 GB
WorkloadLLM
Best proofTest exact runtime
QUANTIZATION
Memory scenarios
MEMORY FILE
Q4_K_M
common balance of size and quality. Estimated requirement: 16.9 GB.
Inspect →MEMORY FILEQ5_K_M
more quality with a larger footprint. Estimated requirement: 20.3 GB.
Inspect →MEMORY FILEQ8_0
near-full quality with high memory use. Estimated requirement: 28.8 GB.
Inspect →MEMORY FILEFP16
full half-precision weights. Estimated requirement: 56.1 GB.
Inspect →CLOSEST FITS
Hardware near the Q4 line
TIGHT
RTX 4060 Ti 16GB
efficient 16 GB consumer option
16 GB VRAMInspect →TIGHTRTX 4070 Ti Super 16GB
faster 16 GB card for mid-size models
16 GB VRAMInspect →TIGHTRTX 4080 Super 16GB
high throughput but still limited to 16 GB
16 GB VRAMInspect →TIGHTRTX 5060 Ti 16GB
current mid-range 16 GB option
16 GB VRAMInspect →TIGHTRTX 5070 Ti 16GB
current performance card with 16 GB ceiling
16 GB VRAMInspect →TIGHTRTX 5080 16GB
fast compute with a 16 GB memory ceiling
16 GB VRAMInspect →Runtime, drivers, memory bandwidth, context, batch size, offloading, and model architecture all matter.