LLM / MODEL FILE
Command R 35B
Cohere model for RAG and tool use. Memory figures are estimates, not measured benchmarks.
Parameters35.0B
Q4 estimate26.0 GB
WorkloadLLM
Best proofTest exact runtime
QUANTIZATION
Memory scenarios
MEMORY FILE
Q4_K_M
common balance of size and quality. Estimated requirement: 26.0 GB.
Inspect →MEMORY FILEQ5_K_M
more quality with a larger footprint. Estimated requirement: 31.4 GB.
Inspect →MEMORY FILEQ8_0
near-full quality with high memory use. Estimated requirement: 45.0 GB.
Inspect →MEMORY FILEFP16
full half-precision weights. Estimated requirement: 88.4 GB.
Inspect →CLOSEST FITS
Hardware near the Q4 line
TIGHT
RTX 3090 24GB
used-market 24 GB local AI favorite
24 GB VRAMInspect →TIGHTRTX 4090 24GB
strong 24 GB inference and image generation
24 GB VRAMInspect →TIGHTRadeon RX 7900 XTX 24GB
24 GB value with a more selective software path
24 GB VRAMInspect →COMFORTABLERTX 5090 32GB
fast 32 GB consumer flagship
32 GB VRAMInspect →OFFLOADRTX 4060 Ti 16GB
efficient 16 GB consumer option
16 GB VRAMInspect →OFFLOADRTX 4070 Ti Super 16GB
faster 16 GB card for mid-size models
16 GB VRAMInspect →Runtime, drivers, memory bandwidth, context, batch size, offloading, and model architecture all matter.