LLM / MODEL FILE
DeepSeek-R1 Distill Qwen 7B
DeepSeek model for compact distilled reasoning. Memory figures are estimates, not measured benchmarks.
Parameters7.6B
Q4 estimate6.8 GB
WorkloadLLM
Best proofTest exact runtime
QUANTIZATION
Memory scenarios
MEMORY FILE
Q4_K_M
common balance of size and quality. Estimated requirement: 6.8 GB.
Inspect →MEMORY FILEQ5_K_M
more quality with a larger footprint. Estimated requirement: 8.0 GB.
Inspect →MEMORY FILEQ8_0
near-full quality with high memory use. Estimated requirement: 10.9 GB.
Inspect →MEMORY FILEFP16
full half-precision weights. Estimated requirement: 20.4 GB.
Inspect →CLOSEST FITS
Hardware near the Q4 line
COMFORTABLE
RTX 3080 10GB
fast older GPU constrained by 10 GB VRAM
10 GB VRAMInspect →COMFORTABLERTX 3060 12GB
budget used GPU with useful 12 GB capacity
12 GB VRAMInspect →COMFORTABLERTX 4060 Ti 16GB
efficient 16 GB consumer option
16 GB VRAMInspect →COMFORTABLERTX 4070 Ti Super 16GB
faster 16 GB card for mid-size models
16 GB VRAMInspect →COMFORTABLERTX 4080 Super 16GB
high throughput but still limited to 16 GB
16 GB VRAMInspect →COMFORTABLERTX 5060 Ti 16GB
current mid-range 16 GB option
16 GB VRAMInspect →Runtime, drivers, memory bandwidth, context, batch size, offloading, and model architecture all matter.