LLM / MODEL FILE
Phi-4 14B
Microsoft model for compact reasoning and coding. Memory figures are estimates, not measured benchmarks.
Parameters14.0B
Q4 estimate11.3 GB
WorkloadLLM
Best proofTest exact runtime
QUANTIZATION
Memory scenarios
MEMORY FILE
Q4_K_M
common balance of size and quality. Estimated requirement: 11.3 GB.
Inspect →MEMORY FILEQ5_K_M
more quality with a larger footprint. Estimated requirement: 13.5 GB.
Inspect →MEMORY FILEQ8_0
near-full quality with high memory use. Estimated requirement: 18.9 GB.
Inspect →MEMORY FILEFP16
full half-precision weights. Estimated requirement: 36.3 GB.
Inspect →CLOSEST FITS
Hardware near the Q4 line
FITS
RTX 3060 12GB
budget used GPU with useful 12 GB capacity
12 GB VRAMInspect →TIGHTRTX 3080 10GB
fast older GPU constrained by 10 GB VRAM
10 GB VRAMInspect →COMFORTABLERTX 4060 Ti 16GB
efficient 16 GB consumer option
16 GB VRAMInspect →COMFORTABLERTX 4070 Ti Super 16GB
faster 16 GB card for mid-size models
16 GB VRAMInspect →COMFORTABLERTX 4080 Super 16GB
high throughput but still limited to 16 GB
16 GB VRAMInspect →COMFORTABLERTX 5060 Ti 16GB
current mid-range 16 GB option
16 GB VRAMInspect →Runtime, drivers, memory bandwidth, context, batch size, offloading, and model architecture all matter.