AI in Practice

Models & Evaluation

What was tried, on what task, how it was judged, and what the published research already says.

Everything in Models & Evaluation

  • A VRAM table comparing Q4_K_M, Q3_K_M and IQ2_XS quantization sizes for a 70B model on a 24GB GPU.

    09/17/2026

    Quantizing a 70B Model for One Consumer GPU

    A 2-bit 70B model fits on one 24GB GPU. Whether it's actually worth running is a different question -- what published quantization research says about that tradeoff, and where it still doesn't settle it.