Quantization reduces numerical precision of model weights or activations to shrink memory use and often speed inference, with possible quality tradeoffs.
AI Glossary
Quantization
Coverage in development
Quantization reduces numerical precision of model weights or activations to shrink memory use and often speed inference, with possible quality tradeoffs.
Definition
Plain English explanation
You store numbers with fewer bits so the model is smaller and cheaper to run.
Related glossary terms
Last reviewed
Sources
- NIST AI Risk Management Framework — risk vocabulary context for AI systems
- Brel Digital company and technology profiles — applied usage evidence where tagged
Correction request
If a technology assignment or hub description is inaccurate, submit a correction via the Corrections Policy.
All technologies → · AI Models → · APIs & SDKs → · Integrations → · Compliance → · Browse all companies → · Explore industries → · Compare →