AI Glossary

Quantization

Coverage in development

Quantization reduces numerical precision of model weights or activations to shrink memory use and often speed inference, with possible quality tradeoffs.

Definition

Quantization reduces numerical precision of model weights or activations to shrink memory use and often speed inference, with possible quality tradeoffs.

Plain English explanation

You store numbers with fewer bits so the model is smaller and cheaper to run.

Last reviewed

Sources

Correction request

If a technology assignment or hub description is inaccurate, submit a correction via the Corrections Policy.

All technologies → · AI Models → · APIs & SDKs → · Integrations → · Compliance → · Browse all companies → · Explore industries → · Compare →