llama.cpp enables efficient LLM inference on CPUs and GPUs with GGUF models and bindings used widely for local deployment.
AI SDK Intelligence
llama.cpp
ggerganov / community · N/A local.
All AI APIs & SDKs → · Official docs →
Editorial overview
Capabilities
- Local LLM inference
- GGUF model support
- Language bindings and server mode
Limitations
- Performance depends on hardware/quantization
- Not identical to original full-precision training stacks
Related technologies
Related glossary terms
Last reviewed
Sources
Correction request
If a technology assignment or hub description is inaccurate, submit a correction via the Corrections Policy.
Knowledge Library → · All technologies → · AI Models → · APIs & SDKs → · Integrations → · Compliance → · Browse all companies → · Explore industries → · Compare →