Inference & serving
llama.cpp
A C/C++ language- and vision-model inference project with quantization, multiple hardware backends, command-line tools and an OpenAI-compatible serving interface.
Overview
Research summary
llama.cpp provides an inference runtime for supported language and vision-language models across a wide range of hardware. Its C/C++ implementation includes quantized execution, CPU and GPU backends, hybrid placement options, and tools for running models interactively or exposing them through a server. Applications can use the library and API surfaces as the model-execution layer beneath a chat interface or agent.
The project itself does not supply a complete coding-agent workflow, and model weights are acquired separately. Supported formats, operators, context sizes, speed, and memory use vary with the chosen model and backend, so hardware fit needs workload-specific validation. The code is MIT licensed, while models and third-party components retain their own terms.
Local inference also does not imply that an application built around the runtime has no external integrations.
Repository summary
- Stars
- 128,414
- Open issues
- Unavailable
- Last push
- 2026-09-16
- Commits, 90 days
- Unavailable
- Repository activity
- Not scored
- Version
- Unavailable
Recorded catalogue figures. View repository data and provenance →
Classification
Pricing & services
Paid services unknown
Whether the provider offers paid products or services has not been established.
Licence scope
Implementation
Recorded implementation details and interfaces for llama.cpp.
Implementation details
C/C++ inference library, hardware backends, CLI and server
- Languages
- C++
- Repository type
- source
Recorded interfaces and capabilities
Licence scope
Repository
Repository snapshots, release information and recorded maintenance signals.
Repository snapshot
- Stars
- 128,414
- Open issues
- Unavailable
- Last push
- 2026-09-16
- Commits, 90 days
- Unavailable
- Repository activity
- Not scored
- Archived
- Not recorded
Repository activity is a snapshot, not a quality or popularity ranking. It combines recent-push freshness (50%), 90-day commits (30%) and issue pressure (20%).
Maintenance and provenance
- Catalogue snapshot
- 2026-10-06
- Stars source
- Recorded fallback
- Last push source
- Recorded fallback
Documentation
Recorded references and research provenance for this entry.
Recorded sources 5
- https://github.com/ggml-org/llama.cpp Project page · Linked repository · Research reference
- https://github.com/ggml-org/llama.cpp/blob/master/README.md Research reference
- https://github.com/ggml-org/llama.cpp/blob/master/LICENSE Research reference
- https://github.com/stars/angelcervera/lists/ai Research reference
- https://api.github.com/repos/ggml-org/llama.cpp Research reference
Research metadata
- Research date
- 2026-09-16
- Catalogue snapshot
- 2026-10-06
Community & social 1 channels
Communities
- GitHub Discussionsgithub.com/ggml-org/llama.cpp/discussionsOfficialDiscussion
Link details for GitHub Discussions
- Access
- Public
- Last checked