meta-llama/llama
Inference code for Llama models
A fast inference library for running LLMs locally on modern consumer-class GPUs
Appears on
Quick read
Latest capture 2026-08-03 03:13
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +64
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Inference code for Llama models
llama.cpp fork with additional SOTA quants and improved performance
Running Llama 2 and other Open-Source LLMs on CPU Inference Locally for Document Q&A
A self-hosted, offline, ChatGPT-like chatbot. Powered by Llama 2. 100% private, with no data leaving your device. New: Code Llama support!
Utilities intended for use with Llama models.
llama.go is like llama.cpp in pure Golang!