github Actively maintained

EricLBuehler/candle-vllm

Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.

1 awesome list

Quick read

Stars
712
Forks
88
Open issues
29
Commits
612

Activity and growth

Latest capture 2026-08-15 03:05

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+45
Stored snapshots
6

Classification

Technology stack

Metadata

Language
Rust
License
MIT
Default branch
master
Created
2023-10-29
First commit
2023-10-29
Last pushed
2026-08-02
GitHub updated
2026-08-11
Last synced
2026-08-15 03:05
Stack scanned
2026-08-15 03:05
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

6 observed captures since 2026-05-27. Observed captures are shown by default.

Stars from first capture +45

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

ikawrakow/ik_llama.cpp

llama.cpp fork with additional SOTA quants and improved performance

2,989 stars
C++ 1 awesome list

NVIDIA/Model-Optimizer

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

3,371 stars
Python 2 awesome lists

kaito-project/aikit

🏗️ Fine-tune, build, and deploy open-source LLMs easily!

535 stars
Go 2 awesome lists