github Archived

FMInference/FlexLLMGen

Running large language models on a single GPU for throughput-oriented scenarios.

2 awesome lists

Quick read

Stars
9,357
Forks
592
Open issues
58
Commits
107

Activity and growth

Latest capture 2026-08-16 03:04

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
-10
Stored snapshots
6

Classification

Metadata

Language
Python
License
Apache-2.0
Default branch
main
Created
2023-02-15
First commit
2023-02-15
Last pushed
2024-10-28
GitHub updated
2026-08-14
Last synced
2026-08-16 03:04
Stack scanned
2026-08-16 03:04
Archived
Yes

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

6 observed captures since 2026-05-27. Observed captures are shown by default.

Stars from first capture -10

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

bobazooba/xllm

🦖 X—LLM: Cutting Edge & Easy LLM Finetuning

410 stars
Python 1 awesome list

Tiiny-AI/PowerInfer

High-speed Large Language Model Serving for Local Deployment

9,692 stars
C++ 2 awesome lists

hpcaitech/ColossalAI

Making large AI models cheaper, faster and more accessible

41,430 stars
Python 4 awesome lists

OptimalScale/LMFlow

An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.

8,486 stars
Python 2 awesome lists

ikawrakow/ik_llama.cpp

llama.cpp fork with additional SOTA quants and improved performance

2,989 stars
C++ 1 awesome list