github Actively maintained

informatico-madrid/blackwell-linux-infra-optimizer

Optimized vLLM deployment for NVIDIA Blackwell (RTX 5090) on Linux Kernel 6.14. Resolves SM_120 kernel incompatibilities, P2P deadlocks, and memory fragmentation for high-performance LLM inference.

1 awesome list

Quick read

Stars
6
Forks
1
Open issues
0
Commits
11

Activity and growth

Latest capture 2026-08-02 03:11

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
+1
Stored snapshots
6

Classification

Metadata

Language
Dockerfile
Default branch
main
Created
2026-01-16
First commit
2026-01-16
Last pushed
2026-01-17
GitHub updated
2026-07-06
Last synced
2026-08-02 03:11
Stack scanned
2026-08-02 03:11
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

6 observed captures since 2026-05-25. Observed captures are shown by default.

Stars from first capture +1

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

llm-d/llm-d

Achieve state of the art inference performance with modern accelerators on Kubernetes

3,959 stars
Shell 2 awesome lists

EricLBuehler/candle-vllm

Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.

712 stars
Rust 1 awesome list

bobazooba/xllm

🦖 X—LLM: Cutting Edge & Easy LLM Finetuning

410 stars
Python 1 awesome list

NVIDIA/Model-Optimizer

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

3,371 stars
Python 2 awesome lists

vllm-project/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

90,045 stars
Python 4 awesome lists

defilantech/LLMKube

Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.

196 stars
Go 1 awesome list