FMInference/FlexLLMGen
Running large language models on a single GPU for throughput-oriented scenarios.
Run Mixtral-8x7B models in Colab or consumer desktops
Appears on
Quick read
Latest capture 2026-09-05 10:53
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
1 observed capture since 2026-09-05. Observed captures are shown by default.
Stars from first capture 0
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Running large language models on a single GPU for throughput-oriented scenarios.
Official inference library for Mistral models
text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
Run LLMs with MLX
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM