github Actively maintained

dvmazur/mixtral-offloading

Run Mixtral-8x7B models in Colab or consumer desktops

1 awesome list

Quick read

Stars
2,334
Forks
225
Open issues
26
Commits
86

Activity and growth

Latest capture 2026-09-05 10:53

Stars · last 7 days
No history
Commits · last 7 days
No history
Stars since tracking
0
Stored snapshots
1

Classification

Metadata

Language
Python
License
MIT
Default branch
master
Created
2023-12-15
First commit
2023-12-15
Last pushed
2024-04-08
GitHub updated
2026-08-31
Last synced
2026-09-05 10:53
Stack scanned
2026-09-05 10:53
Archived
No

AI development signals

0 paths

Agent instructions and tool configuration found in this repository.

No config files detected.

Growth history

Tracked growth

1 observed capture since 2026-09-05. Observed captures are shown by default.

Stars from first capture 0

Chart data

Observed captures only

Time horizon

All tracked data

Custom date range

Stars history

Observed snapshots

Commits history

Observed snapshots

Similar repositories

Nearest indexed repositories by embedding similarity.

FMInference/FlexLLMGen

Running large language models on a single GPU for throughput-oriented scenarios.

9,357 stars
Python 2 awesome lists

zai-org/CogVideo

text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)

13,019 stars
Python 2 awesome lists

NVIDIA/Model-Optimizer

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

3,780 stars
Python 2 awesome lists

vllm-project/llm-compressor

Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM

3,770 stars
Python 2 awesome lists