bitsandbytes-foundation/bitsandbytes
Accessible large language models via k-bit quantization for PyTorch.
PyTorch native quantization and sparsity for training and inference
Accessible large language models via k-bit quantization for PyTorch.
Go ahead and axolotl questions
A unified library of SOTA model optimization techniques like quantization, pruning, distillation, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
QLoRA: Efficient Finetuning of Quantized LLMs
PyTorch native post-training library
A PyTorch native platform for training generative AI models
1 capture since 2026-05-25
AI agent config detected
Key config paths
CLAUDE.md