EricLBuehler/mistral.rs
Fast, flexible LLM inference
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
Quick read
Latest capture 2026-08-03 03:09
10 paths
Agent instructions and tool configuration found in this repository.
Agent instructions
GitHub Copilot 9
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +461
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Fast, flexible LLM inference
PettingZoo and Gymnasium bindings for popular reinforcement learning environments outside of Farama
**CHIMERA v3.0** represents the future of natural language processing. It's the **first framework that runs deep learning entirely on OpenGL**, eliminating traditional token-based, transformer, and backpropagation approaches.
Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.
🏗️ Fine-tune, build, and deploy open-source LLMs easily!
AirLLM 70B inference with single 4GB GPU