jundot/omlx
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.
Quick read
Latest capture 2026-09-22 04:00
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
2 observed captures since 2026-09-21. Observed captures are shown by default.
Stars from first capture +2237
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
llama.cpp fork with additional SOTA quants and improved performance
AirLLM 70B inference with single 4GB GPU
An Open-source Toolkit for LLM Development
Inference code for Llama models
Build, personalize and control your own LLMs. From data pre-processing to fine-tuning, xTuring provides an easy way to personalize open-source LLMs. Join our discord community: https://discord.gg/TgHXuSJEk6