langroid/langroid
Harness LLMs with Multi-Agent Programming
LiteRT-LM is Google's production-ready, high-performance, open-source inference framework for deploying Large Language Models on edge devices.
Appears on
Quick read
Latest capture 2026-08-03 03:03
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
6 observed captures since 2026-05-25. Observed captures are shown by default.
Stars from first capture +888
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Harness LLMs with Multi-Agent Programming
A minimal Python framework for building custom AI inference servers with full control over logic, batching, and scaling.
Fast, flexible LLM inference
🏗️ Fine-tune, build, and deploy open-source LLMs easily!
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.