facebookexperimental/CUTracer
A dynamic binary instrumentation tool for tracing and analyzing CUDA kernel instructions.
Repository profile
eBPF agent and MCP server for GPU causal observability
Repository updates
Get generated ingero-io/ingero development summaries by email, or follow the weekly and monthly RSS feeds.
Sign in to subscribe by email. RSS feeds are public.
Sign in to subscribeTracked growth, recent movement, and commit velocity from stored repository snapshots.
Latest capture 2026-06-05 08:10
1 capture since 2026-06-05
Stars from baseline 0
All tracked data
Frameworks, package managers, ecosystems, and dependency manifests found during catalog scans.
Scanned 2026-06-05 08:10
go.mod
go ecosystem,
42 dependencies
go.sum
go ecosystem,
78 dependencies
python/ingero-annotate/pyproject.toml
python ecosystem,
1 dependency
tests/workloads/requirements.txt
python ecosystem,
6 dependencies
examples/integrations/accelerate/requirements.txt
python ecosystem,
3 dependencies
examples/integrations/deepspeed/requirements.txt
python ecosystem,
3 dependencies
examples/integrations/hf-trainer/requirements.txt
python ecosystem,
3 dependencies
examples/integrations/pytorch-lightning/requirements.txt
python ecosystem,
3 dependencies
Searchable topics, generated tags, and stack labels that explain where this repository fits.
Agent instructions and tool configuration paths found in the repository tree.
Nearest indexed repositories by embedding similarity.
A dynamic binary instrumentation tool for tracing and analyzing CUDA kernel instructions.
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.
Optimized vLLM deployment for NVIDIA Blackwell (RTX 5090) on Linux Kernel 6.14. Resolves SM_120 kernel incompatibilities, P2P deadlocks, and memory fragmentation for high-performance LLM inference.
Performance analysis tools based on Linux perf_events (aka perf) and ftrace
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.