microsoft/LLMLingua
[EMNLP'23, ACL'24] To speed up LLMs' inference and enhance LLM's perceive of key information, compress the prompt and KV-Cache, which achieves up to 20x compression with minimal performance loss.
General technology for enabling AI capabilities w/ LLMs and MLLMs
Appears on
Quick read
Latest capture 2026-09-07 10:55
2 paths
Agent instructions and tool configuration found in this repository.
Agent instructions 2
1 observed capture since 2026-09-07. Observed captures are shown by default.
Stars from first capture 0
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
[EMNLP'23, ACL'24] To speed up LLMs' inference and enhance LLM's perceive of key information, compress the prompt and KV-Cache, which achieves up to 20x compression with minimal performance loss.
Learn AI and LLMs from scratch using free resources
Foundation Architecture for (M)LLMs
Prompt Orchestration Markup Language
Benchmarking large language models' complex reasoning ability with chain-of-thought prompting
Awesome resources for in-context learning and prompt engineering: Mastery of the LLMs such as ChatGPT, GPT-3, and FlanT5, with up-to-date and cutting-edge updates.