FranxYao/chain-of-thought-hub
Benchmarking large language models' complex reasoning ability with chain-of-thought prompting
Leaderboard Comparing LLM Performance at Producing Hallucinations when Summarizing Short Documents
Appears on
Quick read
Latest capture 2026-09-10 10:55
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
1 observed capture since 2026-09-10. Observed captures are shown by default.
Stars from first capture 0
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
Benchmarking large language models' complex reasoning ability with chain-of-thought prompting
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
No description.
GPT-Fathom is an open-source and reproducible LLM evaluation suite, benchmarking 10+ leading open-source and closed-source LLMs as well as OpenAI's earlier models on 20+ curated benchmarks under aligned settings.
DeepSeek LLM: Let there be answers
Rigourous evaluation of LLM-synthesized code - NeurIPS 2023 & COLM 2024