jaberjaber23/rightnow-arabic-llm-corpus
RightNow Arabic LLM Corpus - One of the largest high-quality Arabic text datasets for LLM training
This repo contains a set of Arabic newspaper articles alongwith metadata, extracted from various Saudi newspapers.
Appears on
Quick read
Latest capture 2026-09-04 03:04
0 paths
Agent instructions and tool configuration found in this repository.
No config files detected.
7 observed captures since 2026-05-23. Observed captures are shown by default.
Stars from first capture 0
Observed captures only
All tracked data
Observed snapshots
Observed snapshots
Nearest indexed repositories by embedding similarity.
RightNow Arabic LLM Corpus - One of the largest high-quality Arabic text datasets for LLM training
Saudi Real Estate Open Data: 4.93M transactions from MOJ & REGA (2018-2025). Sales, mortgages, seizures, rentals across 13 regions.
source{d} datasets ("big code") for source code analysis and machine learning on source code
MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集。对标chatGPT训练的40T数据。MNBVC数据集不但包括主流文化,也包括各个小众文化甚至火星文的数据。MNBVC数据集包括新闻、作文、小说、书籍、杂志、论文、台词、帖子、wiki、古诗、歌词、商品介绍、笑话、糗事、聊天记录等一切形式的纯文本中文数据。
P2PCLAW: Training Dataset for Autonomous Scientific Peer Review - Apache 2.0
Disability demographics, web accessibility compliance, assistive technology usage, and related data