8 data sources · updated Aug 7, 14:56 UTC
AI Readiness
Models, datasets, and evaluation benchmarks that show whether a language is trainable and evaluable for AI.
Start from a language to see how every data source in this category covers it, side by side.
349 languages have data in at least one data source.
Hugging Face Hub (models)
95 langLanguage-tagged models on the Hub. Counts need careful filtering and deduplication.
- Languages in this snapshot
- 95
- Filter parameter
- ?filter=
Hugging Face Hub (datasets)
95 langLanguage-tagged datasets on the Hub. Strong signal that a language is trainable.
- Languages in this snapshot
- 95
- Filter parameter
- ?filter=
NLLB-200
196 langMeta's No Language Left Behind translation models covering ~200 languages (FLORES-aligned).
- Languages in NLLB-200
- 196
- Model family
- facebook/nllb-200
FLORES-200
194 langEvaluation benchmark for ~200 languages — the de facto baseline for whether a language is evaluable at all.
- Languages in FLORES-200
- 194
- Role
- MT evaluation benchmark
SIB-200
197 langTopic classification benchmark across ~200 languages; among the widest eval coverage available.
- Languages in SIB-200
- 197
- Task
- Topic classification
Belebele
96 langReading comprehension across 100+ language variants. Presence means evaluable, not merely trainable.
- Language variants in Belebele
- 96
- Task
- Reading comprehension
Global-MMLU
42 langKnowledge evaluation (MMLU-style) in 42 languages from Cohere For AI.
- Languages in Global-MMLU
- 42
- Task
- Knowledge / MMLU-style evaluation
MMS
129 langMassively Multilingual Speech — 1000+ ASR and 1100+ TTS languages (Meta). Strong speech signal.
- Languages in MMS (this harvest)
- 129
- ASR languages (reported)
- 1000+