Open Datasets

Data built from
authentic community sources

All datasets collected with community consent, annotated locally, and released openly under CC-BY-4.0. Available on HuggingFace and AIKosh.

Language corpora

Northeast India Voices

A multilingual speech and text corpus collected across Northeast Indian communities, capturing natural spoken and written language as used in everyday context. Built to support ASR, TTS, and language modelling work across the region’s diverse linguistic landscape.

Languages Covered: Khasi, Garo, Mizo, Nagamese, and Kokborok.

Northeast Languages Test Set

A standardised evaluation benchmark for testing NE language model performance. Used internally across NE-BERT, NE-LID, and related models to ensure consistent, comparable evaluation across languages.

Languages covered: Assamese (asm), Garo (grt), Khasi (kha), Kokborok (trp), Meitei (mni), Mizo (lus), Nagamese (nag), Nyishi (njz), Pnar (pbv)

Pnar Speech Corpus

Native Pnar speech recordings collected directly from speaker communities in Meghalaya. One of the only structured speech datasets available for Pnar, supporting ASR and TTS development for a language with minimal prior digital representation.

Garo-English Parallel Corpus

Aligned sentence pairs in Garo and English, built for training and evaluating machine translation systems. A foundational resource for Garo NLP, where parallel data has historically been scarce.

Meitei Monolingual Corpus

A large-scale monolingual Meitei text corpus, used as the primary training data behind MeiteiRoBERTa. Covers a substantial volume of native Meitei text to support deeper language modelling research.

Assamese Monolingual Corpus

A large-scale monolingual Assamese text corpus assembled for language model pretraining. Provides broad coverage of written Assamese across formal and informal registers.

Mizo Language Corpus

A large open Mizo NLP dataset available, 4 million cleaned and processed sentences derived from a 5.94M sentence source corpus. Built to support NLP research, linguistic equity, and open development for Mizo.

Northeast India Tribes & Subtribes

The first structured open-source dataset documenting tribes and sub-tribes across all 8 states of Northeast India. Built to support cultural and linguistic reference for researchers, developers, and educators working on the region

Northeast India Districts & Villages

A statewise dataset of districts and villages across Northeast India, sourced and cleaned from the Government of India’s LGDirectory. Useful for geographic NLP, geocoding, and administrative reference tasks.

VISION-LANGUAGE DATASET

NE-CLIP Dataset

A multimodal dataset pairing images with parallel English and native-language captions across Northeast Indian languages, including Kokborok. Each row includes an image, English description, native-language description, and language label — built to train NE-CLIP for cross-modal alignment in low-resource languages.

Want to contribute data or partner on collection?

We work with NGOs, universities, and community organisations across Northeast India.