The NE-Stack

Models built for
Northeast India.

From speech to text, translation to vision; every model is trained from authentic community data across Northeast India’s indigenous languages. Open-source, production-ready, and deployable offline.

Flagship models

NE-ASR

Speech Recognition

Multilingual ASR system covering 8 Northeast Indian languages. Fine-tuned on 267+ hours of native speech data collected across Meghalaya, Mizoram, Manipur, Tripura, Arunachal Pradesh, Nagaland and Assam.

NE-OCR

Text Recognition

Unified OCR model for 9 Northeast Indian language-script pairs. Trained on 1.34M text-image pairs from native corpora. Fastest inference at 17.2ms per image.

NE-BERT

Language Models

Regional state-of-the-art open-source encoder for 9 Northeast Indian languages. Built on ModernBERT for superior speed and accuracy in low-resource NLP tasks.

Kren-M

Generative Language Model

Kren is MWire Labs’ flagship generative LM — a bilingual Khasi-English model developed through extensive continued pre-training on Gemma 2 (2B). The talk and reasoning layer of the NE-Stack.

Open Model Family

NLP

BERT Suite + Language ID

Individual encoder models for 8 NE languages. CC-BY-4.0, on HuggingFace.

SPEECH

Speech Models

Multilingual speech understanding and identification across NE languages.

Translation

Machine Translation

11 NE language pairs. NLLB-200 + M2M-100 and IndicTransV2 fine-tuning

Vision Langauge

MultiModels

Cross-modal and vision-language alignment across 9 NE languages.

Embeddings

Text Embeddings

LaBSE fine-tuned on NE languages. Published to AIKosh. pip-installable.

Want to deploy these
in your application?

API access, on-premise licensing, and custom integration support available.