Indic-Transcribe-core INT4 ONNX
Built with Indic-Transcribe-core from Bodhan AI / AI4Bharat.
bodhan-ai/indic-transcribe-core (speech recognition for 25 Indian languages and English) quantized to 4-bit and exported to ONNX for on-device use with ONNX Runtime. 0.70 GB. Unofficial community release, not affiliated with Bodhan AI, AI4Bharat or NVIDIA.
Install
pip install -r requirements.txt # CPU
pip install "onnxruntime-gpu[cuda,cudnn]" # optional: NVIDIA GPU
Usage
python inference.py clip.wav --lang hi
python inference.py clip.wav --lang ta --mode mixed
python inference.py clip.wav --lang hi --device cuda
from inference import IndicTranscribeONNX, load_audio
model = IndicTranscribeONNX() # IndicTranscribeONNX(device="cuda") for GPU
result = model.transcribe(load_audio("clip.wav"), lang="hi")
print(result["text"])
- Language must be given:
as bhb bho bn brx doi en gu hi kn kok ks mai ml mni mr ne or pa sa sat sd ta te ur. - Mode:
native(native script, default),mixed(Latin loanwords, digits),romanized. - Audio: 16 kHz mono, up to 30 s per call; other sample rates are resampled.
Accuracy
FLEURS dev (60 clips per language), word error rate:
| Model | Size | Hindi | Tamil |
|---|---|---|---|
| Original (fp32 checkpoint, run in bf16) | 4.9 GB | 10.2 | 25.3 |
| This model (INT4) | 0.70 GB | 11.0 | 26.9 |
Details
- 4-bit weight-only quantization (ONNX Runtime
MatMulNBits, block 128); activations stay in floating point. - Three graphs in
onnx/: encoder, cross-attention key/value projection (once per utterance), decoder with a KV cache. - Runs on ONNX Runtime CPU (including ONNX Runtime Mobile) and CUDA.
- Like the original, it produces text for non-speech audio: gate the input with a voice activity detector.
License
Licensed under the Indic Open Model License v1.0 (Bodhan_AI_Open_Model_License.md, summary in indic-open-license.md).
Hosting it as an API for third parties requires Bodhan AI's written approval.
Base architecture nvidia/canary-1b-v2 is licensed under CC-BY-4.0.
Attribution and changes: NOTICE.md.
Citation
@misc{indictranscribe2026,
title = {Indic-Transcribe: Built for the way India actually speaks},
author = {Bodhan AI, AI4Bharat},
year = {2026},
url = {https://bodhan.ai/research/blogs/indic-transcribe}
}