Indic-Transcribe-core INT4 ONNX

Built with Indic-Transcribe-core from Bodhan AI / AI4Bharat.

bodhan-ai/indic-transcribe-core (speech recognition for 25 Indian languages and English) quantized to 4-bit and exported to ONNX for on-device use with ONNX Runtime. 0.70 GB. Unofficial community release, not affiliated with Bodhan AI, AI4Bharat or NVIDIA.

Install

pip install -r requirements.txt                 # CPU
pip install "onnxruntime-gpu[cuda,cudnn]"       # optional: NVIDIA GPU

Usage

python inference.py clip.wav --lang hi
python inference.py clip.wav --lang ta --mode mixed
python inference.py clip.wav --lang hi --device cuda
from inference import IndicTranscribeONNX, load_audio

model = IndicTranscribeONNX()                    # IndicTranscribeONNX(device="cuda") for GPU
result = model.transcribe(load_audio("clip.wav"), lang="hi")
print(result["text"])
  • Language must be given: as bhb bho bn brx doi en gu hi kn kok ks mai ml mni mr ne or pa sa sat sd ta te ur.
  • Mode: native (native script, default), mixed (Latin loanwords, digits), romanized.
  • Audio: 16 kHz mono, up to 30 s per call; other sample rates are resampled.

Accuracy

FLEURS dev (60 clips per language), word error rate:

Model Size Hindi Tamil
Original (fp32 checkpoint, run in bf16) 4.9 GB 10.2 25.3
This model (INT4) 0.70 GB 11.0 26.9

Details

  • 4-bit weight-only quantization (ONNX Runtime MatMulNBits, block 128); activations stay in floating point.
  • Three graphs in onnx/: encoder, cross-attention key/value projection (once per utterance), decoder with a KV cache.
  • Runs on ONNX Runtime CPU (including ONNX Runtime Mobile) and CUDA.
  • Like the original, it produces text for non-speech audio: gate the input with a voice activity detector.

License

Licensed under the Indic Open Model License v1.0 (Bodhan_AI_Open_Model_License.md, summary in indic-open-license.md). Hosting it as an API for third parties requires Bodhan AI's written approval. Base architecture nvidia/canary-1b-v2 is licensed under CC-BY-4.0. Attribution and changes: NOTICE.md.

Citation

@misc{indictranscribe2026,
  title  = {Indic-Transcribe: Built for the way India actually speaks},
  author = {Bodhan AI, AI4Bharat},
  year   = {2026},
  url    = {https://bodhan.ai/research/blogs/indic-transcribe}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kazuto098/indic-transcribe-core-int4-onnx

Quantized
(7)
this model