Meta omniASR-CTC-300M Standalone Audio Encoder (bfloat16 SafeTensors)

This repository provides the standalone bfloat16 safetensors Audio Encoder extracted from facebook/omniASR-CTC-300M.

It is $100%$ natively compatible with Hugging Face transformers.Wav2Vec2Model and transformers.AutoModel without requiring Fairseq2 or complex custom dependencies.


πŸ“ Architecture & Technical Specifications

Parameter Value Detailed Description
Audio Tower Architecture 24 Transformer/Conformer Layers (16 heads, FFN 4096) 300M Parameter Omnilingual Audio Encoder
Feature Dimension (Hidden Size) 1024-dim Continuous dense acoustic representations (last_hidden_state)
Frame Rate 50 frames/s (20ms / frame) $320\times$ raw audio temporal downsampling ($16000 / 320 = 50\text{Hz}$)
Input Audio Front-End 16 kHz Raw Waveforms 7-layer 1D Temporal Convolution ($5\times 2\times 2\times 2\times 2\times 2\times 2 = 320\times$)
Language Coverage 1,600+ Languages Pre-trained across global language families
Weight Format & Precision model.safetensors (bfloat16) Lightweight standalone package: ~631 MB

πŸ’» Quick Start with Hugging Face transformers

import torch
import torchaudio
from transformers import AutoFeatureExtractor, AutoModel

# 1. Load Standalone Encoder from Hugging Face
model_id = "giangndm/omniASR-CTC-300M-encoder"
feature_extractor = AutoFeatureExtractor.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id, torch_dtype=torch.bfloat16).to("cuda")
model.eval()

# 2. Load Any 16kHz Audio File
waveform, sr = torchaudio.load("path/to/audio.wav") # [1, T_samples]
if sr != 16000:
    resampler = torchaudio.transforms.Resample(sr, 16000)
    waveform = resampler(waveform)

inputs = feature_extractor(waveform.squeeze(0).numpy(), sampling_rate=16000, return_tensors="pt")
input_values = inputs.input_values.to("cuda").to(torch.bfloat16)

# 3. Extract 50Hz Dense Audio Representations
with torch.no_grad():
    outputs = model(input_values)
    audio_embeds = outputs.last_hidden_state # [1, T_frames, 1024] at 50 Hz

print("Audio Embeddings Shape:", audio_embeds.shape)
# Example: 1-second audio -> torch.Size([1, 49, 1024])

πŸ“¦ Key Components Inside:

  • feature_extractor: 7 temporal convolutional layers with layer normalization downsampling raw audio $320\times$.
  • feature_projection: Linear layer projecting 512-dim conv features to 1024-dim model space.
  • encoder.pos_conv_embed: Depthwise convolutional positional embedding.
  • encoder.layers.0 ... encoder.layers.23: 24 Transformer self-attention layers with Pre-LayerNorm and GELU FFNs.
  • encoder.layer_norm: Final LayerNorm of the Audio Tower.
Downloads last month
12
Safetensors
Model size
0.3B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support