Meta omniASR-CTC-300M Standalone Audio Encoder (bfloat16 SafeTensors)
This repository provides the standalone bfloat16 safetensors Audio Encoder extracted from facebook/omniASR-CTC-300M.
It is $100%$ natively compatible with Hugging Face transformers.Wav2Vec2Model and transformers.AutoModel without requiring Fairseq2 or complex custom dependencies.
π Architecture & Technical Specifications
| Parameter | Value | Detailed Description |
|---|---|---|
| Audio Tower Architecture | 24 Transformer/Conformer Layers (16 heads, FFN 4096) | 300M Parameter Omnilingual Audio Encoder |
| Feature Dimension (Hidden Size) | 1024-dim | Continuous dense acoustic representations (last_hidden_state) |
| Frame Rate | 50 frames/s (20ms / frame) | $320\times$ raw audio temporal downsampling ($16000 / 320 = 50\text{Hz}$) |
| Input Audio Front-End | 16 kHz Raw Waveforms | 7-layer 1D Temporal Convolution ($5\times 2\times 2\times 2\times 2\times 2\times 2 = 320\times$) |
| Language Coverage | 1,600+ Languages | Pre-trained across global language families |
| Weight Format & Precision | model.safetensors (bfloat16) |
Lightweight standalone package: ~631 MB |
π» Quick Start with Hugging Face transformers
import torch
import torchaudio
from transformers import AutoFeatureExtractor, AutoModel
# 1. Load Standalone Encoder from Hugging Face
model_id = "giangndm/omniASR-CTC-300M-encoder"
feature_extractor = AutoFeatureExtractor.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id, torch_dtype=torch.bfloat16).to("cuda")
model.eval()
# 2. Load Any 16kHz Audio File
waveform, sr = torchaudio.load("path/to/audio.wav") # [1, T_samples]
if sr != 16000:
resampler = torchaudio.transforms.Resample(sr, 16000)
waveform = resampler(waveform)
inputs = feature_extractor(waveform.squeeze(0).numpy(), sampling_rate=16000, return_tensors="pt")
input_values = inputs.input_values.to("cuda").to(torch.bfloat16)
# 3. Extract 50Hz Dense Audio Representations
with torch.no_grad():
outputs = model(input_values)
audio_embeds = outputs.last_hidden_state # [1, T_frames, 1024] at 50 Hz
print("Audio Embeddings Shape:", audio_embeds.shape)
# Example: 1-second audio -> torch.Size([1, 49, 1024])
π¦ Key Components Inside:
feature_extractor: 7 temporal convolutional layers with layer normalization downsampling raw audio $320\times$.feature_projection: Linear layer projecting 512-dim conv features to 1024-dim model space.encoder.pos_conv_embed: Depthwise convolutional positional embedding.encoder.layers.0...encoder.layers.23: 24 Transformer self-attention layers with Pre-LayerNorm and GELU FFNs.encoder.layer_norm: Final LayerNorm of the Audio Tower.
- Downloads last month
- 12