Skip to content

Latest commit

 

History

History
 
 

README.md

Voxtral Realtime

Self-contained ExecuTorch implementation of Mistral's Voxtral-Mini-4B-Realtime-2602, a ~4B parameter streaming speech-to-text model. No HuggingFace Transformers dependency — weights are loaded directly from the Mistral checkpoint. See model.md for architecture and implementation details.

Overview

The pipeline has two stages: export (Python, once) and inference (C++ runner, repeated). Export converts the Mistral checkpoint into a model.pte file. A separate preprocessor.pte handles audio-to-mel conversion. At inference time, the C++ runner loads both .pte files and the Tekken tokenizer, then transcribes audio to text.

Two modes are supported: streaming (process 80ms chunks in real time, including live microphone input, with unlimited duration) and offline (encode full audio, then decode, bounded by --max-seq-len). The examples below use streaming mode. Omit --streaming from export and run commands for offline mode.

Demo: streaming on Metal backend with microphone input

IMG_8714.mp4

Also, try a sample standalone macOS app to do real time transcription.

VoxtralApp.-.final.mp4

Prerequisites

  • ExecuTorch installed from source (see building from source)
  • safetensors (pip install safetensors)
  • Model weights downloaded from HuggingFace. The directory should contain params.json, consolidated.safetensors, and tekken.json.

Preprocessor

Export a preprocessor .pte to convert raw audio into the format the model expects:

python -m executorch.extension.audio.mel_spectrogram \
    --feature_size 128 \
    --streaming \
    --output_file ./voxtral_rt_exports/preprocessor.pte

For offline mode:

python -m executorch.extension.audio.mel_spectrogram \
    --feature_size 128 \
    --max_audio_len 300 \
    --output_file ./voxtral_rt_exports/preprocessor.pte

For MLX backend, use --backend mlx:

python -m executorch.extension.audio.mel_spectrogram \
    --feature_size 128 \
    --max_audio_len 300 \
    --backend mlx \
    --output_file ./voxtral_rt_exports/preprocessor.pte

For streaming, use a separate preprocessor with --streaming (no audio length limit):

python -m executorch.extension.audio.mel_spectrogram \
    --feature_size 128 \
    --streaming \
    --output_file ./voxtral_streaming_exports/preprocessor.pte

For streaming with MLX backend:

python -m executorch.extension.audio.mel_spectrogram \
    --feature_size 128 \
    --streaming \
    --backend mlx \
    --output_file ./voxtral_streaming_exports/preprocessor.pte

Export

Export produces a single .pte containing the audio encoder, text decoder, and token embedding.

Tip

Mistral has already published pre-exported .pte files for select backends, including macOS Metal, on their HuggingFace Hub.

XNNPACK (default)

python export_voxtral_rt.py \
    --model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
    --backend xnnpack \
    --streaming \
    --sliding-window 2048 \
    --output-dir ./voxtral_rt_exports \
    --qlinear-encoder 8da4w \
    --qlinear 8da4w \
    --qembedding 8w

Backend support

Backend Offline Streaming Quantization
xnnpack ✓ ✓ 4w, 8w, 8da4w, 8da8w
metal ✓ ✓ none (fp32) or fpa4w (Metal-specific 4-bit)
mlx ✓ ✓ 4w, 8w, nvfp4 (NVIDIA FP4 dtype)
cuda ✓ ✓ 4w, 8w
cuda-windows ✓ ✓ 4w, 8w
rocm ✓ ✓ BF16; packed linear 4w and embedding 8w

MLX and Metal backends provide Apple GPU acceleration. CUDA provides NVIDIA GPU acceleration, and experimental ROCm support provides AMD GPU acceleration, both through AOTInductor.

CUDA export examples

Offline with int4 quantization:

python export_voxtral_rt.py \
    --model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
    --backend cuda \
    --dtype bf16 \
    --output-dir ./voxtral_rt_exports \
    --qlinear-encoder 4w \
    --qlinear-encoder-packing-format tile_packed_to_4d \
    --qlinear 4w \
    --qlinear-packing-format tile_packed_to_4d \
    --qembedding 8w

Streaming with int4 quantization:

python export_voxtral_rt.py \
    --model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
    --backend cuda \
    --dtype bf16 \
    --streaming \
    --output-dir ./voxtral_rt_exports \
    --qlinear-encoder 4w \
    --qlinear-encoder-packing-format tile_packed_to_4d \
    --qlinear 4w \
    --qlinear-packing-format tile_packed_to_4d \
    --qembedding 8w

ROCm export examples

ROCm support is experimental. Manual validation currently covers MI300X (gfx942); CI canaries exercise gfx950 and gfx1100. ROCm is never enabled automatically. Use a ROCm PyTorch build with its matching Triton AMD backend; do not run install_executorch.sh, because its dependency setup can replace ROCm PyTorch with a CPU build.

The validated ROCm configurations use BF16 and optionally packed 4w linear weights with an 8w embedding. The exporter rejects other ROCm dtype and quantization combinations before loading the model.

Start with the BF16 streaming baseline:

python export_voxtral_rt.py \
    --model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
    --backend rocm \
    --dtype bf16 \
    --streaming \
    --sliding-window 2048 \
    --output-dir ./voxtral_rt_rocm_bf16

The preferred W4/BF16 setup nibble-packs TorchAO weight-only INT4 tensors and runs the ExecuTorch Triton W4A16 kernel while keeping activations in BF16:

python export_voxtral_rt.py \
    --model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
    --backend rocm \
    --dtype bf16 \
    --streaming \
    --sliding-window 2048 \
    --output-dir ./voxtral_rt_rocm_w4_bf16 \
    --qlinear-encoder 4w \
    --qlinear 4w \
    --qembedding 8w

Do not use CUDA's tile_packed_to_4d option on ROCm. That format requires the CUDA-only _weight_int4pack_mm fallback shim, which is intentionally not built or advertised by the ROCm backend. The exporter rejects that combination before model loading.

The packed path performs dequantization inside the GPU kernel and does not materialize a full BF16 weight for each invocation.

The ROCm W4 decoder is specialized to the runner's one-token input and uses the packed INT4 matvec kernel. Encoder linears use packed INT4 matmul. ROCm uses the standard SDPA kernel because split-K decode produced non-finite logits for this fixed-shape workload. CUDA and other non-ROCm exports are unchanged.

Metal export examples

Offline:

python export_voxtral_rt.py \
    --model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
    --backend metal \
    --output-dir ./voxtral_rt_exports \
    --qlinear-encoder fpa4w \
    --qlinear fpa4w

Streaming:

python export_voxtral_rt.py \
    --model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
    --backend metal \
    --dtype bf16 \
    --streaming \
    --sliding-window 2048 \
    --output-dir ./voxtral_rt_exports \
    --qlinear-encoder fpa4w \
    --qlinear fpa4w

Metal 4-bit quantization (fpa4w) requires torchao built with experimental MPS ops:

# From the ao repo (third-party/ao/)
USE_CPP=1 TORCHAO_BUILD_EXPERIMENTAL_MPS=1 pip install . --no-build-isolation

# Or while installing ExecuTorch from source
EXECUTORCH_BUILD_KERNELS_TORCHAO=1 TORCHAO_BUILD_EXPERIMENTAL_MPS=1 ./install_executorch.sh

MLX export examples

MLX backend uses the MLX delegate for Apple Silicon GPU acceleration. NVFP4 quantizes weights using NVIDIA's FP4 data type.

Offline (NVFP4):

python export_voxtral_rt.py \
    --model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
    --backend mlx \
    --output-dir ./voxtral_rt_exports \
    --qlinear-encoder nvfp4 \
    --qlinear nvfp4 \
    --qembedding nvfp4

Streaming (NVFP4):

python export_voxtral_rt.py \
    --model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
    --backend mlx \
    --streaming \
    --output-dir ./voxtral_rt_exports \
    --qlinear-encoder nvfp4 \
    --qlinear nvfp4 \
    --qembedding nvfp4

Offline (int4 linear + int8 embedding):

python export_voxtral_rt.py \
    --model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
    --backend mlx \
    --output-dir ./voxtral_rt_exports \
    --qlinear-encoder 4w \
    --qlinear 4w \
    --qembedding 8w \
    --qembedding-group-size 128

Streaming (int4 linear + int8 embedding):

python export_voxtral_rt.py \
    --model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
    --backend mlx \
    --streaming \
    --sliding-window 2048 \
    --output-dir ./voxtral_rt_exports \
    --qlinear-encoder 4w \
    --qlinear 4w \
    --qembedding 8w \
    --qembedding-group-size 128

CUDA-Windows export examples

Requires x86_64-w64-mingw32-g++ on PATH (mingw-w64 cross-compiler) and WINDOWS_CUDA_HOME pointing to the extracted Windows CUDA package directory. See Parakeet README for detailed extraction steps.

export WINDOWS_CUDA_HOME=/opt/cuda-windows/extracted/cuda_cudart/cudart

python export_voxtral_rt.py \
    --model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
    --backend cuda-windows \
    --dtype bf16 \
    --streaming \
    --sliding-window 2048 \
    --output-dir ./voxtral_rt_exports \
    --qlinear-encoder 4w \
    --qlinear-encoder-packing-format tile_packed_to_4d \
    --qlinear 4w \
    --qlinear-packing-format tile_packed_to_4d \
    --qembedding 8w

Note

Omit --streaming from any export command above for offline mode. CUDA, CUDA-Windows, and ROCm exports also produce an aoti_cuda_blob.ptd file alongside model.pte.

Options

Flag Default Description
--model-path (required) Directory with params.json + consolidated.safetensors
--backend xnnpack xnnpack, mlx, metal, cuda, cuda-windows, rocm, or portable
--dtype fp32 Model dtype: fp32 or bf16
--output-dir ./voxtral_rt_exports Output directory
--max-seq-len 4096 KV cache length (offline mode only; ignored with --streaming)
--delay-tokens 6 Transcription delay in tokens (6 = 480ms)

| --qlinear | (none) | Decoder linear layer quantization (4w, 8w, 8da4w, 8da8w, fpa4w, nvfp4) | | --qlinear-group-size | auto | Group size for decoder linear quantization | | --qlinear-packing-format | (none) | Packing format for decoder 4w quantization (tile_packed_to_4d for CUDA) | | --qlinear-encoder | (none) | Encoder linear layer quantization (4w, 8w, 8da4w, 8da8w, fpa4w, nvfp4) | | --qlinear-encoder-group-size | auto | Group size for encoder linear quantization | | --qlinear-encoder-packing-format | (none) | Packing format for encoder 4w quantization (tile_packed_to_4d for CUDA) | | --qembedding | (none) | Embedding layer quantization (4w, 8w, nvfp4) | | --qembedding-group-size | auto | Group size for embedding quantization | | --streaming | off | Export streaming model with ring buffer KV caches (unlimited duration) | | --max-enc-len | 750 | Encoder sliding window size (streaming only) | | --sliding-window | from params.json | Decoder sliding window size (streaming only; ignored in offline mode). Smaller values reduce memory and improve decode speed but limit context | Notes:

  • fpa4w quantization requires --backend metal.
  • The model was trained with --delay-tokens 6. Other values may degrade accuracy.
  • The decoder sliding window controls how far back the decoder can attend. At 80ms/step: 2048 = ~2.7 min, 4096 = ~5.5 min, 8192 = ~10.9 min.

Build

ExecuTorch must be installed from source first (see Prerequisites). The make targets below handle building core libraries and the runner binary.

make voxtral_realtime-cpu      # XNNPACK (CPU)
make voxtral_realtime-metal    # Metal (Apple GPU)
make voxtral_realtime-cuda     # CUDA (NVIDIA GPU)
make voxtral_realtime-rocm     # ROCm (AMD GPU, experimental)

The CPU, CUDA, Metal, and MLX targets produce the runner at cmake-out/examples/models/voxtral_realtime/voxtral_realtime_runner. ROCm uses the isolated path below.

ROCm (experimental)

From the ExecuTorch repository root, the explicit ROCm target builds and installs the LLM runtime and the model runner into a separate directory:

make voxtral_realtime-rocm

The equivalent workflows are:

cmake --workflow --preset llm-release-rocm
cd examples/models/voxtral_realtime
cmake --workflow --preset voxtral-realtime-rocm

The runner is written to cmake-out-rocm-llm/examples/models/voxtral_realtime/voxtral_realtime_runner. The separate directory keeps ROCm configuration out of default and CUDA build caches.

For an end-to-end BF16 or W4/BF16 example:

examples/models/voxtral_realtime/run_rocm_e2e.sh \
    ~/models/Voxtral-Mini-4B-Realtime-2602 \
    /path/to/input-16khz-mono.wav \
    w4-bf16 \
    both \
    /path/to/output

The third argument selects bf16, w4-bf16, or both precision modes. The fourth selects streaming, offline, or both execution modes. Set ROCM_PATH if ROCm is installed outside /opt/rocm. The script reports model export time, PTE/PTD sizes, and RTF computed as runner inference time divided by WAV duration.

The AOTInductor-generated shared objects use the C++ runtime from the active Python environment. Add that runtime before launching the runner directly:

PYTHON_PREFIX="$(python -c 'import sys; print(sys.prefix)')"
export LD_LIBRARY_PATH="$PYTHON_PREFIX/lib:${LD_LIBRARY_PATH:-}"

Without it, systems with an older /lib64/libstdc++.so.6 can fail while loading the extracted delegate library with a missing GLIBCXX version.

CUDA-Windows

On Windows (PowerShell), use CMake workflow presets from the executorch root directory. If you exported with 4-bit quantization, specify your GPU's compute capability to avoid "invalid device function" errors (the int4mm kernels require SM 80+).

cmake --workflow --preset llm-release-cuda
Push-Location examples/models/voxtral_realtime
cmake --workflow --preset voxtral-realtime-cuda
Pop-Location

This builds ExecuTorch with CUDA backend support. The runner binary is at the same path as above. Requires NVIDIA GPU with CUDA toolkit installed.

Metal (Apple GPU)

make voxtral_realtime-metal

This builds ExecuTorch with Metal backend support. The runner binary is at the same path as above. Metal exports can only run on macOS with Apple Silicon.

MLX (Apple GPU)

make voxtral_realtime-mlx

This builds ExecuTorch with MLX backend support. MLX provides GPU acceleration on Apple Silicon via the MLX delegate.

Run

The runner requires:

  • model.pte — exported model (see Export)
  • tekken.json — tokenizer from the model weights directory
  • preprocessor.pte — mel spectrogram preprocessor (see Preprocessor)
  • A 16kHz mono WAV audio file (or live audio via --mic)

Basic usage

cmake-out/examples/models/voxtral_realtime/voxtral_realtime_runner \
    --model_path voxtral_rt_exports/model.pte \
    --tokenizer_path ~/models/Voxtral-Mini-4B-Realtime-2602/tekken.json \
    --preprocessor_path voxtral_rt_exports/preprocessor.pte \
    --audio_path input.wav \
    --streaming

Omit --streaming for offline transcription (requires an offline-exported model and offline preprocessor).

For CUDA and ROCm, add --data_path voxtral_rt_exports/aoti_cuda_blob.ptd.

Windows (PowerShell):

.\cmake-out\examples\models\voxtral_realtime\Release\voxtral_realtime_runner.exe `
    --model_path voxtral_rt_exports\model.pte `
    --data_path voxtral_rt_exports\aoti_cuda_blob.ptd `
    --tokenizer_path C:\path\to\tekken.json `
    --preprocessor_path voxtral_rt_exports\preprocessor.pte `
    --audio_path input.wav `
    --streaming

Live microphone input

Use --mic to read raw 16kHz float32 PCM from stdin. Requires a streaming-exported model and streaming preprocessor. Pipe from any audio capture tool:

# macOS
ffmpeg -f avfoundation -i ":0" -ar 16000 -ac 1 -f f32le -nostats -loglevel error pipe:1 | \
  cmake-out/examples/models/voxtral_realtime/voxtral_realtime_runner \
    --model_path voxtral_rt_exports/model.pte \
    --tokenizer_path ~/models/Voxtral-Mini-4B-Realtime-2602/tekken.json \
    --preprocessor_path voxtral_rt_exports/preprocessor.pte \
    --mic

Ctrl+C stops recording and flushes remaining text.

Options

Flag Default Description
--model_path model.pte Path to exported model
--data_path (none) Path to delegate data file (.ptd, required for CUDA and ROCm)
--tokenizer_path tekken.json Path to Tekken tokenizer
--preprocessor_path (none) Path to mel preprocessor .pte
--audio_path (none) Path to 16kHz mono WAV file
--temperature 0.0 Sampling temperature (0 = greedy)
--offline_max_new_tokens 500 Offline-only: maximum extra tokens after audio embeddings are exhausted
--streaming off Use streaming transcription (from WAV file)
--mic off Live microphone mode (reads raw f32le PCM from stdin)
--mic_chunk_ms 80 Mic read chunk size in ms (multiples of 80 recommended)
--color (none) Output text color: green or red

Troubleshooting

  • Audio format: Input must be 16kHz mono WAV. Convert with ffmpeg -i input.mp3 -ar 16000 -ac 1 output.wav.
  • OOM during export: Reduce --max-seq-len (offline mode) or skip encoder quantization (--qlinear-encoder).
  • "Model was not exported with --streaming": Re-export with the --streaming flag. Both --streaming and --mic runner modes require a streaming-exported model.
  • fpa4w error: This quantization requires --backend metal.
  • Metal runner fails with Library not loaded: @rpath/libc++.1.dylib: The AOTInductor-compiled .so inside the .pte references libc++ via @rpath, which can't be resolved when extracted to a temp directory. Add /usr/lib to DYLD_LIBRARY_PATH so dyld finds it in the shared cache:
    DYLD_LIBRARY_PATH=/usr/lib \
        cmake-out/examples/models/voxtral_realtime/voxtral_realtime_runner ...
  • Metal runner fails with Library not loaded: libomp.dylib: The AOTInductor-compiled .so links against OpenMP. Install it via Homebrew and add it to DYLD_LIBRARY_PATH:
    brew install libomp
    DYLD_LIBRARY_PATH=/usr/lib:$(brew --prefix libomp)/lib \
        cmake-out/examples/models/voxtral_realtime/voxtral_realtime_runner ...