Self-contained ExecuTorch implementation of Mistral's Voxtral-Mini-4B-Realtime-2602, a ~4B parameter streaming speech-to-text model. No HuggingFace Transformers dependency — weights are loaded directly from the Mistral checkpoint. See model.md for architecture and implementation details.
The pipeline has two stages: export (Python, once) and inference
(C++ runner, repeated). Export converts the Mistral checkpoint into a
model.pte file. A separate preprocessor.pte handles audio-to-mel
conversion. At inference time, the C++ runner loads both .pte files
and the Tekken tokenizer, then transcribes audio to text.
Two modes are supported: streaming (process 80ms chunks in real time,
including live microphone input, with unlimited duration) and offline
(encode full audio, then decode, bounded by --max-seq-len). The examples
below use streaming mode. Omit --streaming from export and run commands
for offline mode.
IMG_8714.mp4
Also, try a sample standalone macOS app to do real time transcription.
VoxtralApp.-.final.mp4
- ExecuTorch installed from source (see building from source)
- safetensors (
pip install safetensors) - Model weights downloaded from HuggingFace.
The directory should contain
params.json,consolidated.safetensors, andtekken.json.
Export a preprocessor .pte to convert raw audio into the format the
model expects:
python -m executorch.extension.audio.mel_spectrogram \
--feature_size 128 \
--streaming \
--output_file ./voxtral_rt_exports/preprocessor.pteFor offline mode:
python -m executorch.extension.audio.mel_spectrogram \
--feature_size 128 \
--max_audio_len 300 \
--output_file ./voxtral_rt_exports/preprocessor.pteFor MLX backend, use --backend mlx:
python -m executorch.extension.audio.mel_spectrogram \
--feature_size 128 \
--max_audio_len 300 \
--backend mlx \
--output_file ./voxtral_rt_exports/preprocessor.pteFor streaming, use a separate preprocessor with --streaming (no audio
length limit):
python -m executorch.extension.audio.mel_spectrogram \
--feature_size 128 \
--streaming \
--output_file ./voxtral_streaming_exports/preprocessor.pteFor streaming with MLX backend:
python -m executorch.extension.audio.mel_spectrogram \
--feature_size 128 \
--streaming \
--backend mlx \
--output_file ./voxtral_streaming_exports/preprocessor.pteExport produces a single .pte containing the audio encoder, text decoder,
and token embedding.
Tip
Mistral has already published pre-exported .pte files for select backends, including macOS Metal, on their HuggingFace Hub.
python export_voxtral_rt.py \
--model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
--backend xnnpack \
--streaming \
--sliding-window 2048 \
--output-dir ./voxtral_rt_exports \
--qlinear-encoder 8da4w \
--qlinear 8da4w \
--qembedding 8w| Backend | Offline | Streaming | Quantization |
|---|---|---|---|
xnnpack |
✓ | ✓ | 4w, 8w, 8da4w, 8da8w |
metal |
✓ | ✓ | none (fp32) or fpa4w (Metal-specific 4-bit) |
mlx |
✓ | ✓ | 4w, 8w, nvfp4 (NVIDIA FP4 dtype) |
cuda |
✓ | ✓ | 4w, 8w |
cuda-windows |
✓ | ✓ | 4w, 8w |
rocm |
✓ | ✓ | BF16; packed linear 4w and embedding 8w |
MLX and Metal backends provide Apple GPU acceleration. CUDA provides NVIDIA GPU acceleration, and experimental ROCm support provides AMD GPU acceleration, both through AOTInductor.
Offline with int4 quantization:
python export_voxtral_rt.py \
--model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
--backend cuda \
--dtype bf16 \
--output-dir ./voxtral_rt_exports \
--qlinear-encoder 4w \
--qlinear-encoder-packing-format tile_packed_to_4d \
--qlinear 4w \
--qlinear-packing-format tile_packed_to_4d \
--qembedding 8wStreaming with int4 quantization:
python export_voxtral_rt.py \
--model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
--backend cuda \
--dtype bf16 \
--streaming \
--output-dir ./voxtral_rt_exports \
--qlinear-encoder 4w \
--qlinear-encoder-packing-format tile_packed_to_4d \
--qlinear 4w \
--qlinear-packing-format tile_packed_to_4d \
--qembedding 8wROCm support is experimental. Manual validation currently covers MI300X
(gfx942); CI canaries exercise gfx950 and gfx1100. ROCm is never enabled
automatically. Use a ROCm PyTorch build with its matching Triton AMD backend;
do not run install_executorch.sh, because its dependency setup can replace
ROCm PyTorch with a CPU build.
The validated ROCm configurations use BF16 and optionally packed 4w linear
weights with an 8w embedding. The exporter rejects other ROCm dtype and
quantization combinations before loading the model.
Start with the BF16 streaming baseline:
python export_voxtral_rt.py \
--model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
--backend rocm \
--dtype bf16 \
--streaming \
--sliding-window 2048 \
--output-dir ./voxtral_rt_rocm_bf16The preferred W4/BF16 setup nibble-packs TorchAO weight-only INT4 tensors and runs the ExecuTorch Triton W4A16 kernel while keeping activations in BF16:
python export_voxtral_rt.py \
--model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
--backend rocm \
--dtype bf16 \
--streaming \
--sliding-window 2048 \
--output-dir ./voxtral_rt_rocm_w4_bf16 \
--qlinear-encoder 4w \
--qlinear 4w \
--qembedding 8wDo not use CUDA's tile_packed_to_4d option on ROCm. That format requires the
CUDA-only _weight_int4pack_mm fallback shim, which is intentionally not built
or advertised by the ROCm backend. The exporter rejects that combination
before model loading.
The packed path performs dequantization inside the GPU kernel and does not materialize a full BF16 weight for each invocation.
The ROCm W4 decoder is specialized to the runner's one-token input and uses the packed INT4 matvec kernel. Encoder linears use packed INT4 matmul. ROCm uses the standard SDPA kernel because split-K decode produced non-finite logits for this fixed-shape workload. CUDA and other non-ROCm exports are unchanged.
Offline:
python export_voxtral_rt.py \
--model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
--backend metal \
--output-dir ./voxtral_rt_exports \
--qlinear-encoder fpa4w \
--qlinear fpa4wStreaming:
python export_voxtral_rt.py \
--model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
--backend metal \
--dtype bf16 \
--streaming \
--sliding-window 2048 \
--output-dir ./voxtral_rt_exports \
--qlinear-encoder fpa4w \
--qlinear fpa4wMetal 4-bit quantization (fpa4w) requires torchao built with experimental MPS ops:
# From the ao repo (third-party/ao/)
USE_CPP=1 TORCHAO_BUILD_EXPERIMENTAL_MPS=1 pip install . --no-build-isolation
# Or while installing ExecuTorch from source
EXECUTORCH_BUILD_KERNELS_TORCHAO=1 TORCHAO_BUILD_EXPERIMENTAL_MPS=1 ./install_executorch.shMLX backend uses the MLX delegate for Apple Silicon GPU acceleration. NVFP4 quantizes weights using NVIDIA's FP4 data type.
Offline (NVFP4):
python export_voxtral_rt.py \
--model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
--backend mlx \
--output-dir ./voxtral_rt_exports \
--qlinear-encoder nvfp4 \
--qlinear nvfp4 \
--qembedding nvfp4Streaming (NVFP4):
python export_voxtral_rt.py \
--model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
--backend mlx \
--streaming \
--output-dir ./voxtral_rt_exports \
--qlinear-encoder nvfp4 \
--qlinear nvfp4 \
--qembedding nvfp4Offline (int4 linear + int8 embedding):
python export_voxtral_rt.py \
--model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
--backend mlx \
--output-dir ./voxtral_rt_exports \
--qlinear-encoder 4w \
--qlinear 4w \
--qembedding 8w \
--qembedding-group-size 128Streaming (int4 linear + int8 embedding):
python export_voxtral_rt.py \
--model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
--backend mlx \
--streaming \
--sliding-window 2048 \
--output-dir ./voxtral_rt_exports \
--qlinear-encoder 4w \
--qlinear 4w \
--qembedding 8w \
--qembedding-group-size 128Requires x86_64-w64-mingw32-g++ on PATH (mingw-w64 cross-compiler) and
WINDOWS_CUDA_HOME pointing to the extracted Windows CUDA package directory.
See Parakeet README for detailed extraction steps.
export WINDOWS_CUDA_HOME=/opt/cuda-windows/extracted/cuda_cudart/cudart
python export_voxtral_rt.py \
--model-path ~/models/Voxtral-Mini-4B-Realtime-2602 \
--backend cuda-windows \
--dtype bf16 \
--streaming \
--sliding-window 2048 \
--output-dir ./voxtral_rt_exports \
--qlinear-encoder 4w \
--qlinear-encoder-packing-format tile_packed_to_4d \
--qlinear 4w \
--qlinear-packing-format tile_packed_to_4d \
--qembedding 8wNote
Omit --streaming from any export command above for offline mode.
CUDA, CUDA-Windows, and ROCm exports also produce an aoti_cuda_blob.ptd file alongside model.pte.
| Flag | Default | Description |
|---|---|---|
--model-path |
(required) | Directory with params.json + consolidated.safetensors |
--backend |
xnnpack |
xnnpack, mlx, metal, cuda, cuda-windows, rocm, or portable |
--dtype |
fp32 |
Model dtype: fp32 or bf16 |
--output-dir |
./voxtral_rt_exports |
Output directory |
--max-seq-len |
4096 |
KV cache length (offline mode only; ignored with --streaming) |
--delay-tokens |
6 |
Transcription delay in tokens (6 = 480ms) |
| --qlinear | (none) | Decoder linear layer quantization (4w, 8w, 8da4w, 8da8w, fpa4w, nvfp4) |
| --qlinear-group-size | auto | Group size for decoder linear quantization |
| --qlinear-packing-format | (none) | Packing format for decoder 4w quantization (tile_packed_to_4d for CUDA) |
| --qlinear-encoder | (none) | Encoder linear layer quantization (4w, 8w, 8da4w, 8da8w, fpa4w, nvfp4) |
| --qlinear-encoder-group-size | auto | Group size for encoder linear quantization |
| --qlinear-encoder-packing-format | (none) | Packing format for encoder 4w quantization (tile_packed_to_4d for CUDA) |
| --qembedding | (none) | Embedding layer quantization (4w, 8w, nvfp4) |
| --qembedding-group-size | auto | Group size for embedding quantization |
| --streaming | off | Export streaming model with ring buffer KV caches (unlimited duration) |
| --max-enc-len | 750 | Encoder sliding window size (streaming only) |
| --sliding-window | from params.json | Decoder sliding window size (streaming only; ignored in offline mode). Smaller values reduce memory and improve decode speed but limit context |
Notes:
fpa4wquantization requires--backend metal.- The model was trained with
--delay-tokens 6. Other values may degrade accuracy. - The decoder sliding window controls how far back the decoder can attend. At 80ms/step: 2048 = ~2.7 min, 4096 = ~5.5 min, 8192 = ~10.9 min.
ExecuTorch must be installed from source first (see
Prerequisites). The make targets below handle
building core libraries and the runner binary.
make voxtral_realtime-cpu # XNNPACK (CPU)
make voxtral_realtime-metal # Metal (Apple GPU)
make voxtral_realtime-cuda # CUDA (NVIDIA GPU)
make voxtral_realtime-rocm # ROCm (AMD GPU, experimental)The CPU, CUDA, Metal, and MLX targets produce the runner at
cmake-out/examples/models/voxtral_realtime/voxtral_realtime_runner. ROCm uses
the isolated path below.
From the ExecuTorch repository root, the explicit ROCm target builds and installs the LLM runtime and the model runner into a separate directory:
make voxtral_realtime-rocmThe equivalent workflows are:
cmake --workflow --preset llm-release-rocm
cd examples/models/voxtral_realtime
cmake --workflow --preset voxtral-realtime-rocmThe runner is written to
cmake-out-rocm-llm/examples/models/voxtral_realtime/voxtral_realtime_runner.
The separate directory keeps ROCm configuration out of default and CUDA build
caches.
For an end-to-end BF16 or W4/BF16 example:
examples/models/voxtral_realtime/run_rocm_e2e.sh \
~/models/Voxtral-Mini-4B-Realtime-2602 \
/path/to/input-16khz-mono.wav \
w4-bf16 \
both \
/path/to/outputThe third argument selects bf16, w4-bf16, or both precision modes. The
fourth selects streaming, offline, or both execution modes. Set ROCM_PATH
if ROCm is installed outside /opt/rocm.
The script reports model export time, PTE/PTD sizes, and RTF computed as runner
inference time divided by WAV duration.
The AOTInductor-generated shared objects use the C++ runtime from the active Python environment. Add that runtime before launching the runner directly:
PYTHON_PREFIX="$(python -c 'import sys; print(sys.prefix)')"
export LD_LIBRARY_PATH="$PYTHON_PREFIX/lib:${LD_LIBRARY_PATH:-}"Without it, systems with an older /lib64/libstdc++.so.6 can fail while loading
the extracted delegate library with a missing GLIBCXX version.
On Windows (PowerShell), use CMake workflow presets from the executorch root
directory. If you exported with 4-bit quantization, specify your GPU's compute
capability to avoid "invalid device function" errors (the int4mm kernels
require SM 80+).
cmake --workflow --preset llm-release-cuda
Push-Location examples/models/voxtral_realtime
cmake --workflow --preset voxtral-realtime-cuda
Pop-LocationThis builds ExecuTorch with CUDA backend support. The runner binary is at the same path as above. Requires NVIDIA GPU with CUDA toolkit installed.
make voxtral_realtime-metalThis builds ExecuTorch with Metal backend support. The runner binary is at the same path as above. Metal exports can only run on macOS with Apple Silicon.
make voxtral_realtime-mlxThis builds ExecuTorch with MLX backend support. MLX provides GPU acceleration on Apple Silicon via the MLX delegate.
The runner requires:
model.pte— exported model (see Export)tekken.json— tokenizer from the model weights directorypreprocessor.pte— mel spectrogram preprocessor (see Preprocessor)- A 16kHz mono WAV audio file (or live audio via
--mic)
cmake-out/examples/models/voxtral_realtime/voxtral_realtime_runner \
--model_path voxtral_rt_exports/model.pte \
--tokenizer_path ~/models/Voxtral-Mini-4B-Realtime-2602/tekken.json \
--preprocessor_path voxtral_rt_exports/preprocessor.pte \
--audio_path input.wav \
--streamingOmit --streaming for offline transcription (requires an offline-exported
model and offline preprocessor).
For CUDA and ROCm, add
--data_path voxtral_rt_exports/aoti_cuda_blob.ptd.
Windows (PowerShell):
.\cmake-out\examples\models\voxtral_realtime\Release\voxtral_realtime_runner.exe `
--model_path voxtral_rt_exports\model.pte `
--data_path voxtral_rt_exports\aoti_cuda_blob.ptd `
--tokenizer_path C:\path\to\tekken.json `
--preprocessor_path voxtral_rt_exports\preprocessor.pte `
--audio_path input.wav `
--streamingUse --mic to read raw 16kHz float32 PCM from stdin. Requires a
streaming-exported model and streaming preprocessor. Pipe from any audio
capture tool:
# macOS
ffmpeg -f avfoundation -i ":0" -ar 16000 -ac 1 -f f32le -nostats -loglevel error pipe:1 | \
cmake-out/examples/models/voxtral_realtime/voxtral_realtime_runner \
--model_path voxtral_rt_exports/model.pte \
--tokenizer_path ~/models/Voxtral-Mini-4B-Realtime-2602/tekken.json \
--preprocessor_path voxtral_rt_exports/preprocessor.pte \
--micCtrl+C stops recording and flushes remaining text.
| Flag | Default | Description |
|---|---|---|
--model_path |
model.pte |
Path to exported model |
--data_path |
(none) | Path to delegate data file (.ptd, required for CUDA and ROCm) |
--tokenizer_path |
tekken.json |
Path to Tekken tokenizer |
--preprocessor_path |
(none) | Path to mel preprocessor .pte |
--audio_path |
(none) | Path to 16kHz mono WAV file |
--temperature |
0.0 |
Sampling temperature (0 = greedy) |
--offline_max_new_tokens |
500 |
Offline-only: maximum extra tokens after audio embeddings are exhausted |
--streaming |
off | Use streaming transcription (from WAV file) |
--mic |
off | Live microphone mode (reads raw f32le PCM from stdin) |
--mic_chunk_ms |
80 |
Mic read chunk size in ms (multiples of 80 recommended) |
--color |
(none) | Output text color: green or red |
- Audio format: Input must be 16kHz mono WAV. Convert with
ffmpeg -i input.mp3 -ar 16000 -ac 1 output.wav. - OOM during export: Reduce
--max-seq-len(offline mode) or skip encoder quantization (--qlinear-encoder). - "Model was not exported with --streaming": Re-export with the
--streamingflag. Both--streamingand--micrunner modes require a streaming-exported model. fpa4werror: This quantization requires--backend metal.- Metal runner fails with
Library not loaded: @rpath/libc++.1.dylib: The AOTInductor-compiled.soinside the.ptereferenceslibc++via@rpath, which can't be resolved when extracted to a temp directory. Add/usr/libtoDYLD_LIBRARY_PATHso dyld finds it in the shared cache:DYLD_LIBRARY_PATH=/usr/lib \ cmake-out/examples/models/voxtral_realtime/voxtral_realtime_runner ... - Metal runner fails with
Library not loaded: libomp.dylib: The AOTInductor-compiled.solinks against OpenMP. Install it via Homebrew and add it toDYLD_LIBRARY_PATH:brew install libomp DYLD_LIBRARY_PATH=/usr/lib:$(brew --prefix libomp)/lib \ cmake-out/examples/models/voxtral_realtime/voxtral_realtime_runner ...