Minimal command-line inference for learning-based video motion magnification. It bundles two model configurations and their checkpoints:
FastDMM: Project site · arXiv · OpenReview
DMM / Learning-based Video Motion Magnification: Project site · arXiv
dmm: the Oh et al. DMM architecture (dmm_oh.tar)fast-dmm: the Ha et al. FAST-DMM architecture (fast_dmm_ha.tar)
The implementation supports dynamic, static, Chebyshev-I IIR, FIR, and Butterworth temporal modes; automatically preserves the input frame rate unless overridden; and accepts arbitrary frame dimensions by mirror-padding internally and cropping back to the exact input size.
- Linux (tested) or another OS with Python 3.10+
- FFmpeg with the
libx264rgbencoder - An NVIDIA GPU is recommended. CPU inference is supported but slow.
On Ubuntu/Debian:
sudo apt update
sudo apt install -y python3-venv ffmpeg
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pipInstall PyTorch using the command for your CUDA version from the official PyTorch installer. For example, for a supported CUDA 12.8 environment:
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txtFor CPU-only use:
pip install torch torchvision --index-url https://download.pytorch.org/whl/cpu
pip install -r requirements.txtCheck the installation:
python run.py --helpDynamic DMM inference:
python run.py \
--input inputs/example.avi \
--output-dir outputs \
--model dmm \
--filter dynamic \
--alpha 10Frequency-selective FAST-DMM inference:
python run.py \
--input inputs/example.avi \
--output-dir outputs \
--model fast-dmm \
--filter butter \
--low 0.01 \
--high 0.10 \
--alpha 10Process all supported videos directly under one input directory:
python run.py \
--input-dir inputs \
--output-dir outputs \
--model fast-dmm \
--filter fir \
--low 0.02 \
--high 0.15 \
--alpha 5 \
--fir-taps 31Use --device cpu, --device cuda, or --device cuda:1 to override automatic device selection. Use --checkpoint PATH to override the checkpoint associated with the selected model.
| Mode | Reference / filter | Frequency arguments |
|---|---|---|
dynamic |
Difference from the immediately preceding frame representation | None |
static |
Difference from the first frame representation | None |
iir |
Causal Chebyshev type-I bandpass; order 2 and 0.5 dB ripple by default | Required |
fir |
Centered Hamming-window linear-phase bandpass; 31 taps by default | Required |
butter |
Causal Butterworth bandpass; order 2 by default | Required |
Frequency values are in cycles per sample, not directly in Hz. They must satisfy:
0 < low < high < 0.5
0.5 is the Nyquist frequency. Convert a ratio to Hz using:
frequency_hz = frequency_ratio * sampling_fps
For a 120 fps video, --low 0.01 --high 0.10 selects approximately 1.2–12 Hz. The CLI reads the sampling FPS from each input video, and the output filename records both normalized and Hz values.
The first frame initializes temporal state and is emitted unchanged. Short videos, including a single-frame input, are therefore valid. IIR and Butterworth modes are causal. FIR uses a centered sliding feature window so its group delay is aligned with the source frame; it buffers only taps feature frames rather than the complete video. Temporal edges are replicated. An even --fir-taps value is promoted to the next odd value.
Each input gets its own subdirectory:
outputs/
└── INPUT_VIDEO_NAME/
└── YYYYMMDD_HHMMSS_INPUT_VIDEO_NAME_FILTER_[FREQUENCY]_xALPHA.mp4
Example:
outputs/example/20260807_193015_example_butter_0p01-0p1cyps_1p2-12Hz_x10.mp4
FFmpeg encodes RGB H.264 so odd widths/heights are preserved. Before model inference, frames are symmetrically mirror-padded to at least 32×32 and to the architecture's required spatial multiple. Output is cropped back to the original dimensions.
@misc{ha2024revisiting,
title = {Revisiting Learning-based Video Motion Magnification for Real-time Processing},
author = {Ha, Hyunwoo and Oh, Hyun-Bin and Kim, Jun-Seong and Kwon, Byung-Ki and Kim, Sung-Bin and Tran, Linh-Tam and Kim, Ji-Yun and Bae, Sung-Ho and Oh, Tae-Hyun},
year = {2024},
eprint = {2403.01898},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
doi = {10.48550/arXiv.2403.01898}
}