Skip to content
kaist-amiPublic

About

[TMLR’26] Official PyTorch Implementation of “Revisiting Learning-based Video Motion Magnification for Real-time Processing“

Resources

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

Fast-MM: DMM / FastDMM inference

Minimal command-line inference for learning-based video motion magnification. It bundles two model configurations and their checkpoints:

FastDMM: Project site · arXiv · OpenReview

DMM / Learning-based Video Motion Magnification: Project site · arXiv

  • dmm: the Oh et al. DMM architecture (dmm_oh.tar)
  • fast-dmm: the Ha et al. FAST-DMM architecture (fast_dmm_ha.tar)

The implementation supports dynamic, static, Chebyshev-I IIR, FIR, and Butterworth temporal modes; automatically preserves the input frame rate unless overridden; and accepts arbitrary frame dimensions by mirror-padding internally and cropping back to the exact input size.

Requirements

  • Linux (tested) or another OS with Python 3.10+
  • FFmpeg with the libx264rgb encoder
  • An NVIDIA GPU is recommended. CPU inference is supported but slow.

On Ubuntu/Debian:

sudo apt update
sudo apt install -y python3-venv ffmpeg
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip

Install PyTorch using the command for your CUDA version from the official PyTorch installer. For example, for a supported CUDA 12.8 environment:

pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt

For CPU-only use:

pip install torch torchvision --index-url https://download.pytorch.org/whl/cpu
pip install -r requirements.txt

Check the installation:

python run.py --help

Quick start

Dynamic DMM inference:

python run.py \
  --input inputs/example.avi \
  --output-dir outputs \
  --model dmm \
  --filter dynamic \
  --alpha 10

Frequency-selective FAST-DMM inference:

python run.py \
  --input inputs/example.avi \
  --output-dir outputs \
  --model fast-dmm \
  --filter butter \
  --low 0.01 \
  --high 0.10 \
  --alpha 10

Process all supported videos directly under one input directory:

python run.py \
  --input-dir inputs \
  --output-dir outputs \
  --model fast-dmm \
  --filter fir \
  --low 0.02 \
  --high 0.15 \
  --alpha 5 \
  --fir-taps 31

Use --device cpu, --device cuda, or --device cuda:1 to override automatic device selection. Use --checkpoint PATH to override the checkpoint associated with the selected model.

Temporal modes

Mode Reference / filter Frequency arguments
dynamic Difference from the immediately preceding frame representation None
static Difference from the first frame representation None
iir Causal Chebyshev type-I bandpass; order 2 and 0.5 dB ripple by default Required
fir Centered Hamming-window linear-phase bandpass; 31 taps by default Required
butter Causal Butterworth bandpass; order 2 by default Required

Frequency values are in cycles per sample, not directly in Hz. They must satisfy:

0 < low < high < 0.5

0.5 is the Nyquist frequency. Convert a ratio to Hz using:

frequency_hz = frequency_ratio * sampling_fps

For a 120 fps video, --low 0.01 --high 0.10 selects approximately 1.2–12 Hz. The CLI reads the sampling FPS from each input video, and the output filename records both normalized and Hz values.

The first frame initializes temporal state and is emitted unchanged. Short videos, including a single-frame input, are therefore valid. IIR and Butterworth modes are causal. FIR uses a centered sliding feature window so its group delay is aligned with the source frame; it buffers only taps feature frames rather than the complete video. Temporal edges are replicated. An even --fir-taps value is promoted to the next odd value.

Output layout

Each input gets its own subdirectory:

outputs/
└── INPUT_VIDEO_NAME/
    └── YYYYMMDD_HHMMSS_INPUT_VIDEO_NAME_FILTER_[FREQUENCY]_xALPHA.mp4

Example:

outputs/example/20260807_193015_example_butter_0p01-0p1cyps_1p2-12Hz_x10.mp4

FFmpeg encodes RGB H.264 so odd widths/heights are preserved. Before model inference, frames are symmetrically mirror-padded to at least 32×32 and to the architecture's required spatial multiple. Output is cropped back to the original dimensions.

Citation

@misc{ha2024revisiting,
  title         = {Revisiting Learning-based Video Motion Magnification for Real-time Processing},
  author        = {Ha, Hyunwoo and Oh, Hyun-Bin and Kim, Jun-Seong and Kwon, Byung-Ki and Kim, Sung-Bin and Tran, Linh-Tam and Kim, Ji-Yun and Bae, Sung-Ho and Oh, Tae-Hyun},
  year          = {2024},
  eprint        = {2403.01898},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  doi           = {10.48550/arXiv.2403.01898}
}

About

[TMLR’26] Official PyTorch Implementation of “Revisiting Learning-based Video Motion Magnification for Real-time Processing“

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages