FusionGuard is a lightweight research-oriented Python library for runtime-adaptive fusion and INT8 selection in Transformer MLP inference.
It dynamically benchmarks multiple execution strategies and selects the fastest configuration for the current hardware environment.
Modern inference runtimes rely heavily on static heuristics to decide whether operator fusion or quantization should be applied. However, fusion profitability and quantization performance depend on hardware architecture, memory bandwidth, batch size and model dimensionality. FusionGuard introduces a minimal runtime benchmarking engine that empirically selects the optimal execution strategy for Transformer MLP blocks. The system evaluates fused vs unfused and FP32 vs dynamic INT8 variants and automatically deploys the fastest configuration.
Inference optimization decisions in systems such as:
- NVIDIA TensorRT
- Google XLA
- Intel oneDNN
- PyTorch Inductor
are often made at compile time using heuristic cost models.
However:
- Fusion profitability depends on memory bandwidth regime.
- Quantization speedups depend on backend engine (FBGEMM, QNNPACK).
- Small-batch inference behaves differently than large-batch.
- CPU vs GPU characteristics differ significantly.
- ARM and x86 quantization backends differ.
There is currently no minimal runtime empirical decision layer in PyTorch that:
- Benchmarks fusion vs non-fusion
- Benchmarks quantized vs non-quantized
- Selects fastest variant dynamically
- Remains lightweight and reproducible
FusionGuard fills this gap.
FusionGuard evaluates four execution variants:
| Variant | Precision | Fusion | Description |
|---|---|---|---|
| FP32 Unfused | float32 | No | Baseline Linear → GELU → Linear |
| FP32 Fused | float32 | Yes | Reduced intermediate overhead |
| INT8 Unfused | int8 dynamic | No | Dynamic quantized linear layers |
| INT8 Fused | int8 dynamic | Yes | Fusion + quantization |
FusionGuard measures average per-iteration latency:
where:
-
$T_{\text{start}}$ denotes the timestamp immediately before timed execution, -
$T_{\text{end}}$ denotes the timestamp immediately after timed execution, -
$N$ denotes the number of benchmark iterations.
Procedure:
- Warm-up runs (10 iterations)
- Timed execution loop (default 50 iterations)
- CUDA synchronization if GPU present
- Average latency computation
- Fastest variant selection
- Python ≥ 3.10
- PyTorch ≥ 2.0
git clone https://github.com/YOUR_USERNAME/fusionguard.git
cd fusionguard
python3.10 -m venv venv
source venv/bin/activate
pip install torch
pip install -e .from fusionguard import FusionGuard
import torch
guard = FusionGuard(dim=768, hidden_dim=3072)
x = torch.randn(16, 768)
y = guard(x)
After initialization, guard(x) automatically uses the optimal execution path.
FusionGuard supports:
- CPU-only environments
- CUDA-enabled systems
- ARM and x86 architectures
- Automatic quantization backend selection If a quantization backend is unavailable, FP32 variants remain functional.
- Reproducibility is ensured via:
- Deterministic iteration counts
- Explicit CUDA synchronization
- Fixed input tensor shapes
- No stochastic graph rewriting
- Dynamic INT8 primarily benefits CPU inference.
- CUDA INT8 dynamic quantization is limited.
- Fusion is implemented at the PyTorch module level (not kernel-level fusion).
- No persistent caching across sessions.
- Limited to Transformer MLP blocks (no attention yet).
Fusion is beneficial when:
- Kernel launch overhead dominates
- Intermediate activation writes are costly
- Memory traffic is a bottleneck
Quantization is beneficial when:
- Compute-bound regime dominates
- INT8 backend is optimized
- Memory bandwidth is constrained
These behaviors align with the Roofline performance model framework.
If you use FusionGuard in research, please cite:
@software{fusionguard2026,
title = {FusionGuard: Runtime Adaptive Fusion and INT8 Selection for Transformer Inference},
author = {Girisha Malni N, Syed Ameen G},
year = {2026},
url = {https://github.com/Girisha-Malni-builds01/fusionguard}
}