Vitality-Aware Compression for Efficient Image-to-Shape Diffusion Transformers
Abstract
Image-to-shape Diffusion Transformers (DiTs) achieve high geometric fidelity but their backbones often exceed 2.5 GB, while existing diffusion compression methods designed for image or video transfer poorly to 3D synthesis. We propose the first vitality-aware adaptive compression framework for image-to-shape DiTs, scoring each layer via per-layer ablation with Earth Mover’s Distance on generated point clouds. The resulting non-uniform importance, with distinct double vs. single-block patterns, guides structured pruning, mixed-precision quantization, and selective fine-tuning under block-type-specific thresholds. Across three state-of-the-art models, our method achieves up to 66% model-size reduction while preserving synthesis quality, demonstrating a favorable quality–resource trade-off and a plug-and-play path to efficient 3D foundation models.