DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation
Abstract
Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student's evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adversarial Distillation, which recasts distribution matching as classification and learns the required log-density ratios directly. Two discriminator heads on a shared backbone distinguish real data and teacher samples from the student's, and linear losses on their logits train the student without auxiliary score fitting. We prove that at the discriminator optimum these losses recover the distribution-matching gradient underlying DMD, through the classical identity linking discriminator logits to log-density ratios. We further introduce gap-based reweighting, which adapts teacher supervision across noise levels from the real-data head's empirical logit gap between real and teacher samples. DMAD reaches a Frรฉchet Inception Distance (FID) of 1.04 with one-step generation on ImageNet-64x64, 14.47 with four-step SDXL on COCO-10K, and a VBench total score of 85.15 with four-step Wan2.1-T2V-14B, the best values among the compared few-step methods and the multi-step teachers. On MiniMax-H3-33B, our four-step student achieves overall human preference rates of 79.1% over DMD2 and 84.6% over rCM for joint audio-video generation, excluding ties. Our code, models and demos are available at https://yzmblog.github.io/projects/DMAD.
Community
๐ฌ Video + audio in just 4 sampling steps.
DMAD (Distribution Matching as Adversarial Distillation) reduces MiniMax-H3 generation from 50 to 4 steps.
The key idea: learn distribution-matching gradients through discriminator log-density ratios, without fitting an auxiliary student-score model. We evaluate the approach across image, video, and joint audio-video generation.
Training and inference code, checkpoints, and ComfyUI integration are available. Watch the samples with sound on! ๐
๐ป Code: https://github.com/Yzmblog/DMAD
๐ Project: https://yzmblog.github.io/projects/DMAD
๐ค Checkpoints: https://huggingface.co/ZhengmingYu/DMAD
๐ฎ Try the demo: https://huggingface.co/spaces/hugging-apps/dmad
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- PDMD: Projected Distribution Matching Distillation for Video Diffusion Models (2026)
- Enhancing Autoregressive Video Generation via Representation Adversarial Distillation (2026)
- DMA$^2$: Pixel-space Distribution Matching with Adversarial and Anchor Losses (2026)
- DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models (2026)
- DuoMatching: Joint-Marginal Distribution Matching for Few-Step Video Generation (2026)
- From Scores to Samples: Elastic Forcing for Autoregressive Video Generation (2026)
- VDOT++: Unified Few-Step Video Generation via Unbalanced Optimal Transport Distillation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 2
nixsn/DMAD
Datasets citing this paper 1
ZhengmingYu/DMAD-H3-data
Spaces citing this paper 4
Collections including this paper 0
No Collection including this paper