Consistent Distribution Matching
for Data-Free Diffusion Distillation

1University of British Columbia 2Vector Institute for AI 3Canada CIFAR AI Chair
Particles leave the prior distribution \(p_1=q_1\) along the student flow. On the way, the teacher’s velocity field corrects them through \(\mathcal{L}_{\mathrm{CDMD}}\), until the student distribution \(q_0\) matches the teacher distribution \(p_0\).

Key Features

Data-Free

No distillation dataset. Training states are bootstrapped from the student itself.

Simulation-Free

No teacher rollouts. Each state costs one student pass plus analytic noising.

Two Models, One Loss

A frozen teacher and a student. No fake score network, no GAN. 20 GB vs. 33–37 GB.

Any-Step Sampling

FID 2.04 in one step and 1.37 in four on ImageNet 256×256, from the same student.

Abstract

Flow and diffusion models suffer from slow inference due to computationally expensive numerical integration. Distillation provides a promising way for a student model to learn from a teacher's dynamics, enabling one-step or few-step generation. However, existing methods often depend on curated distillation datasets, costly teacher rollouts, or auxiliary proxy networks, which complicate model training and scaling.

In this work, we propose Consistent Distribution Matching, a simulation-free and data-free distillation method for accelerating diffusion and flow models while preserving strong generative capacity. Our key insight is to unify sample generation and score estimation with one student network. Thus, our framework uses only two models, a frozen teacher and a trainable student, and optimizes one objective. We prove that minimizing our objective indicates Wasserstein convergence of the student flow-map pushforwards to the teacher marginals. On ImageNet 256×256, our method attains an FID of 2.04 with a single function evaluation (1-NFE) and a 4-NFE FID of 1.37 within 40 epochs of training, surpassing the state-of-the-art distillation baselines without data.

Method

Data-based distillation trains the student on noised data according to a prescribed forward path, but those states are not the ones the teacher visits when it generates. Simulating the teacher model fixes the mismatch, at a high price. CDMD needs neither: one student network is both the few-step generator and the estimator of its own velocity, so training involves only the frozen teacher and the student.

From score distillation to CDMD

All three push the student \(q_t\) toward the teacher \(p_t\). The difference is where the fake term comes from.

DMD reverse KL
\[\min_\theta\;\mathbb{E}_t\,D_{\mathrm{KL}}\!\left(q_t\,\|\,p_t\right)\]
\[\nabla_\theta\approx\mathbb{E}\Big[w_t\big(\colorbox{#fde8e5}{$\textcolor{#c4473a}{s_{\text{fake}}}$}-s_{\text{real}}\big)^{\!\top}\tfrac{\partial G_\theta(z)}{\partial\theta}\Big]\]

\(s_{\text{fake}}\) requires training a third model by score matching.

3 models · 2 optimizers
SiD Fisher divergence
\[\min_\theta\;\mathbb{E}_t\,D_{\mathrm{F}}\!\left(q_t\,\|\,p_t\right)\]
\[D_{\mathrm{F}}=\mathbb{E}_{\hat{x}\sim q_t}\big\|\colorbox{#fde8e5}{$\textcolor{#c4473a}{\nabla\log q_t}$}-\nabla\log p_t\big\|^2\]

\(\nabla\log q_t\) requires training a third model by score matching.

3 models · 2 optimizers
CDMDFisher divergence, velocity form
\[\min_\theta\;\mathbb{E}\,\big\|\,v_{\text{fake}}-v_{\text{real}}\big\|^2\]
\[v_{\text{fake}}\approx\colorbox{#e3f3ec}{$\textcolor{#22785f}{u_\theta(\hat{x}_t,t,r)+(t-r)\tfrac{\mathrm{d}}{\mathrm{d}t}u_\theta(\hat{x}_t,t,r)}$}\]

\(v_{\text{fake}}\) comes from the student itself. No third model.

2 models · 1 optimizer

DMD: Yin et al., Improved Distribution Matching Distillation for Fast Image Synthesis, 2024. SiD: Zhou et al., Score identity Distillation, 2024. The score and velocity forms of these divergences are equivalent up to a time-dependent weight (see the paper's appendix).

Pipeline

CDMD pipeline. (a) Bootstrap states: the student maps noise to an earlier time s, then the state is analytically renoised to time t. (b) Distribution matching: the teacher velocity and the student's JVP-corrected velocity are compared to form the CDMD loss.
aBootstrap states

Jump from noise to time \(s\) with the student, then renoise to \(t\):

\[\hat{x}_t=\tfrac{1-t}{1-s}\,f_\theta(z,1,s)+\tfrac{t-s}{1-s}\,\epsilon,\qquad z,\epsilon\sim p_1 .\]

One student pass per state: no data, no teacher rollouts.

bDistribution matching

Regress the student's fake velocity onto the teacher's:

\[\mathcal{L}_{\mathrm{CDMD}}=\mathbb{E}\,\big\|\,u_\theta+(t-r)\,\mathrm{sg}\big[\tfrac{\mathrm{d}}{\mathrm{d}t}u_\theta\big]-v_{\text{real}}(\hat{x}_t,t)\big\|^2 .\]

One JVP per step. Sample in one or a few steps.

What minimizing \(\mathcal{L}_{\mathrm{CDMD}}\) guarantees

Under regularity assumptions, and provided the student's training states cover the teacher's path, we prove two results.

  1. The distributions match. As the CDMD objective approaches its minimum of zero, the student's samples at every intermediate time approach the teacher's marginals in Wasserstein-2 distance, averaged over time:
    \[\int_{[0,1)} W_2^2\!\Big(\big(f_\theta(\cdot,1,r)\big)_{\#}\,p_1,\;p_r\Big)\,\nu(\mathrm{d}r)\;\longrightarrow\;0 .\]
  2. The transport is identified. If the student can represent the teacher's average velocity \(u\), the lowest achievable objective is zero, and every minimizer recovers \(u\) on its training states. The student learns how the teacher moves, not only where it ends up.

Simpler than prior distillation pipelines

Method# Models# OptimizersData-freeSimulation-freeNo GAN
BOOT21✓✗✓
DMD232✓†✓†✗†
SiD-LSG32✓✓✓
SiDA32✗✓✗
FreeFlow32✓✓✓
CTM32✗✓✗
FACM21✗✓✓
MeanFlow / CM / IMM21✗✓✓†
CDMD (ours)21✓✓✓

† depends on the training setup. A GAN objective needs real data for its discriminator.

Results

ImageNet 256×256: uncurated samples

Uncurated samples from the CDMD student distilled from LightningDiT-XL/2, generated without classifier-free guidance, so each step is a single network evaluation (1 step = 1 NFE). Multi-step samples use the Euler sampler.

Best data-free FID with every teacher

MethodNFEFID ↓IS ↑
Teacher250×21.35295.3
FreeFlow14.28264.9
DMD2†13.43317.2
MFD†16.67281.6
MeanFlow☆14.19213.8
sCM☆14.21266.9
FACM☆1 / 22.15 / 1.79303.8 / 319.6
CDMD (ours)1 / 2 / 42.04 / 1.62 / 1.37294.3 / 317.3 / 306.9

Data-free distillation on class-conditional ImageNet 256×256 under matched training budgets. † reproduced by us; ☆ adapted by us to the data-free setting.

1-step FID versus training compute on ImageNet 256 by 256. CDMD reaches the lowest FID with the least compute and 20 GB of training memory, versus 33 GB for DMD2 and 37 GB for FreeFlow.
1-step FID vs. training compute. Labels give training memory on one A100 under the same configuration.
One student, any number of steps
1 step2.04
2 steps1.62
4 steps1.37
Teacher, 250×2 NFE1.35

FID-50K on ImageNet 256×256, LightningDiT-XL/2 teacher.

Inference-time scaling with the LightningDiT-XL/2 teacher. Left: FID against sampling steps from 1 to 64 for Euler and DDIM; Euler drops from 2.04 to near the teacher reference by 4 to 8 steps and stays there up to 64 steps, while DDIM rises above 5 by 64 steps. Right: Inception Score; DDIM keeps increasing to about 420, Euler stays around 310 to 320.
Inference-time scaling with the LightningDiT-XL/2 teacher. FID (left) and IS (right) against the number of sampling steps for Euler and DDIM, using the same distilled student (trained with \(t=r\) for 50% of the pairs). Dashed line: teacher reference.
How does one student take any number of steps?

CDMD learns the average velocity \(u_\theta(x_t,t,r)\) over any time interval \([r,t]\), not only from pure noise to data. The student can therefore jump from noise to an image in one step, or split the trip into shorter jumps by chaining its own flow map over consecutive intervals:

\[\hat{x}_{t_{k+1}}=\hat{x}_{t_k}-(t_k-t_{k+1})\,u_\theta(\hat{x}_{t_k},t_k,t_{k+1}),\qquad 1=t_0>t_1>\dots>t_K=0 .\]

Each jump is one network evaluation, and no retraining is needed to change the number of steps.

  • Euler keeps the teacher's multi-step behavior. FID improves up to 8 steps and then stays close to the teacher reference all the way to 64 steps, instead of drifting away.
  • The sampler matters. DDIM raises IS but degrades FID beyond 4 steps.
  • Prior-anchored students cannot do this. FreeFlow's student only maps from pure noise at \(t=1\), so it has no sub-interval map to take extra steps with.

Text-to-image: distilling SANA-1.6B to 4 steps

Curated 1024 by 1024 text-to-image samples from the 4-step CDMD student distilled from SANA-1.6B.
Curated 1024×1024 samples from our 4-step CDMD generator distilled from SANA-1.6B. Click to enlarge.
ModelNFEAesthetic ↑PickScore ↑HPSv2 ↑ImageReward ↑CLIPScore ↑FID ↓
SANA teacher206.3350.20210.24170.836330.71730.39
SANA teacher45.9760.18720.2173−0.100828.25135.65
SenseFlow46.4820.19980.24450.792529.321–
AYF46.2310.19630.22980.702229.176–
Diff-Instruct46.4530.19850.24340.767729.393–
Consistency Distillation†46.2480.19060.22630.674328.89054.38
DMD2†46.4890.20060.24390.804229.10332.34
MFD†46.5410.20020.24420.848329.43028.77
CDMD (ours)46.9060.20110.23870.886428.99227.83

Bold marks the best learned 4-NFE result; † reproduced by us. FID against COCO-5K 2017. CDMD trains in about 7 GPU hours, vs. 22 for Diff-Instruct, 25 for MFD and 30 for DMD2.

BibTeX

@article{fu2026cdmd,
  title   = {Consistent Distribution Matching for Data-Free Diffusion Distillation},
  author  = {Fu, Yuxiang and Yan, Qi and Wu, Zike and Zhang, Yongxing and
             Abolmaesumi, Purang and Wang, Lele and Liao, Renjie},
  journal = {arXiv preprint},
  year    = {2026}
}