İçeriğe geç / Skip to content / Zum Inhalt

DMAD and PDMD: 4-Step AI Video Generation Explained

Ahmet Balaman

6 min read

AI VideoDMADPDMDByteDanceMiniMax H3Distillation
DMAD and PDMD: 4-Step AI Video Generation Explained

ByteDance released two methods this week that speed up video and image generation: DMAD and PDMD. Both aim at the same thing: running a video model such as MiniMax H3 in just 4 steps instead of dozens. With fewer steps, generation time drops directly. The most striking finding in this article: in the measurement on PDMD's project page, the 4-step model beats the 50-step teacher model's total score.

This article summarizes the relevant sections of the AI Search channel's weekly AI news video on YouTube and extends them with the measurements on the two projects' own pages; the video link is at the end. This is a news roundup, and I have not tested the methods. The numbers are the researchers' own measurements, not independent reproductions.

What is distillation, and why 4 steps?

Diffusion-based video models build an image out of noise step by step. Good quality often takes dozens of steps; on PDMD's page the teacher model uses 50 "network evaluations" (NFE). Distillation teaches the behavior of a large, slow "teacher" model to a "student" model that runs in few steps. The goal is to cut the step count while preserving quality. DMAD and PDMD bring it down to 4 steps; PDMD also shows 2-step examples.

What is DMAD?

DMAD stands for "Distribution Matching as Adversarial Distillation"; it is work by researchers at Texas A&M University and ByteDance. How it works:

  1. Noise passes through the few-step student to produce a sample.
  2. Teacher, student and real samples are each noised and classified by two discriminator heads that share one backbone.
  3. Gap-based weights decide how the student is updated.

Results on the page:

  • One-step images (ImageNet 64×64): FID 1.04.
  • Text to image (SDXL, 4 steps, COCO-10K): FID 14.47.
  • Text to video (Wan2.1-T2V-14B, 4 steps): VBench 85.15.
  • Human preference (MiniMax-H3, against rCM): 84.6% excluding ties.

The same architecture works not only on MiniMax but on video and image models such as Wan2.1 and SDXL. Code, models and a demo are released. According to the video, to use it you download the base MiniMax model and add LoRA files of about 1.4 GB each. Project page: yzmblog.github.io/projects/DMAD.

What is PDMD?

PDMD is "Projected Distribution Matching Distillation"; work by ByteDance and UC San Diego. The name resembles DMAD, but the idea differs: a one-line projection filters the student's training error.

The problem: a critic finds the difference between student and teacher and corrects the student, but the critic is not perfect and its own errors pass into the student. PDMD removes the component of the training signal that is parallel to the critic's error estimate (the video's phrase: "projects it away"). According to the page nothing else changes: no auxiliary loss, no additional network, no extra model pass. The code shown on the page really is one line:

r = x0_critic - x0_student
d = d - (d * r).sum() / r.pow(2).sum() * r

PDMD measurements

VideoGen-Eval (387 prompts, 544p, MiniMax-H3-33B, total score):

Method Steps (NFE) Total Dynamic
MiniMax-H3-33B teacher 50 82.41 66.67
Undistilled model 4 79.48 44.44
H3 Turbo LoRA 4 81.57 54.78
DMD2 4 82.27 59.95
DMD 4 82.76 61.76
rCM 4 81.18 58.40
AnyFlow 4 81.97 64.60
PDMD 4 83.17 71.83

On VBench with Wan2.1-T2V-1.3B, PDMD is also ahead with a total score of 83.73, above the teacher (83.06) and the other 4-step methods; its dynamic score is 89.72. The page states that DMD and PDMD differ only by the projection, so the gain comes directly from that one line. The page also has a user study and qualitative comparisons.

The output samples were produced with MiniMax-H3-33B at 1344×768, 345 frames at 24 fps (about 14 seconds), each with four network evaluations. Weights are published on Hugging Face. According to the video there are two options: the full model of about 66 GB (running in 4 steps) or a LoRA of about 1.4 GB attached on top of a smaller, quantized MiniMax model. Project page: pdmd2026.github.io.

DMAD and PDMD side by side

DMAD PDMD
Researchers Texas A&M, ByteDance UC San Diego, ByteDance
Goal Generation in 4 steps Generation in 4 (and 2) steps
Idea Two discriminator heads, shared backbone One line removing critic error from the training signal
Models shown MiniMax-H3, Wan2.1, SDXL, EDM MiniMax-H3-33B, Wan2.1-T2V-1.3B
Released Code, models, demo; LoRA (~1.4 GB) Code, weights; full model (66 GB) and LoRA (1.4 GB)

What does it mean in practice?

Going from 50 steps to 4, a compute saving of nearly ten times is theoretically possible on the same hardware, but real time depends on the model's other components and on the hardware. The pages give no wall-clock times, so I do not state an exact speed multiplier. Things to watch:

  • The measurements rest on the researchers' own benchmarks; compare both settings side by side with your own prompts.
  • Accelerated models often lose some detail or variety compared with the full-step model; PDMD's high dynamic score suggests it reduces the stiff-motion problem, but that is also its own measurement.
  • The LoRA route is the cheapest way to gain speed without downloading a large model from scratch.

For another way to cut video generation time, see the same week's InSpatio World 1.5, Sol Refiner and PixelUMM article. For installing an open-source video model locally, there are How to install LTX 2.5 and LTX 2.5 requirements; the Comfy Agent article helps with building workflows with an agent in ComfyUI.

Frequently Asked Questions

What are DMAD and PDMD for?

They speed up generation by running a video model in 4 steps. Both are distillation methods.

Is PDMD really better than the 50-step model?

In the VideoGen-Eval measurement on PDMD's page, the 4-step PDMD's total score (83.17) is above the 50-step MiniMax-H3 teacher's score (82.41). That is the researchers' own measurement; it does not mean superiority in every comparison.

Which models do they work on?

The pages show MiniMax-H3, Wan2.1 and (in DMAD) SDXL. DMAD's architecture can also be applied to other video and image models.

How big are the LoRA files?

According to the video, about 1.4 GB each. PDMD's full model is about 66 GB.

What is the difference between DMAD and PDMD?

DMAD uses the student-teacher gap with two discriminator heads. PDMD removes the critic's error from the training signal with a one-line projection.

Source

Comments