From Static to Sunrise: How AI Learned to Paint With Noise
Last Updated on September 25, 2026 by Editorial Team
Author(s): Rajdip Bera
Originally published on Towards AI.
From Static to Sunrise: How AI Learned to Paint With Noise

Picture a photograph sitting on a table — a mountain lake at sunrise, sharp and clear. Now imagine someone takes that photo and sprinkles a tiny bit of random noise onto it, like static on an old television screen. Barely noticeable. Then they do it again. And again. A thousand times.
By the end, the mountain lake is gone. What’s left is pure, meaningless static — the kind of fuzzy gray noise you’d see on a broken TV. There’s no lake, no sunrise, no mountain. Just randomness.
Here’s the strange question that changed the field of generative AI: if you handed someone only that static, could they work backward and recover the original photograph?
It sounds impossible. Noise is, by definition, the thing that destroys information. But researchers studying physics, probability, and thermodynamics had been staring at a similar problem for over a century — how gas particles diffuse through a room, how heat spreads through metal, how order collapses into disorder. And they noticed something interesting: if you understand precisely how a process destroys structure, step by tiny step, you can — at least mathematically — describe how to walk it backward.
That idea, borrowed from physics and dressed up with neural networks, is what we now call a diffusion model. It’s the engine behind image generators like Stable Diffusion, DALL·E’s later versions, and Midjourney’s underlying tech. And despite how futuristic the results look, the core idea is surprisingly intuitive once you see it slowly.
That’s what this article does. No equation appears without an explanation first. No jump from “clean image” to “trained model” without walking through the middle.

Before Diffusion: Why the Earlier Approaches Fell Short
Diffusion models didn’t appear in a vacuum. Before 2020, the two dominant ways to get a neural network to generate images were GANs (Generative Adversarial Networks) and VAEs (Variational Autoencoders). Both were genuinely clever, and both were already producing usable results. But each carried a structural weakness that diffusion models were specifically designed to sidestep.
GANs: brilliant results, miserable to train
A GAN works like a forger and a detective locked in a room together. One network (the generator) tries to produce fake images. The other (the discriminator) tries to catch which images are fake. Over time, the forger gets better at fooling the detective, and the detective gets better at catching fakes — and if the process goes well, the forger ends up producing convincingly real images.
When it works, a GAN can produce razor-sharp, highly realistic images. But “when it works” is doing a lot of heavy lifting in that sentence:
- Training instability. The generator and discriminator are locked in a constant tug-of-war, and that adversarial setup is notoriously delicate to balance. Small changes in learning rate or architecture can tip the whole system into failure.
- Mode collapse. Instead of learning to produce the full variety of images in the training data, the generator sometimes discovers a small handful of outputs that reliably fool the discriminator — and just keeps producing variations of those few, ignoring the rest of the diversity in the dataset.
- No explicit likelihood. A GAN never directly learns “how probable is this image, given everything I’ve seen?” It just learns to fool a critic. That makes GANs hard to evaluate rigorously and hard to reason about mathematically.
- Sensitive, fiddly training. Getting a GAN to converge often meant a lot of trial and error with architecture tweaks, learning rate schedules, and regularization tricks — recipes that didn’t always transfer from one dataset to the next.
VAEs: stable, but blurry
A VAE takes a different approach. It compresses an image down into a compact numerical representation (a “latent code”), then learns to reconstruct the image back out of that code — while also nudging the space of possible codes into a smooth, well-organized shape, so that sampling a random code and decoding it produces something image-like.
VAEs are much more stable to train than GANs — there’s no adversarial tug-of-war — and they come with a clean mathematical foundation for estimating likelihood. But they have their own well-known drawback:
- Blurry outputs. The way VAEs are trained mathematically tends to average over plausible outputs rather than committing to sharp, confident detail, which routinely produces images that look slightly soft or smudged compared to GAN outputs.
- A compression bottleneck. Squeezing an entire image into a compact latent code inevitably throws away fine detail, and that lost detail is difficult to fully recover during decoding.
- A tension in the training objective. VAEs balance two competing goals — reconstructing the image accurately, and keeping the latent space smoothly organized — and pushing harder on one often comes at the expense of the other.
Where that leaves diffusion models
Diffusion models were, in large part, a response to this exact tradeoff. GANs were sharp but unstable and hard to reason about. VAEs were stable but blurry. Diffusion models aimed for a third path: a training process with no adversarial tug-of-war (just a straightforward noise-prediction loss, much like a VAE’s stable setup), but capable of sharp, high-fidelity output (much like a GAN’s strength) — at the cost of needing many sequential steps to generate a single image, which is the tradeoff later methods like DDIM and Latent Diffusion worked hard to fix.
GANs: sharp results, unstable training, prone to mode collapse
VAEs: stable training, blurry results, compression bottleneck
Diffusion: stable training + sharp results, but slower to sample
Meet the Model That Learns by Destroying Things First
Strip away the math for a second. Here’s the entire idea in one sentence:
A diffusion model is a neural network that learns to undo noise, one small step at a time.
To learn that skill, it needs two processes:
- Forward Process (Diffusion): Take a real image and slowly destroy it by adding a little random noise at each step, until nothing recognizable remains.
- Reverse Process (Denoising): Train a neural network to look at a noisy image and guess what noise was added — so that subtracting it, step by step, reconstructs something real.
Here’s an analogy that has nothing to do with images at all. Imagine a drop of black ink released into a glass of clear water. Left alone, it spreads — diffuses — until the water is a uniform gray. That’s the forward process: structure (a concentrated drop) dissolving into disorder (uniform mixture).
Now imagine you had a perfect physical model of exactly how ink diffuses — the precise rules governing how each molecule moves. In principle, you could run that model in reverse: start from the uniform gray water and mathematically “pull” the ink molecules back into a tight, concentrated drop.
Nobody can actually un-mix ink in real life — that would violate the second law of thermodynamics. But in the world of images and numbers, we can build something that plays this trick, because we’re not bound by physical entropy — we’re bound by math, and we get to design the rules ourselves. A neural network doesn’t need to reverse physics; it just needs to learn the pattern of “what noise was probably added here,” which is a solvable problem when the network sees millions of examples.

Step One: Teaching Chaos How to Behave
Let’s slow down and build this from the ground up.
We start with a real image, which we’ll call x₀. Think of x₀ as just a big list of numbers — one number per pixel value (or, more precisely, per color channel of each pixel). For our purposes, imagine we’ve simplified everything down to a single number, like the brightness of one pixel.
At every timestep t (t = 1, 2, 3, … up to some final step T, often 1000), we add a small amount of Gaussian noise to the image from the previous step.
Why Gaussian noise, and not just “random static”?
This is worth pausing on, because it’s not an arbitrary choice — and it’s the detail most explanations skip.
“Gaussian noise” means noise drawn from a bell curve, most sampled values cluster near zero, with values far from zero becoming increasingly rare. That’s different from, say, picking a totally random pixel color with no bias — Gaussian noise has structure, even though it looks structureless.
That structure gives it three properties that make it the only realistic choice here:
- It’s fully described by just two numbers — a mean (the center of the bell curve) and a variance (how spread out it is). Nothing more to track.
- Adding two Gaussian-distributed things together produces another Gaussian. This is the property that matters most. It means you can chain hundreds of tiny noise-adding steps together and the combined effect is still just one clean Gaussian — not some intractable mathematical mess. Without this property, the equations below simply wouldn’t have closed-form solutions.
- It matches how “generic” randomness behaves in the real world. Thanks to the Central Limit Theorem, when lots of small independent random effects pile on top of each other, the result naturally starts looking Gaussian — thermal noise, sensor noise, static, all of it.
So Gaussian noise isn’t chosen because it visually resembles TV static (though it does). It’s chosen because it’s the one kind of randomness that lets a thousand tiny corruption steps stay mathematically tractable, all the way through.
UNDERSTANDING q AND p: THE TWO DISTRIBUTIONS
q - The Forward (Noise-Adding) Distribution
q is the fixed forward process that we design ourselves. It is not learned.
- q(xₜ | xₜ₋₁) - single-step forward transition: add noise to get the next state.
- q(x₁:T | x₀) - full forward trajectory: joint distribution of all noisy versions.
- q(xₜ₋₁ | xₜ, x₀) - true reverse posterior: computable in closed form, serves as the training target.
Key property: q is fully known in advance. It depends only on the noise schedule (βₜ, αₜ, ᾱₜ).
p θ - The Reverse (Denoising) Distribution
p θ is the learned reverse process. The subscript θ denotes network parameters, updated during training.
- p θ(xₜ₋₁ | xₜ) - model's single-step reverse transition: what the network learns.
- p θ(x₀:T) - full reverse trajectory: joint distribution of the denoising chain.
- p θ(x₀) - model's marginal over generated images. The quantity we want to maximize.
Key property: p θ is unknown at the start and is learned by minimizing the loss. It is the generative model.
Why Both Appear Together
ℒ = E q[ D KL( q(xₜ₋₁ | xₜ, x₀) ‖ p θ(xₜ₋₁ | xₜ) ) ]
q on the left of the KL - the target, computed with knowledge of x₀.
p θ on the right of the KL - the prediction, computed by the network from xₜ alone.
Training pushes p θ toward q at every timestep. When they match, the model has learned to denoise.
The forward step equation
Here’s how one single noise-adding step is defined:
q(xₜ | xₜ₋₁) = 𝒩(xₜ ; √(1 − βₜ) · xₜ₋₁ , βₜ I)
Let's take this apart piece by piece -
1) xₜ₋₁ — the (slightly less noisy) image from the previous step.
2) xₜ — the new, slightly noisier image we're producing at this step.
3) t — which step we're on, out of T total steps.
4) βₜ (beta sub t) — a small number, usually somewhere between 0.0001 and 0.02,
controlling how much noise gets added at step t. It's called a "noise
schedule" — researchers typically make βₜ start tiny and grow slightly
larger as t increases.
5) 𝒩( ... ) — this means "a Gaussian (Normal) distribution."
Whatever's inside describes its mean and its spread.
6) I — the identity matrix. In plain terms: "treat every pixel/dimension as
getting independent noise, with no funny correlations between them."
It just means the noise variance is applied uniformly and independently
across every value.
7) √(1 − βₜ) — this is the interesting part. Before adding noise, we shrink
the previous image slightly by multiplying it by this factor (which is just
under 1). Why shrink it at all? Because if we only ever added noise
without ever scaling the signal down, the numbers would keep growing
larger and larger, exploding to infinity after a thousand steps.
By slightly shrinking the image every time we add a bit of noise,
the total "energy" (variance) of xₜ stays controlled and
bounded — the picture fades toward pure noise smoothly, rather than blowing up.
A tiny numerical example
Let’s use one single pixel value instead of a whole image, so we can actually watch the numbers move.
Say our starting pixel value is:
x_0 = 0.8
Let’s pick β_1 = 0.1 for the first step. Then:
sqrt(1 - β_1) = sqrt(0.9) ≈ 0.949
Now suppose the random Gaussian noise we sample happens to be ε = 0.3 (just a random draw). The new noisy value becomes:
x_1 = 0.949 * 0.8 + sqrt(0.1) * 0.3
= 0.759 + 0.316 * 0.3
= 0.759 + 0.095
= 0.854
Notice: x_1 isn’t wildly different from x_0 — it’s a gentle nudge, shrunk slightly and pushed slightly by randomness. That’s the whole point of using small β values: each individual step barely changes anything. It’s the accumulation of a thousand tiny nudges that eventually turns 0.8 into pure noise.
Repeat this process again and again:
x_0 → x_1 → x_2 → x_3 → ... → x_T
And by the time you reach x_T (say, T = 1000), the accumulated shrinking and noise-adding has erased essentially all trace of the original value. Mathematically, we design the schedule so that:
x_T ≈ N(0, I)
By the final step, the image has been transformed into something indistinguishable from pure random Gaussian noise — mean zero, standard Gaussian spread, no information about x_0 left at all.

The Shortcut That Makes This Whole Thing Practical
Here’s a practical problem. If you wanted to train a network using, say, x_237 (the image at timestep 237), would you really need to simulate all 237 individual noise-adding steps every single time? That would be painfully slow, especially when training on millions of images.
Fortunately, because Gaussian noise combines so cleanly (remember, that closure property from earlier), there’s a shortcut that lets us jump straight from x_0 to any x_t in a single calculation.
First, define two helper terms:
αₜ = 1 − βₜ
Then define the cumulative product of these α values up through step t:
ᾱₜ = α₁ · α₂ · α₃ · … · αₜ
This single number captures how much of the original signal survives after t total steps of shrinking.
With that, the direct-jump equation is:
xₜ = √(ᾱₜ) · x₀ + √(1 − ᾱₜ) · ε
Explanations:
DERIVING xₜ = √(ᾱₜ) · x₀ + √(1 − ᾱₜ) · ε
Step 0 - The starting point
q(xₜ | xₜ₋₁) = 𝒩(xₜ ; √(1 − βₜ) · xₜ₋₁ , βₜ I)
Define αₜ = 1 − βₜ, so:
xₜ = √(αₜ) · xₜ₋₁ + √(1 − αₜ) · εₜ₋₁, where εₜ₋₁ ~ 𝒩(0, I).
[HOW WE GET xₜ = √(αₜ) · xₜ₋₁ + √(1 − αₜ) · εₜ₋₁
— Start with the DDPM paper's definition
q(xₜ | xₜ₋₁) = 𝒩(xₜ ; √(1 − βₜ) · xₜ₋₁ , βₜ I)
Mean: μ = √(1 − βₜ) · xₜ₋₁
Variance: σ² = βₜ
— Recall the reparameterization trick
If z ~ 𝒩(μ, σ²), then z = μ + σ · ε, where ε ~ 𝒩(0, 1).
— Apply it to our distribution
μ = √(1 − βₜ) · xₜ₋₁
σ = √(βₜ)
So:
xₜ = √(1 − βₜ) · xₜ₋₁ + √(βₜ) · εₜ₋₁
— Replace βₜ using αₜ
Define αₜ = 1 − βₜ, which means βₜ = 1 − αₜ.
So:
√(1 − βₜ) = √(αₜ)
√(βₜ) = √(1 − αₜ)
— The final form
xₜ = √(αₜ) · xₜ₋₁ + √(1 − αₜ) · εₜ₋₁
Why this structure:
- √(αₜ) · xₜ₋₁ shrinks the previous signal slightly
- √(1 − αₜ) · εₜ₋₁ adds the right amount of noise
- Together they keep the total variance of xₜ at 1 across all steps]
Step 1 - Write out the first two steps
x₁ = √(α₁) · x₀ + √(1 − α₁) · ε₀
x₂ = √(α₂) · x₁ + √(1 − α₂) · ε₁
Step 2 - Substitute x₁ into the x₂ equation
x₂ = √(α₂) · [ √(α₁) · x₀ + √(1 − α₁) · ε₀ ] + √(1 − α₂) · ε₁
x₂ = √(α₁ · α₂) · x₀ + √(α₂) · √(1 − α₁) · ε₀ + √(1 − α₂) · ε₁
Step 3 - Combine the two noise terms
Variance of first noise: α₂ · (1 − α₁)
Variance of second noise: 1 − α₂
Sum: α₂ · (1 − α₁) + (1 − α₂) = 1 − α₁ · α₂
So both noises collapse into one:
√(α₂) · √(1 − α₁) · ε₀ + √(1 − α₂) · ε₁ = √(1 − α₁ · α₂) · ε̄
Step 4 - Rewrite x₂ in clean form
x₂ = √(α₁ · α₂) · x₀ + √(1 − α₁ · α₂) · ε̄
Step 5 - Generalize to any step t
Define ᾱₜ = α₁ · α₂ · α₃ · … · αₜ
Then:
xₜ = √(ᾱₜ) · x₀ + √(1 − ᾱₜ) · ε
where ε ~ 𝒩(0, I).
Each forward step shrinks the signal by sqrt(α_t) and folds in a bit of new noise. If you chain t of these shrink-and-add steps together, and you carefully track how Gaussian noise combines (variances add when you combine independent Gaussians), the shrinking factors multiply into ᾱ_t, and the accumulated noise contribution collapses into a single clean Gaussian term with variance (1 − ᾱ_t). That’s the payoff of choosing Gaussian noise in the first place — a thousand tiny steps compress into one equation.
Numerical example of the shortcut
Suppose after 3 steps, ᾱ_3 works out to 0.7 (just a made-up value for illustration). Using our original x_0 = 0.8, and drawing a fresh noise sample ε = 0.5:
x_3 = sqrt(0.7) * 0.8 + sqrt(1 - 0.7) * 0.5
= 0.837 * 0.8 + 0.548 * 0.5
= 0.669 + 0.274
= 0.943
One line of arithmetic, and we’ve jumped directly to the noise level of step 3 — no need to simulate steps 1 and 2 individually. This is exactly what makes training diffusion models computationally feasible.

The Big Question: Can You Un-Ring a Bell?
We now know exactly how to destroy an image. But destruction was the easy direction — noise genuinely does erase information, and that’s an intentionally one-way street.
So here’s the moment where the real cleverness comes in:
“We can’t simply run the forward equation backward — because to do that, we’d need to know exactly which specific noise was added, and once it’s mixed in, that information looks lost.”
Instead of trying to invert an equation directly, researchers reframed the problem: train a neural network to guess the noise.
Here’s the setup. At training time, the network is shown:
- A noisy image, x_t
- The timestep, t (so it knows roughly how noisy the image currently is)
And it’s asked to predict:
ε_θ(xₜ, t)
“The noise that the network (with parameters θ) predicts, given the noisy image x_t and timestep t.”
Show the network a staticky image and tell it “this is how many noise-adding steps deep we are.” The network’s job is to say, “I think this specific pattern of noise is what got added.” If the network can reliably guess the added noise, then subtracting that guessed noise from x_t gets you closer to the original image.
This reframing is the single most important idea in the entire field. Instead of asking the network to directly produce a clean image in one shot (which is an enormously hard leap), we ask it to solve a much narrower, more learnable problem: “what noise is hiding in this image?”

Inside the Training Loop: How the Network Learns to Spot Noise
This is the part that actually matters — everything before it (the forward process, the shrinking, the noise schedule) exists purely to make this section possible. Let’s go through the full derivation exactly the way the DDPM paper builds it — maximum likelihood → ELBO → KL divergence → mean → noise — one step at a time, with nothing skipped.
Step 1: What are we actually trying to maximize?
Every generative model starts from the same wish given a real image x₀, we want our model to assign it a high probability of having been generated. Formally, our model defines a distribution p_θ(x₀), and we want:
max_θ log p_θ(x_0)
We want to model the true data distribution p_data(x₀), and the way we do that is by tuning our network’s parameters θ so that p_θ(x₀) gets as large as possible for real images.
The trouble is, p_θ(x₀) isn’t something we can just compute. Because our model generates an image through a whole chain of intermediate noisy steps (x₁, x₂, …, x_T), computing the probability of x₀ alone means summing — technically, integrating — over every possible path through that entire chain:
log p_θ(x_0) = log ∫ p_θ(x_0:T) dx_1:T
Here x_0:T means “the whole sequence x₀, x₁, …, x_T,” and this integral is asking: “add up the probability of every possible noisy trajectory that could have ended at x₀.” That’s an impossible sum to compute directly — there are infinitely many such paths. So direct maximum likelihood is a dead end, and we need a workaround.
Step 2: The workaround — building a lower bound (ELBO)
Instead of computing log p_θ(x₀) exactly, we build something we can compute — a guaranteed lower bound on it. If we maximize that lower bound, we’re guaranteed to be pushing the real quantity up too.
We start by multiplying inside the integral by q(x₁:T | x₀) / q(x₁:T | x₀) — which is just multiplying by 1, so nothing changes — but it lets us rewrite the expression as an expectation:
log p_θ(x_0) = log E_(q(x_1:T | x_0)) [ p_θ(x_0:T) / q(x_1:T | x_0) ]
q here is the forward (noise-adding) process we already fully know and defined earlier — it’s fixed, not learned. Now we apply a mathematical tool called Jensen’s inequality, which (for this kind of setup) tells us that “log of an average” is always greater than or equal to “average of the logs.” Applying that gives:
log p_θ(x_0) ≥ E_(q(x_1:T | x_0)) [ log ( p_θ(x_0:T) / q(x_1:T | x_0) ) ]
The right-hand side is called the Evidence Lower Bound, or ELBO:
ELBO = E_(q(x_1:T | x_0)) [ log ( p_θ(x_0:T) / q(x_1:T | x_0) ) ]
This is the quantity we can actually work with. Since ELBO ≤ log p_θ(x₀) always, pushing ELBO up during training pushes the true likelihood up too — that’s the whole trick.
Step 3: Turning “maximize ELBO” into “minimize a loss”
Training neural networks is normally framed as minimizing a loss, not maximizing a bound — so we just flip the sign:
max_θ ELBO ⟺ min_θ (−ELBO)
min_θ (−ELBO) = min_θ E_(q(x_1:T | x_0)) [ −log ( p_θ(x_0:T) / q(x_1:T | x_0) ) ]
This is now a loss function ℒ that we want to make as small as possible. Following the derivation from the DDPM paper’s Appendix A, this single expression can be carefully expanded — splitting the whole chain x₁…x_T apart timestep by timestep — into a sum of separate, more manageable KL-divergence terms:
ℒ = E_q [
D_KL( q(x_T | x_0) || p(x_T) )
+ Σ_(t=2)^(T) D_KL( q(x_(t-1) | x_t, x_0) || p_θ(x_(t-1) | x_t) )
− log p_θ(x_0 | x_1)
]
Each piece here has a name and a job:
- The first term — “matching the prior” — checks that our final noised-out distribution q(x_T | x₀) actually matches the pure-noise distribution p(x_T) = N(0, I) we start generation from. Since our noise schedule is designed so x_T really is nearly pure Gaussian noise, this term is essentially fixed and contributes almost nothing to training.
- The middle sum — “the denoising term” — is the heart of everything. For every timestep t from 2 to T, it compares the true reverse step (computed by secretly peeking at x₀) against the model’s reverse step (which never gets to see x₀). This is where nearly all the actual learning signal comes from.
- The last term — “the reconstruction term” — handles the very last step, turning x₁ back into the final image x₀. It’s handled as a special case at the end of the chain.
Almost the entire training signal lives in that middle sum, so that’s what we unpack next.
Step 4: What are we comparing? KL divergence between two Gaussians
Each term in that middle sum is a KL divergence between two Gaussian distributions — q(x_(t-1) | x_t, x₀), the true reverse step, and p_θ(x_(t-1) | x_t), the model’s reverse step. To actually compute a KL divergence between two Gaussians, there’s a standard formula:
KL( N(μ_1, Σ_1) || N(μ_2, Σ_2) )
= (1/2) [ tr(Σ_2^(-1) Σ_1) + (μ_2 − μ_1)^T Σ_2^(-1) (μ_2 − μ_1) − k + ln( det Σ_2 / det Σ_1 ) ]
This looks intimidating, but here’s the intuition: this formula measures “how different are two bell curves” by combining (a) how different their spreads are, and (b) how far apart their centers (means) are. k here is just the dimensionality of the distribution — how many numbers make up each data point.
Now here’s where things simplify enormously. In our diffusion setting, both distributions being compared are defined like this:
q(x_(t-1) | x_t, x_0) = N( μ̃_t, β̃_t I )
p_θ(x_(t-1) | x_t) = N( μ_θ(x_t, t), σ_t^2 I )
Both share the same shape of covariance — a constant times the identity matrix — and critically, the DDPM paper fixes σ_t² to a known constant rather than learning it (that’s a design choice they made to simplify and stabilize training). Because both covariances are simple known constants, almost every term in the general KL formula above cancels out or becomes a constant we don’t need to differentiate through. All that survives is the part depending on the difference between the two means:
KL = (1/2) [ (β̃_t / σ_t^2) · k + (1/σ_t^2) ||μ_θ(x_t, t) − μ̃_t||^2 − k + k ln(σ_t^2 / β̃_t) ]
Dropping everything that’s just a fixed constant (doesn’t depend on θ, so it doesn’t affect training), we’re left with a remarkably clean objective:
L_(t-1) = E_(q(x_t, x_0)) [ (1 / (2 σ_t^2)) · ||μ̃_t(x_t, x_0) − μ_θ(x_t, t)||^2 ] + C
In plain English: minimizing the KL divergence at each step reduces to simply making the model’s predicted mean μ_θ match the true mean μ̃_t as closely as possible, measured by squared distance. Everything about probability distributions and divergences has boiled down to ordinary squared error between two numbers (or vectors).
Step 5: What is the true mean, μ̃_t?
We still need to know what μ̃_t actually is. This comes directly from the forward process — since q is a fixed process we designed ourselves, we can derive its reverse conditional exactly :
q(x_(t-1) | x_t, x_0) = N( x_(t-1) ; μ̃_t(x_t, x_0), β̃_t I )
μ̃_t(x_t, x_0) = ( sqrt(α_t) (1 − ᾱ_(t-1)) / (1 − ᾱ_t) ) x_t
+ ( sqrt(ᾱ_(t-1)) β_t / (1 − ᾱ_t) ) x_0β̃_t = ( (1 − ᾱ_(t-1)) / (1 − ᾱ_t) ) β_t
Notice this needs x₀ — the original clean image — plugged directly into it. That’s fine during training (we have x₀, since it’s the real training image), but it’s a problem for generation, when x₀ is exactly the unknown thing we’re trying to produce. So this formula, as written, can’t be used by the network at generation time. We need to rewrite it using only things the network actually has access to.
Step 6: Swapping x₀ for the noise ε
Recall the forward-process shortcut equation:
x_t = sqrt(ᾱ_t) x_0 + sqrt(1 − ᾱ_t) ε
Rearranging this to solve for x₀:
x_0 = (1 / sqrt(ᾱ_t)) ( x_t − sqrt(1 − ᾱ_t) ε )
Now substitute this expression for x₀ into the μ̃_t formula from Step 5. It looks messy at first, but a fair amount of algebra cancels cleanly (because of how α_t and ᾱ_t relate to each other), and it collapses down to:
μ̃_t(x_t, ε) = (1 / sqrt(α_t)) ( x_t − (β_t / sqrt(1 − ᾱ_t)) ε )
This is a genuinely important moment in the derivation: the true mean can now be written using only x_t (which the network has) and ε (the specific noise that was mixed in) — no x₀ required anywhere.
Step 7: Defining the model’s mean the same way
Since the true mean can be expressed purely in terms of ε, the natural move is to have the network predict ε directly — call this prediction ε_θ(x_t, t) — and then build the model’s mean using the exact same formula, just with the predicted noise swapped in for the true noise:
μ_θ(x_t, t) = (1 / sqrt(α_t)) ( x_t − (β_t / sqrt(1 − ᾱ_t)) ε_θ(x_t, t) )
This is exactly the reverse-process mean formula introduced earlier in this article — and now you can see precisely where it comes from. It isn’t an arbitrary design choice; it’s the true posterior mean, algebraically rewritten so that the one unknown piece (x₀, or equivalently ε) is something the network is directly trained to guess.
Step 8: Plugging back in — the loss becomes pure noise-comparison
Now substitute both μ̃_t and μ_θ — written in their ε-based forms from Steps 6 and 7 — back into the L_(t-1) loss from Step 4:
L_(t-1) = E_(x_0, ε) [ (1/(2σ_t^2)) ||μ̃_t(x_t, x_0) − μ_θ(x_t, t)||^2 ] + C
Because both μ̃_t and μ_θ share the exact same outer structure — (1/sqrt(α_t)) times (x_t minus a scaled noise term) — nearly everything cancels except the gap between the true noise and the predicted noise. What’s left, after the algebra settles, is:
L_(t-1) = E_(x_0, ε) [ β_t^2 / (2 σ_t^2 α_t (1 − ᾱ_t)) · ||ε − ε_θ(x_t, t)||^2 ]
This is a fully valid, mathematically justified training loss — but notice that clunky weighting coefficient out front, β_t² / (2σ_t²α_t(1−ᾱ_t)). It changes at every timestep t, which means early steps (small t, close to a clean image) and late steps (large t, close to pure noise) get weighted very differently.
Step 9: The simplification that makes DDPM work so well
Here’s the paper’s key empirical finding: if you simply drop that entire weighting coefficient and train with plain, unweighted squared error instead, the model trains just as well — often better, because the model is no longer implicitly told to care less about certain noise levels. That gives the final, simplified objective used throughout the paper:
L_simple(θ) = E_(t, x_0, ε) [ ||ε − ε_θ(x_t, t)||^2 ]
where, as always:
x_t = sqrt(ᾱ_t) x_0 + sqrt(1 − ᾱ_t) ε, ε ~ N(0, I), t ~ Uniform({1, ..., T})
Every equation from Step 1 through Step 9 was really just one long process of reshaping an impossible quantity (log p_θ(x₀)) into something progressively easier to compute — until it became nothing more than “predict the noise, measure squared error.”
Step 10: The training algorithm, exactly as in the paper
Putting the whole derivation into a runnable loop (this matches “Algorithm 1” in the DDPM paper):
repeat:
x_0 ~ sample a real image from the training data
t ~ Uniform({1, 2, ..., T})
ε ~ N(0, I)
x_t = sqrt(ᾱ_t) * x_0 + sqrt(1 - ᾱ_t) * ε
loss = || ε - ε_θ(x_t, t) ||^2
θ ← update network weights via gradient descent on loss
until converged
All equations in one go
INSIDE THE TRAINING LOOP: HOW THE NETWORK LEARNS TO SPOT NOISE
Step 1: What are we actually trying to maximize?
max_θ log p_θ(x₀)
log p_θ(x₀) = log ∫ p_θ(x₀:T) dx₁:T
Step 2: The workaround — building a lower bound (ELBO)
log p_θ(x₀) = log E_(q(x₁:T | x₀))[ p_θ(x₀:T) / q(x₁:T | x₀) ]
log p_θ(x₀) ≥ E_(q(x₁:T | x₀))[ log ( p_θ(x₀:T) / q(x₁:T | x₀) ) ]
ELBO = E_(q(x₁:T | x₀))[ log ( p_θ(x₀:T) / q(x₁:T | x₀) ) ]
Step 3: Turning "maximize ELBO" into "minimize a loss"
max_θ ELBO ⟺ min_θ (−ELBO)
min_θ (−ELBO) = min_θ E_(q(x₁:T | x₀))[ −log ( p_θ(x₀:T) / q(x₁:T | x₀) ) ]
ℒ = E_q[
D_KL( q(x_T | x₀) ‖ p(x_T) )
+ Σ_(t=2)^(T) D_KL( q(x_(t−1) | x_t, x₀) ‖ p_θ(x_(t−1) | x_t) )
− log p_θ(x₀ | x₁)
]
Step 4: What are we comparing? KL divergence between two Gaussians
KL( 𝒩(μ₁, Σ₁) ‖ 𝒩(μ₂, Σ₂) )
= (1/2)[ tr(Σ₂⁻¹ Σ₁) + (μ₂ − μ₁)ᵀ Σ₂⁻¹ (μ₂ − μ₁) − k + ln( det Σ₂ / det Σ₁ ) ]
q(x_(t−1) | x_t, x₀) = 𝒩( μ̃_t, β̃_t I )
p_θ(x_(t−1) | x_t) = 𝒩( μ_θ(x_t, t), σ_t² I )
KL = (1/2)[ (β̃_t / σ_t²) · k + (1/σ_t²) ‖μ_θ(x_t, t) − μ̃_t‖² − k + k ln(σ_t² / β̃_t) ]
L_(t−1) = E_(q(x_t, x₀))[ (1 / (2σ_t²)) · ‖μ̃_t(x_t, x₀) − μ_θ(x_t, t)‖² ] + C
Step 5: What is the true mean, μ̃_t?
q(x_(t−1) | x_t, x₀) = 𝒩( x_(t−1) ; μ̃_t(x_t, x₀), β̃_t I )
μ̃_t(x_t, x₀) = ( √(α_t)(1 − ᾱ_(t−1)) / (1 − ᾱ_t) ) x_t
+ ( √(ᾱ_(t−1)) β_t / (1 − ᾱ_t) ) x₀
β̃_t = ( (1 − ᾱ_(t−1)) / (1 − ᾱ_t) ) β_t
Step 6: Swapping x₀ for the noise ε
x_t = √(ᾱ_t) x₀ + √(1 − ᾱ_t) ε
x₀ = (1 / √(ᾱ_t)) ( x_t − √(1 − ᾱ_t) ε )
μ̃_t(x_t, ε) = (1 / √(α_t)) ( x_t − (β_t / √(1 − ᾱ_t)) ε )
Step 7: Defining the model's mean the same way
μ_θ(x_t, t) = (1 / √(α_t)) ( x_t − (β_t / √(1 − ᾱ_t)) ε_θ(x_t, t) )
Step 8: Plugging back in — the loss becomes pure noise-comparison
L_(t−1) = E_(x₀, ε)[ (1/(2σ_t²)) ‖μ̃_t(x_t, x₀) − μ_θ(x_t, t)‖² ] + C
L_(t−1) = E_(x₀, ε)[ β_t² / (2σ_t² α_t (1 − ᾱ_t)) · ‖ε − ε_θ(x_t, t)‖² ]
Step 9: The simplification that makes DDPM work so well
L_simple(θ) = E_(t, x₀, ε)[ ‖ε − ε_θ(x_t, t)‖² ]
x_t = √(ᾱ_t) x₀ + √(1 − ᾱ_t) ε, ε ~ 𝒩(0, I), t ~ Uniform({1, ..., T})
Step 10: The training algorithm, exactly as in the paper
repeat:
x₀ ~ sample a real image from the training data
t ~ Uniform({1, 2, ..., T})
ε ~ 𝒩(0, I)
x_t = √(ᾱ_t) * x₀ + √(1 - ᾱ_t) * ε
loss = ‖ ε - ε_θ(x_t, t) ‖²
θ ← update network weights via gradient descent on loss
until converged


Back to the Static
Return, for a moment, to that mountain lake photograph dissolving into static.
What this whole journey shows is that the dissolving wasn’t really a one-way trip into nothingness. It was a precisely characterized process — every step governed by a known, small amount of Gaussian noise. And because it was precisely characterized, a neural network could study thousands of examples of it happening and learn the pattern of the destruction well enough to guess, at any point, what noise had just been added.
Once it can guess that, subtracting it — again and again, from pure static all the way to a clean image — becomes not just possible, but strikingly reliable.
The three-line summary of everything above:
Forward Process: Image → Noise
Learning: Noisy Image → Predict the Noise
Reverse Process: Noise → Image
A simple idea — learn exactly how something was destroyed, so you can learn exactly how to build it back — turned out to be one of the deepest and most productive insights in modern generative AI. From a single equation about Gaussian noise schedules, an entire generation of image-making tools was born.
Core DDPM Paper
- Ho, J., Jain, A., & Abbeel, P. (2020). Denoising Diffusion Probabilistic Models. NeurIPS 2020. https://arxiv.org/abs/2006.11239
Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.
Published via Towards AI
Towards AI Academy
We Build Enterprise-Grade AI. We'll Teach You to Master It Too.
15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.
Start free — no commitment:
→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day
→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages
Our courses:
→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.
→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.
→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.
Note: Article content contains the views of the contributing authors and not Towards AI.