Sora's Diffusion Transformer (DiT) by Hand ✍️
Calculating AI by Hand: 26 of 28
Library › Calculating AI by Hand ✍️
Sora's Diffusion Transformer (DiT) by Hand ✍️
Reinforcement Learning with Human Feedback (RLHF) by Hand ✍️
Sora, OpenAI's text-to-video model, is built on the Diffusion Transformer (DiT), developed by William Peebles and Saining Xie in 2023.
How does DiT work?
Goal: Generate a video conditioned by a text prompt and a series of diffusion steps
Setup
Step 1 of 14: Given
Video
Prompt: "sora is sky"
Diffusion step: t = 3
Step 2 of 14: Video → Patches
Divide all pixels in all frames into 4 spacetime patches
Step 3 of 14: Visual Encoder
Multiply the patches with weights and biases, followed by ReLU
The result is a latent feature vector per patch
The purpose is dimension reduction from 4 (2x2x1) to 2 (2x1).
In the paper, the reduction is 196,608 (256x256x3)→ 4096 (32x32x4)





