Put the person from a still into a driving video. No pose estimation, no segmentation, no masks — the only preprocessing is repainting the clip's first frame. 4 sampling steps. Takes 4–10 s of driving video; anything longer is cut to 10.1 s. If the clip carries a soundtrack it is pinned into the render and the mouth tracks it; upload a silent clip and the mouth moves without saying anything.
No sign-in. The first-frame repaint is paid by this Space — 5 a day per visitor. The render runs on your own ZeroGPU quota and asks for more than a free day holds, so in practice PRO and up.
examples — click one for its driving clip, its painted first frame and the render
driving video — motion, framing and background are kept; its soundtrack, if it has one, is what the mouth follows
…or frame 0, already painted (skips gpt-image-2)
step 2: render
What it produces
Every example above, rendered at 4 steps with seed 42. Each is a single paint and a single render; the first and the last three swap in a person. Only the first one has audio — it is the only one whose driving clip had any.