Video thumbnail for LoRA Model Trained on My Face | The Results Are Insane

LoRA Model Trained on My Face | The Results Are Insane

Aug 6, 2026
I trained a character LoRA of myself for 8000 steps and saved a checkpoint every 250 — 32 versions of the same model, from barely-trained to catastrophically overcooked. Then I rendered the same prompt through every one of them with a fixed seed to find out where the sweet spot actually is. It peaked at step 1750. That was 39 hours into a 164-hour run. Everything after that made it worse, and the loss curve never once said so. Four things break, and they break in a specific order: - Identity locks in at ~1750, and the voice gets there roughly 3x faster than the face — by step 500 the voice was already 72% of the way from "not him" to "actually him" while the picture still looked like a stranger. - Motion collapses next. Frame-to-frame movement peaks at 1750, is down 71% by 2500, and averages 79% below peak across the last quarter of the run. He still looks like me. He just stops moving like a person. - Speech goes off a cliff, not a slope. Fourteen checkpoints in a row say the line perfectly. At 3750 it cracks. After that it's the right voice, the right rhythm, the right confidence — and no words. - Promptability dies last, around 4000. I described my actual room; by step 8000 it ignored every word and rendered the background from my old videos instead. It didn't learn me. It learned that photograph. Also in here: the bit where I nearly published something false. I measured the speech degradation with Whisper, and Whisper told me that at step 7750 I said a sentence that does not exist — there are no words in that audio at all. Whisper is a language model. It does not output gibberish, it outputs plausible English, always. If you use it for your subtitles it is doing that to you right now on every inaudible line. The fix is in the video. The dataset explains all of it. 28 clips, 2 minutes 48 seconds of source, all cut from one video — one background, one camera, one framing. That's 286 passes over the same material when about 63 would have done. Captioning doesn't create separability. Variance does.
#Science