After nine failed dialogue shots, the author changed one thing: lock the audio duration first, then generate the AI video around the performance.