The first test was not whether the video looked impressive. It was whether my family could watch it and still feel the people in the photographs were theirs.
I made the film for my gran's 85th birthday. The photographs covered roughly a century of family life. I wanted to clean them up, add a little movement and turn them into something we could sit down and watch together.
Animating an old photograph is easy enough. Preserving the person in it is not. A model can sharpen a face and quietly turn it into someone else. It can colourise a dress with complete confidence and no historical basis, or invent a smile that was never there. The work was worthwhile only if the people still felt recognisable.
This was a private, non-commercial family project. The animations are AI-assisted reconstructions, not records of movement or colour. I kept the original scans unchanged.
Protecting the people in the photographs
I set the rules before choosing the final tools. There would be no invented speech or lip-sync and no large camera movements that turned a photograph into a fake scene. Every shot would begin with the restored still before moving into the animation. Anything that looked impressive but changed the person's identity or mood would be cut.
I started with high-resolution 16-bit TIFF scans. The preprocessing read each scanner's embedded colour profile, converted the image into sRGB, preserved a higher-precision intermediate and applied the recorded orientation. A small face-detection pass checked each possible rotation so that sideways scans did not confuse the later models.
I then built a local restoration stack with damage masking, face restoration, colourisation, super-resolution and a test harness. Some stages worked, but the complete route did not. The local colouriser failed on the black-and-white photographs that needed it most. The intended high-end super-resolution route produced disappointing results on the hardware available, while several model weights were hidden behind gated repositories, missing checksums or dead links.
I stopped trying to rescue that architecture and used hosted FLUX 2 [pro] to restore the colour-managed scans. Most of the local cascade was bypassed in the film that shipped. That was the right decision because my gran's film mattered more than keeping the original technical plan.
The pivot did not discard the rest of the work. I kept the colour-managed preprocessing, the registry used to track people across photographs, prompts that favoured preservation over reinvention and the review passes for identity and mood. The assembly system also held each restored photograph on screen before the animation began.
The first animations moved smoothly, but that did not make them usable. Faces drifted towards different people, expressions changed and group photographs became slightly too polished to believe. A technically clean clip could still be an immediate rejection.
I ended up using Kling 3.0 for the production animation run. I kept its prompt guidance low so the model trusted the input image more than the text. The prompts became less poetic and more operational over five rounds: preserve face, hair, age, gaze, expression, and pose; no lip-sync; no new people; no face morphing; no exaggerated movement.
Kling could also centre-crop photographs with unusual aspect ratios, sometimes removing a person standing near the edge. I fixed that by padding each image into a supported ratio with a soft extension of its own edges, generating the clip and then cropping back to the recorded frame.
The finished film runs 8:34 and contains 82 shots. Reaching that cut took 147 billed Kling generations. The recorded video-generation spend was $82.32, but I did not record the hosted restoration spend in the same ledger, so there is no reliable total project cost.
The technical work made the film possible, but the family reaction was still the test that mattered. They could watch the photographs move and continue to recognise the people in them.
