From 19th-century photo composites to today's Telegram bots. The technology is new, the impulse isn't, and the ethical line has always been consent, not the medium.
Telegram pausedPutting one face in another image is older than photography itself — painters had been doing it for centuries. Photography made it mechanical. Working photographers and commercial retouchers developed a bag of tricks:
Throughout this era, the ethical line was the same as today: consent from the subject, no impersonation for fraud, no non-consensual intimate imagery. The technology changed; the line didn't.
Academic computer-vision research in the 1990s and 2000s developed the foundations of automated face-swap. Major contributions came from graphics and vision research groups — face detection (the Viola–Jones framework, 2001), face landmark estimation, and the first learned face-swap models.
Through the 2000s and early 2010s, the research was mostly in academic papers. The models worked on still images, took significant compute to run, and were not consumer-accessible. The output was a single image, not a video, and the use cases were mostly in film VFX and academic research.
2017 was the watershed year. Three things happened in quick succession:
This is also when the safety tooling started: facial-age estimation, NSFW classification, configurable local known-hash checks, and watermarks. The category learned that misuse was the first-order problem and built the defenses in response.
2020 onward: consumer face-swap tools become a real product category. Reface, FacePlay, Wombo, and a long tail of mobile apps brought the technology to a mass audience. The shape:
The current wave is the consolidation phase. The technology is mature, the safety stack is established, and the differentiators are now form factor (chat-native vs app vs web) and use case (cinematic vs meme vs marketing), not the model quality.
Three things, in order of importance:
The next phase of the history is going to be about provenance — the technical standards for marking model output as model output, so the "reasonable viewer would believe" test is settled by the file itself, not by a watermark. The C2PA standard is one direction. The category is moving.
Standin's website can create a private, visibly watermarked Action Hero preview when intake is enabled. A clean-result unlock is offered only when the website Stripe offer is configured, and its current amount is shown before purchase. Telegram is a separate native flow without external payment buttons.
The impulse is 19th-century. The computer-vision work is 1990s. The consumer era is 2017 onward.
No single inventor. The darkroom era was working photographers and retouchers. The computer-vision era was academic research groups. The consumer era was open-source releases in 2017.
Three waves: 2017 (open-source release + the 'deepfake' term), 2018–2020 (misuse cases + legislative response), 2020 onward (consumer tools + safety tooling).
The AI model is new (2017). The impulse isn't. 19th-century photo composites did the same thing with cut-and-paste darkroom work.
Open-source face-swap models, the 'deepfake' term, and the first major misuse cases. The legislative response dates from 2018 onward.