From 19th-century photo composites to today's Telegram bots. The technology is new, the impulse isn't, and the ethical line has always been consent, not the medium.
Try Standin in Telegram →Putting one face in another image is older than photography itself — painters had been doing it for centuries. Photography made it mechanical. Working photographers and commercial retouchers developed a bag of tricks:
Throughout this era, the ethical line was the same as today: consent from the subject, no impersonation for fraud, no non-consensual intimate imagery. The technology changed; the line didn't.
Academic computer-vision research in the 1990s and 2000s developed the foundations of automated face-swap. Major contributions came from graphics and vision research groups — face detection (the Viola–Jones framework, 2001), face landmark estimation, and the first learned face-swap models.
Through the 2000s and early 2010s, the research was mostly in academic papers. The models worked on still images, took significant compute to run, and were not consumer-accessible. The output was a single image, not a video, and the use cases were mostly in film VFX and academic research.
2017 was the watershed year. Three things happened in quick succession:
This is also when the safety tooling started: facial-age estimation, NSFW classification, CSAM hash matching, and watermarks. The category learned that misuse was the first-order problem and built the defenses in response.
2020 onward: consumer face-swap tools become a real product category. Reface, FacePlay, Wombo, and a long tail of mobile apps brought the technology to a mass audience. The shape:
The current wave is the consolidation phase. The technology is mature, the safety stack is established, and the differentiators are now form factor (chat-native vs app vs web) and use case (cinematic vs meme vs marketing), not the model quality.
Three things, in order of importance:
The next phase of the history is going to be about provenance — the technical standards for marking model output as model output, so the "reasonable viewer would believe" test is settled by the file itself, not by a watermark. The C2PA standard is one direction. The category is moving.
Standin is the chat-native, pay-per-clip wave of the consumer era. Telegram-native (no install), pay-per-clip ($9+), one free render (watermarked), and the most publicly documented safety stack in the category. The bot is a 2025 product; the underlying model is 2017; the impulse to put your face in a different scene is older than photography.
The impulse is 19th-century. The computer-vision work is 1990s. The consumer era is 2017 onward.
No single inventor. The darkroom era was working photographers and retouchers. The computer-vision era was academic research groups. The consumer era was open-source releases in 2017.
Three waves: 2017 (open-source release + the 'deepfake' term), 2018–2020 (misuse cases + legislative response), 2020 onward (consumer tools + safety tooling).
The AI model is new (2017). The impulse isn't. 19th-century photo composites did the same thing with cut-and-paste darkroom work.
Open-source face-swap models, the 'deepfake' term, and the first major misuse cases. The legislative response dates from 2018 onward.