How to spot deepfakes: visual signals, voice-clone scam defenses, and what detection tools can honestly promise.
A deepfake is media where a person's likeness — face, voice, or both — has been synthesized or swapped using machine learning. The technology spans a spectrum: face swaps that transplant one person onto another's performance, voice clones that reproduce speech from seconds of audio, and fully generated personas that never existed. The through-line is that modern synthesis is good enough to fool casual viewers, which has collapsed the old assumption that seeing (or hearing) is believing.
| Signal | Where to look | Reliability |
|---|---|---|
| Edge blending | Face boundary — subtle mismatch in skin tone or blur | Moderate |
| Blink and eye behavior | Too-frequent or too-rare blinking; dead eye contact | Moderate |
| Lighting inconsistency | Face light direction vs the scene's shadows | High when visible |
| Jewelry and teeth | Earrings that fuse, teeth with impossible geometry | High when visible |
| Audio-video sync | Voice clones drift slightly out of lip sync | Moderate |
These manual checks survive because synthesis models still struggle with exactly the things humans notice subconsciously. But they are fading as models improve — the honest position is that manual detection is getting harder every generation, and provenance (verifying where media came from) is overtaking detection (guessing whether it's fake) as the reliable approach.
The most financially damaging deepfakes today are not videos but voice: cloned voices used in phone scams impersonating family members in distress, or executives authorizing transfers. The defenses are procedural, not technical: agree on a family code word, verify any unexpected money request by calling back on a known number, and treat urgency itself as the red flag — the pressure to act now is the scam's load-bearing element. These cost nothing and defeat every voice clone, no matter how good.
Automated detectors exist, from research classifiers to platform-level systems, and their honest accuracy numbers on benchmark sets are decent — against last-generation fakes, on clean uploads. Against current-generation synthesis, on compressed social video, performance drops sharply, and adversarial users actively evade known detectors. Treat any tool claiming certainty with suspicion; treat a combination of signals (visual checks, source verification, context plausibility) as the working method. For images specifically, our AI image detector page covers the metadata half of the problem.
Creators and public figures face deepfakes as a business risk: synthesized content damaging reputations, fake profiles monetizing their likeness, scraped photos feeding custom synthesis. The protective stack is monitoring (finding copies and fakes as they appear), provenance (signed originals establishing what's real), and enforcement (takedowns at hosts and platforms). That integrated protection layer for image creators is exactly what authAspect builds — detection is one input to it, not the whole answer.
The major platforms now label or downrank media their own systems flag as synthetic, and policy pressure is pushing toward mandatory provenance labeling for AI content in several jurisdictions. Two practical consequences: first, platform detection is free and runs automatically, so one cheap move for any suspicious viral video is checking whether other platforms or fact-checkers have already labeled it; second, labeling regimes remain inconsistent, so absence of a label means nothing. The regulatory direction of travel is toward verifiable provenance, but the transition years — now — are exactly when manual skepticism carries the load.
When something looks off, a fixed sequence beats ad-hoc suspicion: pause before sharing (the single highest-value action — most harm comes from amplification, not creation), check whether established sources or fact-checkers have covered it, search for the original context the clip allegedly came from, apply the visual checks above, and for anything involving money or family, verify out-of-band no matter how real it looks. The protocol's virtue is that it costs ninety seconds and it fails safe — the worst outcome of applying it is a slightly delayed share of something true.
Everything in this guide is transitional technique for a transitional era. The destination is media that carries its own verifiable history — capture signatures, edit trails, cryptographic manifests — such that the question shifts from "can I spot a fake" to "can this file prove what it is." Cameras shipping with signing hardware, platforms adopting content credentials, and regulation pushing labeling are all convergent signals. Learn today's manual checks anyway — the transition will take years, unverified media will circulate for decades, and calm skepticism is the skill that carries across both eras.
The voice-scam defenses in particular — code word, call-back rule — protect the people around you more than they protect you, because scammers target your parents and children with your voice. Share them at the next family dinner; it is a two-minute conversation with an asymmetric payoff. The rest of this guide you can keep for yourself, but that part is worth forwarding.
Skeptical media literacy, like any skill, decays without use. The lightweight way to keep it: once a month, pick one viral media item and run the full protocol on it, just for practice. You will be wrong sometimes — everyone is, including the tools — and being wrong on a practice item, in private, with the truth surfacing later, is exactly how the judgment for the item that matters gets built. Treat it like a fire drill: the value is in the rehearsal, not the alarm.