Industry

    Fingerprinting the Real Instead of the Fake

    Detecting fakes is a race with no finish line. The alternative gaining ground is signing real media at capture — which flips what 'unlabeled' means.

    Versely Team8 min read

    There are two ways to solve "how do we know if this is real," and for most of the last few years the industry has poured its effort into the harder one. Detection tries to look at a finished piece of media and decide whether it was generated — a classifier trained to spot the tells of synthesis. Provenance does the opposite: it doesn't examine the output at all, it signs the input, at the moment of capture, before anything downstream happens to it. Instagram's own head of product has now said out loud what the detection numbers have been implying for a while: the second approach is the one that scales, and the first one structurally can't.

    Why detection doesn't have a finish line

    The problem with training a detector to catch synthetic media is that you're building the exact target the next generation of generative models gets optimized against. Every detectable tell a classifier learns to flag — a texture artifact, an inconsistent shadow, a statistical fingerprint in the pixels — is information about exactly what the next model needs to fix to stop being caught. That's not a stable arms race with occasional setbacks; it's an adversarial loop where the defender's every success trains the attacker's next move, and the attacker gets to try again for free, as many times as it wants, until the tell is gone. However accurate a detector measures against today's generators, that figure tells you almost nothing about its accuracy against next quarter's, because next quarter's model was very possibly trained against detectors resembling this one.

    Provenance doesn't have that problem, because it isn't trying to characterize synthetic output at all. It's a one-time infrastructure decision made at the moment a photon actually hits a sensor, and it doesn't need to keep pace with anything a generative model does later, because it was never in that race to begin with.

    The inversion: sign the real thing, not the fake one

    Adam Mosseri put this plainly: it will be more practical to fingerprint real media than fake media, because as AI becomes ubiquitous, distinguishing it from reality gets harder in exactly the direction detection can't fix — and the fix he points to is camera manufacturers cryptographically signing images at the moment of capture, creating a chain of custody from that point forward. That's the flip: instead of asking "does this file look synthetic," the question becomes "does this file carry a cryptographic signature proving it came from a real sensor at a specific time," a question that doesn't get harder to answer as generators improve, because the answer is decided at capture, once, and never has to be re-litigated against whatever model shipped after.

    A signature like that isn't just a single "verified" stamp, either — done properly, it's a chain of custody. Every edit, crop, or compression pass along the way can be logged against the original signed capture, so what a viewer eventually checks isn't just "was this real at some point" but "here's the specific history of what happened to it between the sensor and your screen."

    What's already shipping

    This isn't a proposal sitting in a whitepaper. Content Credentials — the consumer-facing implementation of this capture-time signing — now ships on Google's Pixel 10 and on Sony's PXW-Z300 professional video camera, which matters because it moves the starting point from "add provenance in editing software after the fact" to "the file is signed before a human ever touches it." A phone that a consumer already owns and a broadcast-grade camera used in professional production are two very different adoption paths converging on the same standard, which is the pattern that tends to actually stick rather than stay a demo.

    The honest catch: unsigned doesn't mean fake

    Here's the part worth being precise about, because it's the difference between a genuinely useful shift and an overclaimed one. A capture-signing system doesn't label fakes. It proves reals. Those sound like mirror images of the same thing and they are not.

    The practical consequence is that "no signature present" lands in an ambiguous middle, not in a "presumed synthetic" bucket. Every photo taken before this standard existed is unsigned. Every camera that hasn't adopted it yet produces unsigned files. A perfectly real photo that gets re-encoded by a platform, screenshotted, or passed through an app that strips metadata can lose its signature along the way even though nothing about the actual scene it depicts was ever fake. None of that is evidence of synthesis — it's just the absence of a proof that, for most of photography's history, never existed to lose. A system built around this needs to resist the temptation to treat "unsigned" as a verdict, because right now, and for a long transition period ahead, it's simply the default state of almost everything.

    What does change, gradually, is the social meaning of the gap. As signed capture becomes normal on flagship hardware, an unsigned file in a context where signing was obviously available starts to read as a choice rather than a limitation — not proof of anything, but a missing opportunity to have removed the doubt. That's a slow cultural shift, not a technical one, and it's worth not overselling how far along it already is.

    Two different jobs, not one system

    It's worth being clear that this doesn't replace the disclosure tools built for the opposite direction — marking content as synthetic rather than proving it's real — because they're solving different halves of the same overall problem, not competing for the same job.

    A watermark — visible or embedded in the pixels — is applied to generated media to identify where it came from. That's a provenance signal running in the opposite direction from capture-signing: it marks the fake, rather than proving the real. Synthetic media disclosure and a platform's AI content label sit downstream of that same idea — obligations and toggles that tell a viewer "this was AI-made or altered," which matters regardless of whether capture-signing ever reaches ubiquity, because a creator who generates a video with Versely and posts it isn't claiming it came from a camera at all. Signing real media at capture is the long-term infrastructure fix for the ambient "is anything real anymore" problem; disclosing your own synthetic content honestly is the thing you're responsible for today, on every post, independent of whether the capture side of this ever catches up.

    A Versely walkthrough: disclosing what you generated, today

    None of the capture-signing infrastructure above is something a generation tool participates in — Versely doesn't produce camera-captured footage, so it has nothing to sign. What it can do is make sure your own output is honestly marked before it leaves your hands, which is the half of this problem that's actually yours to solve:

    1. Burn a visible disclosure onto the finished video rather than relying on a platform's own toggle alone. Ask the agent to add a fixed text overlay — "AI-GENERATED" or your own phrasing, positioned wherever it won't collide with your other on-screen text — and it's burned into the frame itself, which survives a re-upload or repost in a way a platform-side metadata flag doesn't always.
    2. Still toggle the platform's own AI-content label at the point of posting — a burned-in overlay and a platform toggle answer the same question in two different places, and using both is the actual standard most platforms expect rather than a belt-and-suspenders overcorrection.
    3. Make it the default, not a per-post decision. Tell the agent once, as a lasting preference, that every generated video gets the disclosure line before it's finished — the same durable-memory mechanism that holds any other standing instruction — so it's applied without having to remember to ask for it on the next brief.

    What this means right now

    Neither piece of this technology changes what a creator or brand should actually do this week, and that's worth saying directly: don't treat "detection is losing" as a reason to loosen up on disclosure, and don't treat "capture-signing exists" as something that affects generated content at all — it's built for the opposite category. The durable move, regardless of which infrastructure eventually wins, is the one entirely within your own control: label your own synthetic media clearly, at the point you publish it, rather than waiting to see whether some future detector or some future camera-signing standard would have caught it if you hadn't. Provenance infrastructure is a multi-year industry shift. Whether your next post carries an honest AI label is a decision you make today.