Back to all blogs
Why detecting fraud in video interactions requires locating suspicious moments, not simply classifying an entire meeting as real or fake.

Abhishek Kaushik
Most video fraud systems are built around a classification question whether Is this video real or fake? For high-stakes video interactions, we think that framing is incomplete. A 45-minute interview can be authentic for 44 minutes and still be compromised by 30 seconds of synthetic video, generated audio, replayed content, outside assistance, or a change in the participant controlling the session. The fraud does not need to dominate the interaction. It only needs to appear at the moment that matters. This makes sparse fraud fundamentally a localization problem.
Consider a detector producing a fake probability for every frame:
0.04 → 0.06 → 0.05 → 0.81 → 0.89 → 0.77 → 0.04
If these predictions are compressed too early into one video-level score, a short anomalous segment can be diluted by thousands of legitimate frames. Recent evidence suggests this is a real limitation. FakeI2V-Bench, a KDD 2026 study evaluating deepfake detection across 97,548 videos, introduced a short-term forgery experiment in which only 20% of the frames in fake videos were manipulated. Performance deteriorated substantially across both video-native and adapted image-based detectors. The strongest method in this sparse setting reached only 77.15% AUC.
The paper does not study live meetings specifically, but the result motivates an important question for interaction fraud: What happens when the fraudulent event occupies only seconds of a much longer legitimate session?
A study conducted by Sherlock AI's team suggests that for interactive fraud, the system should not immediately collapse evidence into : meeting risk: 31%. It should preserve something closer to:
00:00–18:31 : baseline behavior
18:32–18:47 : synthetic-video signal increases
18:35 : audiovisual consistency changes
18:48 onward : signals return to baseline
The object being detected is no longer simply a “fake video”. It is an integrity-breaking event occurring at a particular point in time. That changes both the architecture and the evaluation problem. A useful fraud-intelligence system should answer:
When did the anomaly begin?
How long did it persist?
Which signals changed together?
How quickly could it have been detected?
What evidence supports the conclusion?
For live interactions, the problem becomes even harder because the detector has to make these judgments sequentially, without seeing the future. From video classification to temporal forensics. Our research hypothesis at Sherlock is that future video fraud systems will increasingly look like:
continuous observations → state changes → suspicious intervals → evidence aggregation → interaction verdict rather than: video → model → real/fake
The final verdict still matters. But it should be the consequence of localized evidence, not a replacement for it. A 60-minute interaction containing 20 seconds of fraud is still a compromised interaction. The difficult problem is finding those 20 seconds. And that may ultimately matter more than correctly classifying the other 59 minutes and 40 seconds.
Request Sherlock Research
Sherlock is actively studying sparse fraud, temporal localization, liveness, synthetic media, identity continuity, and interaction integrity in real-world video interactions.
Request access to proprietary research findings from the Sherlock research team.



