Four detectors, one for each kind of video
They work the same way, but each one learned from a different kind of video. Pick the one that sounds most like yours, and the answer will be more reliable.
| Veni HQ | Veni LQ | Vidi | Vici | |
|---|---|---|---|---|
| Use it for | For video straight from a camera or phone, or a clean download. | For video sent through chat apps or downloaded from social media. | For a face that looks smooth and polished, but something feels off. | For poorly lit, grainy or rough recordings. |
| Works best with | Sharp video where you can see skin and hair detail. | Blocky or blurry video where fine detail is lost. | Convincing face swaps with no obvious blur or colour mismatch. | Video with blur, noise, odd colours or low light. |
| Not ideal for | Videos that have been shared many times and look blocky. | Sharp, high-quality video. Veni HQ sees more there. | Very low-quality or badly damaged video. | Clean, well-lit video, where the others have more to work with. |
| Learned from | FaceForensics++light compression (C23) | FaceForensics++heavy compression (C40) | Celeb-DF v2 | DeeperForensics-1.0 |
Still not sure which one?
You can't break anything by picking the wrong one. A few tips:
- If you're unsure, start with the first option, clear and good quality.
- If the video came from WhatsApp, Messenger, Facebook or TikTok, try the forwarded option.
- If the result says it's a close call, check the same video again with a second option.
How they compare with other detectors
SwinFusion++, the model of our thesis, next to well-known detectors from other research: first on each collection, then on videos from a collection it never learned from.
FaceForensics++ HQ (C23)
| Detector | AUC | Accuracy | F1 |
|---|---|---|---|
| SwinFusion++Ours | 0.9969 | 99.29% | 0.9928 |
| CNN-LSTMReproduced | 0.7876 | 71.43% | 0.7452 |
| XceptionNetReproduced | 0.7725 | 69.64% | 0.6886 |
| SFormerPublished | – | 99.9% | – |
| XceptionNetPublished | – | 98.9% | – |
| TSFF-NetPublished | – | 97.7% | – |
| CASTPublished | 1.00 | – | – |
FaceForensics++ LQ (C40)
| Detector | AUC | Accuracy | F1 |
|---|---|---|---|
| SwinFusion++Ours | 0.9937 | 95.71% | 0.9580 |
| CNN-LSTMReproduced | 0.7984 | 70.71% | 0.7515 |
| XceptionNetReproduced | 0.8109 | 74.29% | 0.7429 |
| SFormerPublished | – | 98.15% | – |
| XceptionNetPublished | – | 96.8% | – |
| TSFF-NetPublished | – | 85.1% | – |
Celeb-DF v2
| Detector | AUC | Accuracy | F1 |
|---|---|---|---|
| SwinFusion++Ours | 0.9982 | 98.67% | 0.9922 |
| XceptionNetReproduced | 0.6480 | – | – |
| SFormerPublished | – | 99.1% | – |
| XceptionNetPublished | – | 99.4% | – |
DeeperForensics-1.0
| Detector | AUC | Accuracy | F1 |
|---|---|---|---|
| SwinFusion++Ours | 0.9969 | 99.25% | 0.9925 |
| XceptionNetReproduced | 0.9978 | 46.00% | 0.5821 |
| SFormerPublished | – | 100% | – |
Learned on FF++ HQ, tested on Celeb-DF v2
| Detector | AUC | Accuracy | F1 |
|---|---|---|---|
| SwinFusion++Ours | 0.8743 | – | – |
| TSFF-NetPublished | 0.771 | – | – |
| CAST-B0Published | 0.7698 | – | – |
| CAST-B5Published | 0.7492 | – | – |
Ours and Reproduced were tested by our team on the same held-out videos, cut out and scored the same way. Published is the figure each paper reports under its own data split and settings, shown for context rather than as a ranking. AUC (ROC-AUC) is how well a detector tells real from fake, where 1.0 is perfect and 0.5 is a coin toss; Acc is accuracy; F1 balances catching fakes against false alarms.
Found the one for your video?
Choose it on the Scan a video page and add your clip. Most checks take about a minute.