---
license: cc-by-nc-4.0
library_name: keras
pipeline_tag: video-classification
thumbnail: https://huggingface.co/mertkayacs/xdfdet/resolve/main/assets/card.png
tags:
- deepfake-detection
- explainability
- grad-cam
- efficientnet
- faceforensics
- tensorflow
---
# xdfdet: where does a deepfake detector look?
**Explainable deepfake detection from senior AI engineer Mert Kaya's research, accepted at UBMK 2026 (IEEE):** eight trained EfficientNet-B4 detectors with Grad-CAM region analysis ([configurations](#configurations)).
Eight EfficientNet-B4 video deepfake detectors from the paper *Augmentation and Cutout in Deepfake Detection: A Comparative Study of Accuracy, Calibration, and Attention* (UBMK 2026) and the MSc thesis behind it, by Mert Kaya at TED University (thesis advisor: Venera Adanova). The eight share one architecture and recipe and differ only in data augmentation and cutout.
[Code](https://github.com/mertkayacs/xdfdet) · [Kaggle](https://www.kaggle.com/models/mertilovski/xdfdet) · [Project page](https://xdfdet.mertkayacs.com) ([TR](https://xdfdet.mertkayacs.com/tr/), [DE](https://xdfdet.mertkayacs.com/de/)) · [Thesis](https://doi.org/10.5281/zenodo.18998566) · Games: [Real or AI?](https://xdfdet.mertkayacs.com/game/), [Estimate](https://xdfdet.mertkayacs.com/estimate/)
*Grad-CAM of four setups on the same real portrait (Pexels). All four call it real, but they look at different places.*
## Use
```bash
pip install git+https://github.com/mertkayacs/xdfdet
xdfdet predict video.mp4 --gradcam cam.png
```
Prints the real-probability and verdict and saves a Grad-CAM image. Default model: `aug-cutout-black`; pick another with `--model baseline` and so on.
Loading a model directly
```python
import xdfdet
model = xdfdet.load_model("aug-cutout-black")
```
Input: 12 face crops per video, 224×224 RGB, normalized with ImageNet mean and std, shape `(batch, 12, 224, 224, 3)`. Output: the probability that each frame is **real**, shape `(batch, 12, 1)`. Average over frames for the video score; below 0.5 means fake. Use the `xdfdet` package for face detection, frame sampling and normalization, because the models are sensitive to all three.
## Configurations
| File | Augmentation and cutout | AUC |
|---|---|---|
| `aug-cutout-black.keras` | standard augmentation, black-fill cutout | **0.8981** |
| `aug-cutout-random.keras` | standard augmentation, random-fill cutout | 0.8820 |
| `aug-cutout-white.keras` | standard augmentation, white-fill cutout | 0.8734 |
| `cutout-white.keras` | white-fill cutout | 0.8700 |
| `baseline.keras` | none | 0.8684 |
| `cutout-black.keras` | black-fill cutout | 0.8669 |
| `cutout-random.keras` | random-fill cutout | 0.8642 |
| `aug-standard.keras` | standard augmentation | 0.8616 |
AUC of each checkpoint's own test run (1.0 is perfect separation, 0.5 is chance). `aug-cutout-black` scores highest of the eight.
All four scores and how the numbers were measured
| File | AUC | F1 | Brier | LogLoss |
|---|---|---|---|---|
| `aug-cutout-black.keras` | **0.8981** | **0.8431** | **0.1247** | 0.4710 |
| `aug-cutout-random.keras` | 0.8820 | 0.7950 | 0.1451 | **0.4656** |
| `aug-cutout-white.keras` | 0.8734 | 0.7883 | 0.1451 | 0.4761 |
| `cutout-white.keras` | 0.8700 | 0.7703 | 0.1463 | 0.5244 |
| `baseline.keras` | 0.8684 | 0.7781 | 0.1523 | 0.4827 |
| `cutout-black.keras` | 0.8669 | 0.7911 | 0.1537 | 0.4989 |
| `cutout-random.keras` | 0.8642 | 0.7774 | 0.1526 | 0.5241 |
| `aug-standard.keras` | 0.8616 | 0.8025 | 0.1576 | 0.5719 |
Each checkpoint's own run on its 150-pair FaceForensics++ test split, taken from the original training notebook. The paper reports mean ± std over three runs per setting; see the [code README](https://github.com/mertkayacs/xdfdet#what-we-found). A ninth setting, `aug-intense`, is in the paper, but its checkpoint was lost.
## How they were trained
1,000 real FaceForensics++ videos, each paired with one manipulated copy. One EfficientNet-B4 is trained under nine setups of augmentation and cutout. Cutout blanks a face region on **fake frames only**, following the winning Deepfake Detection Challenge solution ([Seferbekov, 2020](https://github.com/selimsef/dfdc_deepfake_challenge)); real frames get a small star with the same fill, so a blank patch alone never means fake. Grad-CAM then measures which of eight face regions drive each decision.
Training details
- **Data:** 1,000 real/fake pairs, one fake per real video, with FaceSwap, Face2Face, FaceShifter and Deepfakes in rotation. Split 70/15/15.
- **Faces:** MTCNN detection with eye alignment, 32 frames per video, 12 used per sample.
- **Cutout:** an SSIM map locates where the fake frame is most similar to its real source, and a landmark polygon over that area (2 to 5 percent of the frame) is filled on the fake frames with black, white or random pixels. Real frames receive a star-shaped cutout (outer radius 8 to 16 pixels) with the same fill. Each is applied with probability 0.5.
- **Augmentation:** Albumentations with noise, blur, colour shifts, flips and small rotations, at a standard and an intense probability level.
- **Model:** EfficientNet-B4 (ImageNet init) on each frame, global average pooling, dropout 0.55 (0.25 for `aug-cutout-random`), sigmoid output with L2 1e-3. Binary cross-entropy on the frame-averaged prediction.
- **Optimization:** Adam with cosine decay (1e-3 over 1,000 steps), gradient clipping at 1.0, batches of 4 pairs, early stopping on validation loss (patience 5), at most 20 epochs.
- **Environment:** TensorFlow 2.19 and Keras 3.10 on a Colab T4 in mixed precision. The released files are float32 copies with bit-identical weights; scores can differ slightly from the float16 runs (see `conversion.json`). For the original precision on a GPU, use `xdfdet.load_model(name, mixed_precision=True)`.
## Limits
- Trained and tested only on FaceForensics++. On 398 unseen [DFDC](https://www.kaggle.com/code/mertilovski/xdfdet-on-dfdc) videos AUC drops to 0.60 to 0.66 and most fakes pass as real. Expect the same on other datasets, newer generators and heavily compressed video.
- Each setting used its own random data split, so do not re-score these checkpoints on one shared FaceForensics++ split and compare them.
- Performance across age, sex and skin tone was not measured.
- These are research models. A score is not evidence that a video is real or fake and should not be used on its own for decisions about people.
## License
Models: CC BY-NC 4.0, because FaceForensics++ allows non-commercial research and educational use only. Code: MIT.
## Citation
```bibtex
@inproceedings{kaya2026augmentation,
title = {Augmentation and Cutout in Deepfake Detection: A Comparative Study of
Accuracy, Calibration, and Attention},
author = {Kaya, Mert and Adanova, Venera},
booktitle = {11th International Conference on Computer Science and Engineering (UBMK 2026)},
year = {2026},
note = {To appear}
}
@mastersthesis{kaya2025xdfdet,
title = {Explainable Deepfake Detection Using Frame Level CNN Models:
A Comparative Study of Augmentation and Cutout Techniques},
author = {Kaya, Mert},
school = {TED University},
year = {2025},
doi = {10.5281/zenodo.18998566}
}
```
---
An [Eschatia Labs](https://eschatialabs.com) project. [Built by Mert Kaya](https://mertkayacs.com).