--- license: cc-by-nc-4.0 library_name: keras pipeline_tag: video-classification thumbnail: https://huggingface.co/mertkayacs/xdfdet/resolve/main/assets/card.png tags: - deepfake-detection - explainability - grad-cam - efficientnet - faceforensics - tensorflow --- # xdfdet: where does a deepfake detector look? **Explainable deepfake detection from senior AI engineer Mert Kaya's research, accepted at UBMK 2026 (IEEE):** eight trained EfficientNet-B4 detectors with Grad-CAM region analysis ([configurations](#configurations)). Eight EfficientNet-B4 video deepfake detectors from the paper *Augmentation and Cutout in Deepfake Detection: A Comparative Study of Accuracy, Calibration, and Attention* (UBMK 2026) and the MSc thesis behind it, by Mert Kaya at TED University (thesis advisor: Venera Adanova). The eight share one architecture and recipe and differ only in data augmentation and cutout. [Code](https://github.com/mertkayacs/xdfdet) · [Kaggle](https://www.kaggle.com/models/mertilovski/xdfdet) · [Project page](https://xdfdet.mertkayacs.com) ([TR](https://xdfdet.mertkayacs.com/tr/), [DE](https://xdfdet.mertkayacs.com/de/)) · [Thesis](https://doi.org/10.5281/zenodo.18998566) · Games: [Real or AI?](https://xdfdet.mertkayacs.com/game/), [Estimate](https://xdfdet.mertkayacs.com/estimate/) Grad-CAM of four setups on the same real portrait *Grad-CAM of four setups on the same real portrait (Pexels). All four call it real, but they look at different places.* ## Use ```bash pip install git+https://github.com/mertkayacs/xdfdet xdfdet predict video.mp4 --gradcam cam.png ``` Prints the real-probability and verdict and saves a Grad-CAM image. Default model: `aug-cutout-black`; pick another with `--model baseline` and so on.
Loading a model directly ```python import xdfdet model = xdfdet.load_model("aug-cutout-black") ``` Input: 12 face crops per video, 224×224 RGB, normalized with ImageNet mean and std, shape `(batch, 12, 224, 224, 3)`. Output: the probability that each frame is **real**, shape `(batch, 12, 1)`. Average over frames for the video score; below 0.5 means fake. Use the `xdfdet` package for face detection, frame sampling and normalization, because the models are sensitive to all three.
## Configurations | File | Augmentation and cutout | AUC | |---|---|---| | `aug-cutout-black.keras` | standard augmentation, black-fill cutout | **0.8981** | | `aug-cutout-random.keras` | standard augmentation, random-fill cutout | 0.8820 | | `aug-cutout-white.keras` | standard augmentation, white-fill cutout | 0.8734 | | `cutout-white.keras` | white-fill cutout | 0.8700 | | `baseline.keras` | none | 0.8684 | | `cutout-black.keras` | black-fill cutout | 0.8669 | | `cutout-random.keras` | random-fill cutout | 0.8642 | | `aug-standard.keras` | standard augmentation | 0.8616 | AUC of each checkpoint's own test run (1.0 is perfect separation, 0.5 is chance). `aug-cutout-black` scores highest of the eight.
All four scores and how the numbers were measured | File | AUC | F1 | Brier | LogLoss | |---|---|---|---|---| | `aug-cutout-black.keras` | **0.8981** | **0.8431** | **0.1247** | 0.4710 | | `aug-cutout-random.keras` | 0.8820 | 0.7950 | 0.1451 | **0.4656** | | `aug-cutout-white.keras` | 0.8734 | 0.7883 | 0.1451 | 0.4761 | | `cutout-white.keras` | 0.8700 | 0.7703 | 0.1463 | 0.5244 | | `baseline.keras` | 0.8684 | 0.7781 | 0.1523 | 0.4827 | | `cutout-black.keras` | 0.8669 | 0.7911 | 0.1537 | 0.4989 | | `cutout-random.keras` | 0.8642 | 0.7774 | 0.1526 | 0.5241 | | `aug-standard.keras` | 0.8616 | 0.8025 | 0.1576 | 0.5719 | Each checkpoint's own run on its 150-pair FaceForensics++ test split, taken from the original training notebook. The paper reports mean ± std over three runs per setting; see the [code README](https://github.com/mertkayacs/xdfdet#what-we-found). A ninth setting, `aug-intense`, is in the paper, but its checkpoint was lost.
## How they were trained Aligned face crop, cutout on a fake frame, star on a real frame, eight face regions 1,000 real FaceForensics++ videos, each paired with one manipulated copy. One EfficientNet-B4 is trained under nine setups of augmentation and cutout. Cutout blanks a face region on **fake frames only**, following the winning Deepfake Detection Challenge solution ([Seferbekov, 2020](https://github.com/selimsef/dfdc_deepfake_challenge)); real frames get a small star with the same fill, so a blank patch alone never means fake. Grad-CAM then measures which of eight face regions drive each decision.
Training details - **Data:** 1,000 real/fake pairs, one fake per real video, with FaceSwap, Face2Face, FaceShifter and Deepfakes in rotation. Split 70/15/15. - **Faces:** MTCNN detection with eye alignment, 32 frames per video, 12 used per sample. - **Cutout:** an SSIM map locates where the fake frame is most similar to its real source, and a landmark polygon over that area (2 to 5 percent of the frame) is filled on the fake frames with black, white or random pixels. Real frames receive a star-shaped cutout (outer radius 8 to 16 pixels) with the same fill. Each is applied with probability 0.5. - **Augmentation:** Albumentations with noise, blur, colour shifts, flips and small rotations, at a standard and an intense probability level. - **Model:** EfficientNet-B4 (ImageNet init) on each frame, global average pooling, dropout 0.55 (0.25 for `aug-cutout-random`), sigmoid output with L2 1e-3. Binary cross-entropy on the frame-averaged prediction. - **Optimization:** Adam with cosine decay (1e-3 over 1,000 steps), gradient clipping at 1.0, batches of 4 pairs, early stopping on validation loss (patience 5), at most 20 epochs. - **Environment:** TensorFlow 2.19 and Keras 3.10 on a Colab T4 in mixed precision. The released files are float32 copies with bit-identical weights; scores can differ slightly from the float16 runs (see `conversion.json`). For the original precision on a GPU, use `xdfdet.load_model(name, mixed_precision=True)`.
## Limits - Trained and tested only on FaceForensics++. On 398 unseen [DFDC](https://www.kaggle.com/code/mertilovski/xdfdet-on-dfdc) videos AUC drops to 0.60 to 0.66 and most fakes pass as real. Expect the same on other datasets, newer generators and heavily compressed video. - Each setting used its own random data split, so do not re-score these checkpoints on one shared FaceForensics++ split and compare them. - Performance across age, sex and skin tone was not measured. - These are research models. A score is not evidence that a video is real or fake and should not be used on its own for decisions about people. ## License Models: CC BY-NC 4.0, because FaceForensics++ allows non-commercial research and educational use only. Code: MIT. ## Citation ```bibtex @inproceedings{kaya2026augmentation, title = {Augmentation and Cutout in Deepfake Detection: A Comparative Study of Accuracy, Calibration, and Attention}, author = {Kaya, Mert and Adanova, Venera}, booktitle = {11th International Conference on Computer Science and Engineering (UBMK 2026)}, year = {2026}, note = {To appear} } @mastersthesis{kaya2025xdfdet, title = {Explainable Deepfake Detection Using Frame Level CNN Models: A Comparative Study of Augmentation and Cutout Techniques}, author = {Kaya, Mert}, school = {TED University}, year = {2025}, doi = {10.5281/zenodo.18998566} } ``` --- Eschatia Labs An [Eschatia Labs](https://eschatialabs.com) project. [Built by Mert Kaya](https://mertkayacs.com).