mertkayacs commited on
Commit
f0d4984
·
verified ·
1 Parent(s): f38a70e

Release 8 xdfdet checkpoints with model card

Browse files
.gitattributes CHANGED
@@ -33,3 +33,13 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
 
 
 
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ assets/card.png filter=lfs diff=lfs merge=lfs -text
37
+ assets/gradcam.png filter=lfs diff=lfs merge=lfs -text
38
+ aug-cutout-black.keras filter=lfs diff=lfs merge=lfs -text
39
+ aug-cutout-random.keras filter=lfs diff=lfs merge=lfs -text
40
+ aug-cutout-white.keras filter=lfs diff=lfs merge=lfs -text
41
+ aug-standard.keras filter=lfs diff=lfs merge=lfs -text
42
+ baseline.keras filter=lfs diff=lfs merge=lfs -text
43
+ cutout-black.keras filter=lfs diff=lfs merge=lfs -text
44
+ cutout-random.keras filter=lfs diff=lfs merge=lfs -text
45
+ cutout-white.keras filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,95 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: cc-by-nc-4.0
3
+ library_name: keras
4
+ pipeline_tag: video-classification
5
+ thumbnail: https://huggingface.co/mertkayacs/xdfdet/resolve/main/assets/card.png
6
+ tags:
7
+ - deepfake-detection
8
+ - explainability
9
+ - grad-cam
10
+ - efficientnet
11
+ - faceforensics
12
+ - tensorflow
13
+ ---
14
+
15
+ # xdfdet: deepfake detectors with Grad-CAM region analysis
16
+
17
+ Eight EfficientNet-B4 video classifiers from the paper **"Augmentation and Cutout in Deepfake Detection: A Comparative Study of Accuracy, Calibration, and Attention"** (UBMK 2026) and the MSc thesis behind it. Each one was trained on FaceForensics++ under a different augmentation and cutout setting, so you can compare what the preprocessing changes, down to which face regions the model looks at.
18
+
19
+ [Code](https://github.com/mertkayacs/xdfdet) · [Project site](https://xdfdet.mertkayacs.com) · [Thesis](https://doi.org/10.5281/zenodo.18998566)
20
+
21
+ <img src="assets/gradcam.png" alt="Grad-CAM of the aug-cutout-black model on a public-domain NASA portrait: activation concentrates on the left eye and the nose" width="320">
22
+
23
+ *Grad-CAM of `aug-cutout-black` on a public-domain NASA portrait. Real-probability 0.9996; strongest activation on the left eye (81%) and the nose (63%).*
24
+
25
+ ## Use
26
+
27
+ ```bash
28
+ pip install git+https://github.com/mertkayacs/xdfdet
29
+ xdfdet predict video.mp4 --gradcam cam.png # default model: aug-cutout-black
30
+ xdfdet predict video.mp4 --model baseline
31
+ ```
32
+
33
+ ```python
34
+ import xdfdet
35
+ model = xdfdet.load_model("aug-cutout-black")
36
+ ```
37
+
38
+ Input is a batch of 12 face crops of 224×224 RGB, normalized with ImageNet mean and std: shape `(batch, 12, 224, 224, 3)`. Output is the probability that each frame is **real**, shape `(batch, 12, 1)`. Average over frames for the video score; below 0.5 means fake. The `xdfdet` package handles face cropping (MTCNN), frame sampling and normalization. Use it rather than feeding raw frames, because the preprocessing has to match training exactly.
39
+
40
+ ## Models
41
+
42
+ | File | Augmentation | Cutout fill | AUC | F1 | Brier | LogLoss |
43
+ |---|---|---|---|---|---|---|
44
+ | `aug-cutout-black.keras` | standard | black | **0.8981** | **0.8431** | **0.1247** | 0.4710 |
45
+ | `aug-cutout-random.keras` | standard | random | 0.8820 | 0.7950 | 0.1451 | **0.4656** |
46
+ | `aug-cutout-white.keras` | standard | white | 0.8734 | 0.7883 | 0.1451 | 0.4761 |
47
+ | `cutout-white.keras` | flip only | white | 0.8700 | 0.7703 | 0.1463 | 0.5244 |
48
+ | `baseline.keras` | none | none | 0.8684 | 0.7781 | 0.1523 | 0.4827 |
49
+ | `cutout-black.keras` | flip only | black | 0.8669 | 0.7911 | 0.1537 | 0.4989 |
50
+ | `cutout-random.keras` | flip only | random | 0.8642 | 0.7774 | 0.1526 | 0.5241 |
51
+ | `aug-standard.keras` | standard | none | 0.8616 | 0.8025 | 0.1576 | 0.5719 |
52
+
53
+ Scores are each checkpoint's own run on its 150-pair FaceForensics++ test split, taken from the original training notebook. The paper reports mean ± std over three runs per configuration; those numbers are in the [code README](https://github.com/mertkayacs/xdfdet#results). The ninth configuration, `aug-intense`, is described in the paper, but its checkpoint was lost.
54
+
55
+ ## Training
56
+
57
+ - **Data:** 1,000 real/fake pairs from FaceForensics++, one fake per real video, with FaceSwap, Face2Face, FaceShifter and Deepfakes in rotation. Split 70/15/15 into train, validation and test.
58
+ - **Preprocessing:** MTCNN face detection with eye alignment, 32 frames per video, 12 used per sample.
59
+ - **Model:** EfficientNet-B4 (ImageNet init) applied to each frame, then global average pooling, dropout 0.55 (0.25 for `aug-cutout-random`), and a sigmoid unit with L2 1e-3.
60
+ - **Loss:** binary cross-entropy on the frame-averaged prediction.
61
+ - **Optimization:** Adam with cosine decay (1e-3 over 1,000 steps) and gradient clipping at 1.0. Batch of 4 pairs, early stopping on validation loss with patience 5, at most 20 epochs.
62
+ - **Environment:** TensorFlow 2.19 / Keras 3.10 on a Colab T4, in mixed precision. The released files are float32 copies with bit-identical weights. Scores can differ from the float16 runs, most near the 0.5 boundary; `conversion.json` lists both for two test images. For the original precision on a GPU, load with `xdfdet.load_model(name, mixed_precision=True)`. `scripts/convert_checkpoints.py` in the code repo shows how each Colab checkpoint was matched to its configuration.
63
+
64
+ ## Limits
65
+
66
+ - Trained and tested only on FaceForensics++. Expect lower accuracy on other datasets, on newer generators such as diffusion-based face swaps, and on heavily compressed video.
67
+ - Each configuration used its own random split, so a checkpoint may have seen, during training, videos that are in another configuration's test set. Do not re-score these checkpoints on a shared FaceForensics++ split and compare them.
68
+ - Performance across age, sex and skin tone was not measured.
69
+ - These are research models. A score from them is not evidence that a video is real or fake, and it should not be used on its own for decisions about people.
70
+
71
+ ## License
72
+
73
+ CC BY-NC 4.0. The models were trained on FaceForensics++, whose terms allow non-commercial research and educational use only. The code is MIT.
74
+
75
+ ## Citation
76
+
77
+ ```bibtex
78
+ @inproceedings{kaya2026augmentation,
79
+ title = {Augmentation and Cutout in Deepfake Detection: A Comparative Study of
80
+ Accuracy, Calibration, and Attention},
81
+ author = {Kaya, Mert and Adanova, Venera},
82
+ booktitle = {11th International Conference on Computer Science and Engineering (UBMK 2026)},
83
+ year = {2026},
84
+ note = {To appear}
85
+ }
86
+
87
+ @mastersthesis{kaya2025xdfdet,
88
+ title = {Explainable Deepfake Detection Using Frame Level CNN Models:
89
+ A Comparative Study of Augmentation and Cutout Techniques},
90
+ author = {Kaya, Mert},
91
+ school = {TED University},
92
+ year = {2025},
93
+ doi = {10.5281/zenodo.18998566}
94
+ }
95
+ ```
assets/card.png ADDED

Git LFS Details

  • SHA256: b2508b83f7e7f03afcaf5dc5963bbde52531e9cbb672bfd81b6c93e2460a919d
  • Pointer size: 131 Bytes
  • Size of remote file: 232 kB
assets/gradcam.png ADDED

Git LFS Details

  • SHA256: 091b95b3f437d167430f0eca9280ca6d07d33bbd340de6a5c1c03398632bd86d
  • Pointer size: 131 Bytes
  • Size of remote file: 478 kB
aug-cutout-black.keras ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:cfb4ad8a40edfc22a75e8910cc28f113b432b80c5ba50647641ad622eb125490
3
+ size 72364819
aug-cutout-random.keras ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c3b5e1a3f230f247fb0d53e8f1f840e7024166395f07f95cc5fc9795888f04e1
3
+ size 72364819
aug-cutout-white.keras ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e9d10cc5eb99d378da6d29ca3cfdf0564dc83fdc751ba413836c61b4ce4f57e5
3
+ size 72364819
aug-standard.keras ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:aa5903362048d62e4292cd2f23ff0a201f779a869b45751f3ce7b30e6aaf2251
3
+ size 72364798
baseline.keras ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a79f7560dda198c180929076214b919f85251d3915b21d333ae2e71a6ad6e9fb
3
+ size 72364786
conversion.json ADDED
@@ -0,0 +1,98 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "baseline": {
3
+ "source": "baseline_no_aug_no_cutout.h5",
4
+ "weights_identical": true,
5
+ "video_score_mixed_float16": [
6
+ 0.468505859375,
7
+ 0.493896484375
8
+ ],
9
+ "video_score_release": [
10
+ 0.6664999723434448,
11
+ 0.10909999907016754
12
+ ]
13
+ },
14
+ "aug-standard": {
15
+ "source": "justaug.h5",
16
+ "weights_identical": true,
17
+ "video_score_mixed_float16": [
18
+ 1.0,
19
+ 1.0
20
+ ],
21
+ "video_score_release": [
22
+ 1.0,
23
+ 1.0
24
+ ]
25
+ },
26
+ "cutout-random": {
27
+ "source": "noaug.h5",
28
+ "weights_identical": true,
29
+ "video_score_mixed_float16": [
30
+ 0.0,
31
+ 0.0
32
+ ],
33
+ "video_score_release": [
34
+ 0.0,
35
+ 0.0
36
+ ]
37
+ },
38
+ "cutout-black": {
39
+ "source": "noaugcb.h5",
40
+ "weights_identical": true,
41
+ "video_score_mixed_float16": [
42
+ 0.0,
43
+ 0.0011997222900390625
44
+ ],
45
+ "video_score_release": [
46
+ 0.0,
47
+ 0.007699999958276749
48
+ ]
49
+ },
50
+ "cutout-white": {
51
+ "source": "noaugcw.h5",
52
+ "weights_identical": true,
53
+ "video_score_mixed_float16": [
54
+ 1.0,
55
+ 1.0
56
+ ],
57
+ "video_score_release": [
58
+ 1.0,
59
+ 1.0
60
+ ]
61
+ },
62
+ "aug-cutout-random": {
63
+ "source": "rc12.h5",
64
+ "weights_identical": true,
65
+ "video_score_mixed_float16": [
66
+ 0.52197265625,
67
+ 0.4453125
68
+ ],
69
+ "video_score_release": [
70
+ 0.40709999203681946,
71
+ 0.45080000162124634
72
+ ]
73
+ },
74
+ "aug-cutout-black": {
75
+ "source": "bzeroc12.h5",
76
+ "weights_identical": true,
77
+ "video_score_mixed_float16": [
78
+ 0.99365234375,
79
+ 0.9990234375
80
+ ],
81
+ "video_score_release": [
82
+ 0.9980000257492065,
83
+ 0.9994999766349792
84
+ ]
85
+ },
86
+ "aug-cutout-white": {
87
+ "source": "whitec12.h5",
88
+ "weights_identical": true,
89
+ "video_score_mixed_float16": [
90
+ 1.0,
91
+ 0.9990234375
92
+ ],
93
+ "video_score_release": [
94
+ 0.9998999834060669,
95
+ 0.9979000091552734
96
+ ]
97
+ }
98
+ }
cutout-black.keras ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0d2e5cd08e681c8f9324a6f480d98549b66cffbaf710378581cb982b93d2663b
3
+ size 72364812
cutout-random.keras ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:acda23932604cc38fabb217b035617d90c4f7e79185495ad1397e37af06d5524
3
+ size 72364812
cutout-white.keras ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f26b6d96d772b53e2bc8e1abdfb19c275eb93a592cf0c52bb43381424399ba1f
3
+ size 72364812