ananthu-aniraj commited on
Commit
a1c197f
·
verified ·
1 Parent(s): 9046a41

Upload folder using huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +4 -4
README.md CHANGED
@@ -12,17 +12,17 @@ arxiv: 2506.08915
12
  library_name: generic
13
  ---
14
 
15
- # IFAM (metashift-k8) Model Checkpoint
16
 
17
- This is the official pre-trained checkpoint of the **IFAM (Iterative Focus and Attention Masking)** framework, proposed in the paper **"Two-stage Vision Transformers and Hard Masking offer Robust Object Representations"** (accepted as an oral presentation at ICPR 2026).
18
 
19
  - **Paper:** [Two-stage Vision Transformers and Hard Masking offer Robust Object Representations](https://arxiv.org/abs/2506.08915)
20
  - **Repository:** [GitHub - ananthu-aniraj/ifam](https://github.com/ananthu-aniraj/ifam)
21
 
22
  ## Model Description
23
- IFAM model trained on the Metashifts dataset with 8 parts (K=8).
24
 
25
- The IFAM framework is a two-stage approach:
26
  1. **Stage 1 (Selector):** Processes the full image to discover object parts and identify task-relevant regions.
27
  2. **Stage 2 (Predictor):** Restricts its receptive field to the selected regions using input attention masking, preventing spurious background details from affecting the classification.
28
 
 
12
  library_name: generic
13
  ---
14
 
15
+ # iFAM (metashift-k8) Model Checkpoint
16
 
17
+ This is the official pre-trained checkpoint of the **iFAM (Inherently Faithful Attention Maps for Vision Transformers)** framework, proposed in the paper **"Two-stage Vision Transformers and Hard Masking offer Robust Object Representations"** (accepted as an oral presentation at ICPR 2026).
18
 
19
  - **Paper:** [Two-stage Vision Transformers and Hard Masking offer Robust Object Representations](https://arxiv.org/abs/2506.08915)
20
  - **Repository:** [GitHub - ananthu-aniraj/ifam](https://github.com/ananthu-aniraj/ifam)
21
 
22
  ## Model Description
23
+ iFAM model trained on the Metashifts dataset with 8 parts (K=8).
24
 
25
+ The iFAM framework is a two-stage approach:
26
  1. **Stage 1 (Selector):** Processes the full image to discover object parts and identify task-relevant regions.
27
  2. **Stage 2 (Predictor):** Restricts its receptive field to the selected regions using input attention masking, preventing spurious background details from affecting the classification.
28