Image Classification
Transformers
Safetensors
nula
computer-vision
cnn
cifar10
adversarial-robustness
stress-test
downsampling
anti-aliasing
custom_code
Instructions to use MamaPearl/nula-cifar10-robust-v0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MamaPearl/nula-cifar10-robust-v0 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-classification", model="MamaPearl/nula-cifar10-robust-v0", trust_remote_code=True) pipe("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/hub/parrots.png")# Load model directly from transformers import AutoModelForImageClassification model = AutoModelForImageClassification.from_pretrained("MamaPearl/nula-cifar10-robust-v0", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Update README.md
Browse files
README.md
CHANGED
|
@@ -50,38 +50,24 @@ The result is an image perceptually identical to the original, with a manipulate
|
|
| 50 |
[stem β s1 β s2 β s3 β head, with channel dims]
|
| 51 |
[BlurPool replaces strided conv β note this explicitly]
|
| 52 |
|
| 53 |
-
##
|
| 54 |
|
| 55 |
-
|
| 56 |
-
the model never saw test distribution during training.
|
| 57 |
|
| 58 |
-
| Perturbation | Accuracy | Drop |
|
| 59 |
-
|---|---|---|
|
| 60 |
| Clean | 91.95% | β |
|
| 61 |
| Resize Γ0.5 (bilinear) | 59.83% | β32.12% |
|
| 62 |
| Resize Γ0.25 (bilinear) | 24.82% | β67.13% |
|
| 63 |
| Decimate Γ2 | 30.03% | β61.92% |
|
| 64 |
-
| Checkerboard
|
| 65 |
-
| Checkerboard
|
| 66 |
|
| 67 |
-
|
| 68 |
|
| 69 |
-
|
| 70 |
|
| 71 |
-
|
| 72 |
-
|---|---|---|
|
| 73 |
-
| Clean | 89.42% | β2.53% |
|
| 74 |
-
| Resize Γ0.5 | 85.37% | +25.54% |
|
| 75 |
-
| Resize Γ0.25 | 71.80% | +46.98% |
|
| 76 |
-
| Decimate Γ2 | 85.02% | +54.99% |
|
| 77 |
-
| Checkerboard Ξ΅=0.03 | 89.43% | +13.96% |
|
| 78 |
-
| Checkerboard Ξ΅=0.05 | 89.39% | +44.40% |
|
| 79 |
-
|
| 80 |
-
The adversarial variant trades 2.53% clean accuracy for substantial robustness across all tested perturbations.
|
| 81 |
-
The checkerboard attack βa direct null-space exploit against stride-2 downsampling β drops from 44.99% to near-clean 89.39%.
|
| 82 |
-
## Robust training under information-destroying transformations
|
| 83 |
-
|
| 84 |
-
The robust variant of NULA was trained from scratch under a modified training distribution.
|
| 85 |
|
| 86 |
During training, image were stochastically exposed to resolution-degrading transformations such as:
|
| 87 |
|
|
@@ -92,6 +78,27 @@ During training, image were stochastically exposed to resolution-degrading trans
|
|
| 92 |
The objective becomes invariance learning.
|
| 93 |
The model is forced to reduce sensitivity to unstable fine-scale structure and form representations that survive frequency loss and sampling artifacts.
|
| 94 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 95 |
## Usage
|
| 96 |
|
| 97 |
Nula is hosted on the HuggingFace Hub and can be loaded directly via the transformers library.
|
|
@@ -120,30 +127,19 @@ print(f"Label: {model.config.id2label[predicted_class]}")
|
|
| 120 |
Input tensors should be shape (B, C, H, W).
|
| 121 |
|
| 122 |
## Architecture
|
|
|
|
|
|
|
|
|
|
| 123 |
|
| 124 |
-
|
| 125 |
-
|
| 126 |
-
|
| 127 |
-
|
| 128 |
-
|
| 129 |
-
|
| 130 |
-
|
| 131 |
-
|
| 132 |
-
\[
|
| 133 |
-
(128, 256, 512)
|
| 134 |
-
\]
|
| 135 |
-
|
| 136 |
-
with a classifier hidden dimension of:
|
| 137 |
-
|
| 138 |
-
\[
|
| 139 |
-
512
|
| 140 |
-
\]
|
| 141 |
|
| 142 |
-
|
| 143 |
-
- **Stage 1:** residual block at constant spatial resolution
|
| 144 |
-
- **Stage 2:** residual block with anti-aliased downsampling, \(128 \to 256\)
|
| 145 |
-
- **Stage 3:** residual block with anti-aliased downsampling, \(256 \to 512\)
|
| 146 |
-
- **Head:** global average pooling, linear layer, SiLU, dropout, final classifier
|
| 147 |
|
| 148 |
## Citation
|
| 149 |
|
|
|
|
| 50 |
[stem β s1 β s2 β s3 β head, with channel dims]
|
| 51 |
[BlurPool replaces strided conv β note this explicitly]
|
| 52 |
|
| 53 |
+
### FIRST EVALUATION (Base)
|
| 54 |
|
| 55 |
+
The first evaluation was trained for clean classification performance without the robust training procedure described above.
|
|
|
|
| 56 |
|
| 57 |
+
| Perturbation | Accuracy | Drop from clean |
|
| 58 |
+
|---|---:|---:|
|
| 59 |
| Clean | 91.95% | β |
|
| 60 |
| Resize Γ0.5 (bilinear) | 59.83% | β32.12% |
|
| 61 |
| Resize Γ0.25 (bilinear) | 24.82% | β67.13% |
|
| 62 |
| Decimate Γ2 | 30.03% | β61.92% |
|
| 63 |
+
| Checkerboard \( \varepsilon = 0.03 \) | 75.47% | β16.48% |
|
| 64 |
+
| Checkerboard \( \varepsilon = 0.05 \) | 44.99% | β46.96% |
|
| 65 |
|
| 66 |
+
This model achieves high clean accuracy, but its performance collapses under resolution loss and aliasing-sensitive perturbations. In other words, it is accurate but fragile.
|
| 67 |
|
| 68 |
+
### SECOND EVALUATION (Robust)
|
| 69 |
|
| 70 |
+
The second evaluation of NULA was trained from scratch under a modified training distribution.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 71 |
|
| 72 |
During training, image were stochastically exposed to resolution-degrading transformations such as:
|
| 73 |
|
|
|
|
| 78 |
The objective becomes invariance learning.
|
| 79 |
The model is forced to reduce sensitivity to unstable fine-scale structure and form representations that survive frequency loss and sampling artifacts.
|
| 80 |
|
| 81 |
+
| Perturbation | Accuracy | Change vs. baseline |
|
| 82 |
+
|---|---:|---:|
|
| 83 |
+
| Clean | 89.42% | β2.53% |
|
| 84 |
+
| Resize Γ0.5 | 85.37% | +25.54% |
|
| 85 |
+
| Resize Γ0.25 | 71.80% | +46.98% |
|
| 86 |
+
| Decimate Γ2 | 85.02% | +54.99% |
|
| 87 |
+
| Checkerboard \( \varepsilon = 0.03 \) | 89.43% | +13.96% |
|
| 88 |
+
| Checkerboard \( \varepsilon = 0.05 \) | 89.39% | +44.40% |
|
| 89 |
+
|
| 90 |
+
The robust variant trades 2.53 percentage points of clean accuracy for a substantial increase in robustness across all tested perturbations.
|
| 91 |
+
|
| 92 |
+
Most notably, under the checkerboard perturbation with \( \varepsilon = 0.05 \), performance improves from 44.99% to 89.39%, which is nearly clean-level behavior.
|
| 93 |
+
|
| 94 |
+
## Interpretation
|
| 95 |
+
|
| 96 |
+
The baseline model seemingly relies on brittle high-frequency cues. When these cues are removed, aliased, or perturbed, classification performance degrades sharply.
|
| 97 |
+
|
| 98 |
+
The robust variant instead learns representations that are much less sensitive to these fine-scale disturbances. It does not become perfectly invariant to severe information destruction, but it shifts the model toward more stable, lower-frequency, structurally meaningful features.
|
| 99 |
+
|
| 100 |
+
This is the central result of NULA: not merely high clean accuracy, but a substantially better tradeoff between accuracy and robustness under downsampling-related perturbations.
|
| 101 |
+
|
| 102 |
## Usage
|
| 103 |
|
| 104 |
Nula is hosted on the HuggingFace Hub and can be loaded directly via the transformers library.
|
|
|
|
| 127 |
Input tensors should be shape (B, C, H, W).
|
| 128 |
|
| 129 |
## Architecture
|
| 130 |
+
```
|
| 131 |
+
stem β s1 β s2 β s3 β global pool β head
|
| 132 |
+
```
|
| 133 |
|
| 134 |
+
| Component | Details |
|
| 135 |
+
|---|---|
|
| 136 |
+
| Stem | 3 β 128, Conv3Γ3, BatchNorm, SiLU |
|
| 137 |
+
| Stage 1 | 128 β 128, residual, no downsample |
|
| 138 |
+
| Stage 2 | 128 β 256, residual, BlurPool downsample |
|
| 139 |
+
| Stage 3 | 256 β 512, residual, BlurPool downsample |
|
| 140 |
+
| Head | GlobalAvgPool β Linear(512, 512) β SiLU β Dropout(0.3) β Linear(512, 10) |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 141 |
|
| 142 |
+
SE blocks applied at each stage with reduction factor 16.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 143 |
|
| 144 |
## Citation
|
| 145 |
|