Create README.md
Browse files
README.md
ADDED
|
@@ -0,0 +1,93 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: other
|
| 3 |
+
tags:
|
| 4 |
+
- vision-transformer
|
| 5 |
+
- radar
|
| 6 |
+
- signal-processing
|
| 7 |
+
- electronic-warfare
|
| 8 |
+
- modulation-recognition
|
| 9 |
+
---
|
| 10 |
+
|
| 11 |
+
# GC-ViT: AI-Based Global Context Vision Transformer for Radar Signal Modulation Recognition
|
| 12 |
+
An AI-based Global Context Vision Transformer (GC-ViT) that leverages the
|
| 13 |
+
short-time Fourier transform (STFT) phase spectrum for feature extraction to
|
| 14 |
+
identify phase-coded radar waveforms β a key capability for electronic
|
| 15 |
+
warfare (EW) systems facing the growing use of low-probability-of-intercept
|
| 16 |
+
(LPI) radars. Combines local and global self-attention to improve
|
| 17 |
+
recognition robustness at low SNR.
|
| 18 |
+
|
| 19 |
+
**Dataset:** (https://www.kaggle.com/datasets/sidrabhatti/latestdataset-cnn)
|
| 20 |
+
|
| 21 |
+
## Role & Attribution
|
| 22 |
+
Sidra Ghayour Bhatti β first author; conceived, implemented, and evaluated
|
| 23 |
+
the GC-ViT model. Co-authored with Mohsin Ullah (Dept. of Electrical
|
| 24 |
+
Engineering, Capital University of Science and Technology, Islamabad).
|
| 25 |
+
|
| 26 |
+
## Method
|
| 27 |
+
|
| 28 |
+
**Input representation.** Each intercepted radar pulse is converted to its
|
| 29 |
+
short-time Fourier transform (STFT) phase spectrum rather than the more
|
| 30 |
+
commonly used magnitude spectrum β the deliberate choice is because
|
| 31 |
+
phase-coded waveforms (Barker codes, P1βP4 polyphase codes, etc.) are
|
| 32 |
+
*defined* by their intrapulse phase-modulation pattern, not by how their
|
| 33 |
+
energy is distributed across frequency. Two different phase codes can have
|
| 34 |
+
near-identical magnitude spectra while differing sharply in phase, so a
|
| 35 |
+
magnitude-only spectrogram throws away the very information that
|
| 36 |
+
distinguishes the waveform classes; the phase spectrum keeps it. The
|
| 37 |
+
resulting 2D time-frequency phase map is cropped to its informative region
|
| 38 |
+
and resized to 224Γ224Γ3, turning a 1D IQ waveform recognition problem into
|
| 39 |
+
a 2D image classification problem that a vision architecture can exploit.
|
| 40 |
+
|
| 41 |
+
**Model.** A GC-ViT (Global Context Vision Transformer, Tiny variant)
|
| 42 |
+
backbone, pretrained then fine-tuned with a 6-class softmax classification
|
| 43 |
+
head for phase-coded waveform families. GC-ViT combines local window-based
|
| 44 |
+
self-attention (fine-grained spectrogram texture) with a global query token
|
| 45 |
+
(long-range structure across the full time-frequency map) in each block β
|
| 46 |
+
the architectural reason it outperforms CNN baselines at low SNR, where
|
| 47 |
+
discriminative structure is spread across the spectrogram rather than
|
| 48 |
+
localized.
|
| 49 |
+
|
| 50 |
+
**Training.** Fine-tuned with SGD (lr=0.001) and sparse categorical
|
| 51 |
+
cross-entropy loss on top of GC-ViT's pretrained backbone.
|
| 52 |
+
|
| 53 |
+
**Evaluation protocol.** Tested across a wide SNR sweep (roughly β14 dB to
|
| 54 |
+
+8 dB in 2 dB steps) with held-out test sets per SNR level, so performance
|
| 55 |
+
is reported as a function of noise level rather than a single aggregate
|
| 56 |
+
number β directly relevant to EW receivers that must operate across
|
| 57 |
+
unknown, often very low, intercept SNRs.
|
| 58 |
+
|
| 59 |
+
## Results
|
| 60 |
+
- ~80% recognition accuracy at β12 dB SNR, substantially outperforming prior
|
| 61 |
+
methods at that noise level
|
| 62 |
+
- Performance degrades gracefully across the full tested SNR sweep rather
|
| 63 |
+
than collapsing sharply, consistent with the global self-attention
|
| 64 |
+
mechanism recovering structure that local-only (CNN) models miss at low SNR
|
| 65 |
+
- Demonstrates robustness for EW situational-awareness applications in
|
| 66 |
+
complex electromagnetic environments
|
| 67 |
+
|
| 68 |
+
## Code
|
| 69 |
+
`code/Global_Context_ViT_for_RadarSiganls.ipynb` β the training/evaluation
|
| 70 |
+
notebook used for this paper. It's Colab-specific (mounts Google Drive and
|
| 71 |
+
loads a locally-trained `.h5` weights file not included here), so it won't
|
| 72 |
+
run standalone, but it documents the exact preprocessing, model setup, and
|
| 73 |
+
SNR-sweep evaluation protocol behind the reported results. Training data:
|
| 74 |
+
see the Kaggle dataset link above.
|
| 75 |
+
|
| 76 |
+
**Published:** Bhatti, S.G. & Ullah, M. (2024). "Radar signal modulation
|
| 77 |
+
identification using global context vision transformer." *Engineering
|
| 78 |
+
Research Express*, 6(4), 045331.
|
| 79 |
+
[doi.org/10.1088/2631-8695/ad8b96](https://doi.org/10.1088/2631-8695/ad8b96)
|
| 80 |
+
(subscription required β not open access, so no PDF is included here; link only)
|
| 81 |
+
|
| 82 |
+
## Citation
|
| 83 |
+
```bibtex
|
| 84 |
+
@article{bhatti2024radar,
|
| 85 |
+
author = {Bhatti, Sidra Ghayour and Ullah, Mohsin},
|
| 86 |
+
title = {Radar signal modulation identification using global context vision transformer},
|
| 87 |
+
journal = {Engineering Research Express},
|
| 88 |
+
volume = {6},
|
| 89 |
+
number = {4},
|
| 90 |
+
pages = {045331},
|
| 91 |
+
year = {2024},
|
| 92 |
+
doi = {10.1088/2631-8695/ad8b96}
|
| 93 |
+
}
|