Feature Extraction
Transformers
Safetensors
PerceptionEncoder
English
pe_audio_video
audio-text-retrieval
image-text-retrieval
zero-shot-audio-classification
multimodal-embeddings
Eval Results (legacy)
Instructions to use MC7ever/pe-single-L14 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MC7ever/pe-single-L14 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="MC7ever/pe-single-L14")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModel processor = AutoProcessor.from_pretrained("MC7ever/pe-single-L14") model = AutoModel.from_pretrained("MC7ever/pe-single-L14", device_map="auto") - PerceptionEncoder
How to use MC7ever/pe-single-L14 with PerceptionEncoder:
# Use any PE model as a vision encoder import core.vision_encoder.pe as pe model = pe.VisionTransformer.from_config("MC7ever/pe-single-L14", pretrained=True) - Notebooks
- Google Colab
- Kaggle
card: certification status, voided runs, layer-sweep negative result
Browse files
README.md
CHANGED
|
@@ -56,6 +56,19 @@ are width-1024 — cross-width tensors cannot be weight-averaged. L/14
|
|
| 56 |
(1024-wide, 24 layers, patch 14) is the scale where Core + Spatial + AV-video
|
| 57 |
share width/depth/patch, so true weight fusion is possible.
|
| 58 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 59 |
## How it was built
|
| 60 |
|
| 61 |
1. Audio tower: exact-shape suffix-mapped average of AV-audio and A-Frame-audio
|
|
|
|
| 56 |
(1024-wide, 24 layers, patch 14) is the scale where Core + Spatial + AV-video
|
| 57 |
share width/depth/patch, so true weight fusion is possible.
|
| 58 |
|
| 59 |
+
## Status (2026-10-03): this tag is the grid-search winner, still current best
|
| 60 |
+
|
| 61 |
+
- Evolution rerun in progress after two voided attempts (documented below) —
|
| 62 |
+
if it certifies a better genome, that becomes the next revision.
|
| 63 |
+
- Layer-depth probe: naive mid-network readout does NOT beat the joint output
|
| 64 |
+
(t2i R@1 ≤ 0.06 vs 0.79) — mid-network features need trained per-layer
|
| 65 |
+
pooling (fine-tune phase), not a free readout change.
|
| 66 |
+
- Voided runs: (1) blend-of-blend contamination via shared-storage state_dict;
|
| 67 |
+
(2) silent no-op `load_state_dict` on this model's dual-prefix key layout;
|
| 68 |
+
(3) joint-path audio mismatch that faked ESC collapse at high audio weight.
|
| 69 |
+
All fixed (pristine clones, `copy_` by param name, standalone-tower-only
|
| 70 |
+
audio mapping) before trusting any evolution number.
|
| 71 |
+
|
| 72 |
## How it was built
|
| 73 |
|
| 74 |
1. Audio tower: exact-shape suffix-mapped average of AV-audio and A-Frame-audio
|