Feature Extraction
Transformers
Safetensors
PerceptionEncoder
English
pe_audio_video
audio-text-retrieval
image-text-retrieval
zero-shot-audio-classification
multimodal-embeddings
Eval Results (legacy)
Instructions to use MC7ever/pe-single-L14 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MC7ever/pe-single-L14 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="MC7ever/pe-single-L14")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModel processor = AutoProcessor.from_pretrained("MC7ever/pe-single-L14") model = AutoModel.from_pretrained("MC7ever/pe-single-L14", device_map="auto") - PerceptionEncoder
How to use MC7ever/pe-single-L14 with PerceptionEncoder:
# Use any PE model as a vision encoder import core.vision_encoder.pe as pe model = pe.VisionTransformer.from_config("MC7ever/pe-single-L14", pretrained=True) - Notebooks
- Google Colab
- Kaggle
File size: 6,478 Bytes
5f5a478 6b632dd 824091a 6baebab 2aed196 6baebab 5f5a478 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 | ---
license: apache-2.0
library_name: transformers
pipeline_tag: feature-extraction
tags:
- pe_audio_video
- perception-encoder
- audio-text-retrieval
- image-text-retrieval
- zero-shot-audio-classification
- multimodal-embeddings
language:
- en
metrics:
- recall@1
- recall@5
- accuracy
model-index:
- name: pe-single-L14-a025
results:
- task:
type: text-to-image-retrieval
dataset:
name: COCO slice (150, MMEB MSCOCO_i2t rows)
metrics:
- type: recall@1
value: 0.766
name: t2i R@1 (search slice n=64)
- type: recall@5
value: 0.95
name: t2i R@5 (approx, search slice)
- task:
type: zero-shot-audio-classification
dataset:
name: ESC-50 slice (100 clips)
metrics:
- type: accuracy
value: 0.96
name: ESC-50 accuracy (search slice)
---
# pe-single-L14 — one fused Meta Perception Encoder, no router
Single-file joint embedding model (audio + image/video + text, 1024-d shared
space, 1.03B params) fusing four Meta PE family members:
| source | HF id | what was fused |
|---|---|---|
| PE-AV base (anchor) | `facebook/pe-av-base` | full joint model kept as scaffold |
| PE A-Frame base | `facebook/pe-a-frame-base` | audio tower, weight 0.25 |
| PE-Core-L14-336 | `facebook/PE-Core-L14-336` | vision trunk (near-identical to anchor video tower — init copy, ~no-op) |
| PE-Spatial-L14-448 | `facebook/PE-Spatial-L14-448` | vision trunk, weight 0.0 in this tag (hurt COCO alignment; see below) |
Why L-scale: PE-Core/Spatial-G14 are width-1536 × 50 layers while PE-AV/A-Frame
are width-1024 — cross-width tensors cannot be weight-averaged. L/14
(1024-wide, 24 layers, patch 14) is the scale where Core + Spatial + AV-video
share width/depth/patch, so true weight fusion is possible.
## Status (2026-10-04): v3 TUNED — local MPS fine-tune, new SOTA
Heads-only contrastive tune (COCO-train 4000 pairs x 2 epochs, frozen towers,
local Apple-Silicon MPS, fp32). Full-slice eval (COCO n=150, ESC-50 n=200):
| model | COCO t2i R@1 / R@5 | COCO i2t R@1 / R@5 | ESC-50 acc |
|---|---|---|---|
| `facebook/pe-av-base` (anchor) | 0.720 / 0.953 | 0.013 / 0.073 | 0.885 |
| frozen fusion (this repo, previous tag) | 0.713 / 0.953 | 0.013 / 0.047 | 0.895 |
| **v3 tuned (this revision)** | **0.827 / 0.980** | **0.727 / 0.973** | 0.895 |
The tune unlocked the text direction (i2t 0.013 -> 0.727) that no weight
blend touched, and lifted t2i +0.11. ESC unchanged (v1 trained video<->text
only; audio<->text tuning is next). Previous certification notes below remain
the record for the frozen init.
Full-slice validation (COCO n=150, ESC-50 n=200) vs anchor:
| model | COCO t2i R@1 / R@5 | COCO i2t R@1 | ESC-50 acc |
|---|---|---|---|
| `facebook/pe-av-base` (anchor) | 0.720 / 0.953 | 0.013 | 0.885 |
| this tag | 0.713 / 0.953 | 0.013 | 0.895 |
Verdict: statistical tie (diffs within SE). Weight fusion preserves anchor
quality while carrying A-Frame/Core/Spatial weights in one file — it does not
beat the anchor on these benches by itself. Measured gains are deferred to
alignment fine-tuning, for which this tag is the frozen init. Certification:
9-build grid + 18-genome evolution (fixed machinery) + full-slice validation,
all agreeing (audio light, vision ~zero).
15-metric suite (COCO n=300, ESC-50 5x120, MPS fp32) vs anchor:
vision tied (V1 t2i R@1 0.613 vs 0.617, V5 0.637 vs 0.643, V4 min 0.527 vs
0.547; V2/V3 at chance for both — text embeds are dummy-video dominated in
this harness, a comparative-only handicap); audio 4/5 folds +0.017..+0.025
(fold3 -0.017), prompt-mean +0.008; joint diagnostics mixed tiny deltas
(J5 +0.033). No regression anywhere; consistent small audio edge. Full table:
suite.json in the GitHub repo.
- Evolution rerun in progress after two voided attempts (documented below) —
if it certifies a better genome, that becomes the next revision.
- Layer-depth probe: naive mid-network readout does NOT beat the joint output
(t2i R@1 ≤ 0.06 vs 0.79) — mid-network features need trained per-layer
pooling (fine-tune phase), not a free readout change.
- Voided runs: (1) blend-of-blend contamination via shared-storage state_dict;
(2) silent no-op `load_state_dict` on this model's dual-prefix key layout;
(3) joint-path audio mismatch that faked ESC collapse at high audio weight.
All fixed (pristine clones, `copy_` by param name, standalone-tower-only
audio mapping) before trusting any evolution number.
## How it was built
1. Audio tower: exact-shape suffix-mapped average of AV-audio and A-Frame-audio
(both `pe_audio_encoder`, 1024/16L/8H — 422/422 tensors matched).
2. Video trunk: explicit native-PE → timm key map (patch_embed, cls_token,
pos_embed@336px, norms, QKV, MLP ×24 blocks). Pool/head/proj kept from anchor.
3. Blend weights chosen by measured 3×3 grid search over COCO-retrieval +
ESC-50 slices (script `search_pe_weights.py`), not by vibes. Vision donor
weight degraded COCO t2i monotonically (0.766 → 0.64), so the winning tag
keeps anchor vision and blends audio at 0.25.
## Usage
```python
from transformers import AutoModel, AutoProcessor
m = AutoModel.from_pretrained("MC7ever/pe-single-L14", trust_remote_code=True)
p = AutoProcessor.from_pretrained("MC7ever/pe-single-L14", trust_remote_code=True)
inputs = p(videos=[frames], audio=[waveform], text=["a cat playing in rain"],
return_tensors="pt")
out = m(**inputs) # out.video_embeds, out.audio_embeds, out.text_video_embeds, ...
```
Retrieval: dot-product `video_embeds @ text_video_embeds.T`.
Zero-shot audio: argmax over `audio_embeds @ text_audio_embeds.T` with
`"the sound of {label}"` prompts.
## Honest limits
- Numbers above are CPU-run slices (COCO n=64, ESC n=100), comparative
anchor-vs-variant only — not official MMEB/PE-paper figures.
- i2t (text→image caption ranking over 150 similar COCO captions) sits near
floor for anchor and variants alike; t2i and audio carry the signal.
- Encoder only: no generation, detection, or segmentation heads.
- Next: evolution-based per-tower weights, intermediate-layer readout
(PE paper: best features are mid-network), GPU alignment fine-tune.
## Reproduce
`fuse_pe_single.py` (fusion) · `eval_pe_bench.py` (COCO+ESC-50 bench) ·
`search_pe_weights.py` (grid) · `evo_fuse.py` (evolution) · `layer_sweep.py`
(depth probe). Benchmark slice metadata: MMEB `MSCOCO_i2t` rows + COCO val2014
images + `ashraq/esc50` streaming clips.
|