File size: 6,478 Bytes
5f5a478
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
6b632dd
 
 
 
 
 
 
 
 
 
 
 
 
 
 
824091a
 
 
 
 
 
 
 
 
 
 
 
 
 
6baebab
2aed196
 
 
 
 
 
 
 
6baebab
 
 
 
 
 
 
 
 
 
 
5f5a478
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
---
license: apache-2.0
library_name: transformers
pipeline_tag: feature-extraction
tags:
- pe_audio_video
- perception-encoder
- audio-text-retrieval
- image-text-retrieval
- zero-shot-audio-classification
- multimodal-embeddings
language:
- en
metrics:
- recall@1
- recall@5
- accuracy
model-index:
- name: pe-single-L14-a025
  results:
  - task:
      type: text-to-image-retrieval
      dataset:
        name: COCO slice (150, MMEB MSCOCO_i2t rows)
    metrics:
    - type: recall@1
      value: 0.766
      name: t2i R@1 (search slice n=64)
    - type: recall@5
      value: 0.95
      name: t2i R@5 (approx, search slice)
  - task:
      type: zero-shot-audio-classification
      dataset:
        name: ESC-50 slice (100 clips)
    metrics:
    - type: accuracy
      value: 0.96
      name: ESC-50 accuracy (search slice)
---

# pe-single-L14 — one fused Meta Perception Encoder, no router

Single-file joint embedding model (audio + image/video + text, 1024-d shared
space, 1.03B params) fusing four Meta PE family members:

| source | HF id | what was fused |
|---|---|---|
| PE-AV base (anchor) | `facebook/pe-av-base` | full joint model kept as scaffold |
| PE A-Frame base | `facebook/pe-a-frame-base` | audio tower, weight 0.25 |
| PE-Core-L14-336 | `facebook/PE-Core-L14-336` | vision trunk (near-identical to anchor video tower — init copy, ~no-op) |
| PE-Spatial-L14-448 | `facebook/PE-Spatial-L14-448` | vision trunk, weight 0.0 in this tag (hurt COCO alignment; see below) |

Why L-scale: PE-Core/Spatial-G14 are width-1536 × 50 layers while PE-AV/A-Frame
are width-1024 — cross-width tensors cannot be weight-averaged. L/14
(1024-wide, 24 layers, patch 14) is the scale where Core + Spatial + AV-video
share width/depth/patch, so true weight fusion is possible.

## Status (2026-10-04): v3 TUNED — local MPS fine-tune, new SOTA

Heads-only contrastive tune (COCO-train 4000 pairs x 2 epochs, frozen towers,
local Apple-Silicon MPS, fp32). Full-slice eval (COCO n=150, ESC-50 n=200):

| model | COCO t2i R@1 / R@5 | COCO i2t R@1 / R@5 | ESC-50 acc |
|---|---|---|---|
| `facebook/pe-av-base` (anchor) | 0.720 / 0.953 | 0.013 / 0.073 | 0.885 |
| frozen fusion (this repo, previous tag) | 0.713 / 0.953 | 0.013 / 0.047 | 0.895 |
| **v3 tuned (this revision)** | **0.827 / 0.980** | **0.727 / 0.973** | 0.895 |

The tune unlocked the text direction (i2t 0.013 -> 0.727) that no weight
blend touched, and lifted t2i +0.11. ESC unchanged (v1 trained video<->text
only; audio<->text tuning is next). Previous certification notes below remain
the record for the frozen init.

Full-slice validation (COCO n=150, ESC-50 n=200) vs anchor:

| model | COCO t2i R@1 / R@5 | COCO i2t R@1 | ESC-50 acc |
|---|---|---|---|
| `facebook/pe-av-base` (anchor) | 0.720 / 0.953 | 0.013 | 0.885 |
| this tag | 0.713 / 0.953 | 0.013 | 0.895 |

Verdict: statistical tie (diffs within SE). Weight fusion preserves anchor
quality while carrying A-Frame/Core/Spatial weights in one file — it does not
beat the anchor on these benches by itself. Measured gains are deferred to
alignment fine-tuning, for which this tag is the frozen init. Certification:
9-build grid + 18-genome evolution (fixed machinery) + full-slice validation,
all agreeing (audio light, vision ~zero).

15-metric suite (COCO n=300, ESC-50 5x120, MPS fp32) vs anchor:
vision tied (V1 t2i R@1 0.613 vs 0.617, V5 0.637 vs 0.643, V4 min 0.527 vs
0.547; V2/V3 at chance for both — text embeds are dummy-video dominated in
this harness, a comparative-only handicap); audio 4/5 folds +0.017..+0.025
(fold3 -0.017), prompt-mean +0.008; joint diagnostics mixed tiny deltas
(J5 +0.033). No regression anywhere; consistent small audio edge. Full table:
suite.json in the GitHub repo.

- Evolution rerun in progress after two voided attempts (documented below) —
  if it certifies a better genome, that becomes the next revision.
- Layer-depth probe: naive mid-network readout does NOT beat the joint output
  (t2i R@1 ≤ 0.06 vs 0.79) — mid-network features need trained per-layer
  pooling (fine-tune phase), not a free readout change.
- Voided runs: (1) blend-of-blend contamination via shared-storage state_dict;
  (2) silent no-op `load_state_dict` on this model's dual-prefix key layout;
  (3) joint-path audio mismatch that faked ESC collapse at high audio weight.
  All fixed (pristine clones, `copy_` by param name, standalone-tower-only
  audio mapping) before trusting any evolution number.

## How it was built

1. Audio tower: exact-shape suffix-mapped average of AV-audio and A-Frame-audio
   (both `pe_audio_encoder`, 1024/16L/8H — 422/422 tensors matched).
2. Video trunk: explicit native-PE → timm key map (patch_embed, cls_token,
   pos_embed@336px, norms, QKV, MLP ×24 blocks). Pool/head/proj kept from anchor.
3. Blend weights chosen by measured 3×3 grid search over COCO-retrieval +
   ESC-50 slices (script `search_pe_weights.py`), not by vibes. Vision donor
   weight degraded COCO t2i monotonically (0.766 → 0.64), so the winning tag
   keeps anchor vision and blends audio at 0.25.

## Usage

```python
from transformers import AutoModel, AutoProcessor
m = AutoModel.from_pretrained("MC7ever/pe-single-L14", trust_remote_code=True)
p = AutoProcessor.from_pretrained("MC7ever/pe-single-L14", trust_remote_code=True)
inputs = p(videos=[frames], audio=[waveform], text=["a cat playing in rain"],
           return_tensors="pt")
out = m(**inputs)  # out.video_embeds, out.audio_embeds, out.text_video_embeds, ...
```

Retrieval: dot-product `video_embeds @ text_video_embeds.T`.
Zero-shot audio: argmax over `audio_embeds @ text_audio_embeds.T` with
`"the sound of {label}"` prompts.

## Honest limits

- Numbers above are CPU-run slices (COCO n=64, ESC n=100), comparative
  anchor-vs-variant only — not official MMEB/PE-paper figures.
- i2t (text→image caption ranking over 150 similar COCO captions) sits near
  floor for anchor and variants alike; t2i and audio carry the signal.
- Encoder only: no generation, detection, or segmentation heads.
- Next: evolution-based per-tower weights, intermediate-layer readout
  (PE paper: best features are mid-network), GPU alignment fine-tune.

## Reproduce

`fuse_pe_single.py` (fusion) · `eval_pe_bench.py` (COCO+ESC-50 bench) ·
`search_pe_weights.py` (grid) · `evo_fuse.py` (evolution) · `layer_sweep.py`
(depth probe). Benchmark slice metadata: MMEB `MSCOCO_i2t` rows + COCO val2014
images + `ashraq/esc50` streaming clips.