pe-single-L14 / README.md
MC7ever's picture
card: v3 tuned numbers (i2t 0.727, t2i 0.827)
6b632dd verified
|
Raw History Blame Contribute Delete
6.48 kB
metadata
license: apache-2.0
library_name: transformers
pipeline_tag: feature-extraction
tags:
  - pe_audio_video
  - perception-encoder
  - audio-text-retrieval
  - image-text-retrieval
  - zero-shot-audio-classification
  - multimodal-embeddings
language:
  - en
metrics:
  - recall@1
  - recall@5
  - accuracy
model-index:
  - name: pe-single-L14-a025
    results:
      - task:
          type: text-to-image-retrieval
          dataset:
            name: COCO slice (150, MMEB MSCOCO_i2t rows)
        metrics:
          - type: recall@1
            value: 0.766
            name: t2i R@1 (search slice n=64)
          - type: recall@5
            value: 0.95
            name: t2i R@5 (approx, search slice)
      - task:
          type: zero-shot-audio-classification
          dataset:
            name: ESC-50 slice (100 clips)
        metrics:
          - type: accuracy
            value: 0.96
            name: ESC-50 accuracy (search slice)

pe-single-L14 — one fused Meta Perception Encoder, no router

Single-file joint embedding model (audio + image/video + text, 1024-d shared space, 1.03B params) fusing four Meta PE family members:

source HF id what was fused
PE-AV base (anchor) facebook/pe-av-base full joint model kept as scaffold
PE A-Frame base facebook/pe-a-frame-base audio tower, weight 0.25
PE-Core-L14-336 facebook/PE-Core-L14-336 vision trunk (near-identical to anchor video tower — init copy, ~no-op)
PE-Spatial-L14-448 facebook/PE-Spatial-L14-448 vision trunk, weight 0.0 in this tag (hurt COCO alignment; see below)

Why L-scale: PE-Core/Spatial-G14 are width-1536 × 50 layers while PE-AV/A-Frame are width-1024 — cross-width tensors cannot be weight-averaged. L/14 (1024-wide, 24 layers, patch 14) is the scale where Core + Spatial + AV-video share width/depth/patch, so true weight fusion is possible.

Status (2026-10-04): v3 TUNED — local MPS fine-tune, new SOTA

Heads-only contrastive tune (COCO-train 4000 pairs x 2 epochs, frozen towers, local Apple-Silicon MPS, fp32). Full-slice eval (COCO n=150, ESC-50 n=200):

model COCO t2i R@1 / R@5 COCO i2t R@1 / R@5 ESC-50 acc
facebook/pe-av-base (anchor) 0.720 / 0.953 0.013 / 0.073 0.885
frozen fusion (this repo, previous tag) 0.713 / 0.953 0.013 / 0.047 0.895
v3 tuned (this revision) 0.827 / 0.980 0.727 / 0.973 0.895

The tune unlocked the text direction (i2t 0.013 -> 0.727) that no weight blend touched, and lifted t2i +0.11. ESC unchanged (v1 trained video<->text only; audio<->text tuning is next). Previous certification notes below remain the record for the frozen init.

Full-slice validation (COCO n=150, ESC-50 n=200) vs anchor:

model COCO t2i R@1 / R@5 COCO i2t R@1 ESC-50 acc
facebook/pe-av-base (anchor) 0.720 / 0.953 0.013 0.885
this tag 0.713 / 0.953 0.013 0.895

Verdict: statistical tie (diffs within SE). Weight fusion preserves anchor quality while carrying A-Frame/Core/Spatial weights in one file — it does not beat the anchor on these benches by itself. Measured gains are deferred to alignment fine-tuning, for which this tag is the frozen init. Certification: 9-build grid + 18-genome evolution (fixed machinery) + full-slice validation, all agreeing (audio light, vision ~zero).

15-metric suite (COCO n=300, ESC-50 5x120, MPS fp32) vs anchor: vision tied (V1 t2i R@1 0.613 vs 0.617, V5 0.637 vs 0.643, V4 min 0.527 vs 0.547; V2/V3 at chance for both — text embeds are dummy-video dominated in this harness, a comparative-only handicap); audio 4/5 folds +0.017..+0.025 (fold3 -0.017), prompt-mean +0.008; joint diagnostics mixed tiny deltas (J5 +0.033). No regression anywhere; consistent small audio edge. Full table: suite.json in the GitHub repo.

  • Evolution rerun in progress after two voided attempts (documented below) — if it certifies a better genome, that becomes the next revision.
  • Layer-depth probe: naive mid-network readout does NOT beat the joint output (t2i R@1 ≤ 0.06 vs 0.79) — mid-network features need trained per-layer pooling (fine-tune phase), not a free readout change.
  • Voided runs: (1) blend-of-blend contamination via shared-storage state_dict; (2) silent no-op load_state_dict on this model's dual-prefix key layout; (3) joint-path audio mismatch that faked ESC collapse at high audio weight. All fixed (pristine clones, copy_ by param name, standalone-tower-only audio mapping) before trusting any evolution number.

How it was built

  1. Audio tower: exact-shape suffix-mapped average of AV-audio and A-Frame-audio (both pe_audio_encoder, 1024/16L/8H — 422/422 tensors matched).
  2. Video trunk: explicit native-PE → timm key map (patch_embed, cls_token, pos_embed@336px, norms, QKV, MLP ×24 blocks). Pool/head/proj kept from anchor.
  3. Blend weights chosen by measured 3×3 grid search over COCO-retrieval + ESC-50 slices (script search_pe_weights.py), not by vibes. Vision donor weight degraded COCO t2i monotonically (0.766 → 0.64), so the winning tag keeps anchor vision and blends audio at 0.25.

Usage

from transformers import AutoModel, AutoProcessor
m = AutoModel.from_pretrained("MC7ever/pe-single-L14", trust_remote_code=True)
p = AutoProcessor.from_pretrained("MC7ever/pe-single-L14", trust_remote_code=True)
inputs = p(videos=[frames], audio=[waveform], text=["a cat playing in rain"],
           return_tensors="pt")
out = m(**inputs)  # out.video_embeds, out.audio_embeds, out.text_video_embeds, ...

Retrieval: dot-product video_embeds @ text_video_embeds.T. Zero-shot audio: argmax over audio_embeds @ text_audio_embeds.T with "the sound of {label}" prompts.

Honest limits

  • Numbers above are CPU-run slices (COCO n=64, ESC n=100), comparative anchor-vs-variant only — not official MMEB/PE-paper figures.
  • i2t (text→image caption ranking over 150 similar COCO captions) sits near floor for anchor and variants alike; t2i and audio carry the signal.
  • Encoder only: no generation, detection, or segmentation heads.
  • Next: evolution-based per-tower weights, intermediate-layer readout (PE paper: best features are mid-network), GPU alignment fine-tune.

Reproduce

fuse_pe_single.py (fusion) · eval_pe_bench.py (COCO+ESC-50 bench) · search_pe_weights.py (grid) · evo_fuse.py (evolution) · layer_sweep.py (depth probe). Benchmark slice metadata: MMEB MSCOCO_i2t rows + COCO val2014 images + ashraq/esc50 streaming clips.