Feature Extraction
Transformers
Safetensors
PerceptionEncoder
English
pe_audio_video
audio-text-retrieval
image-text-retrieval
zero-shot-audio-classification
multimodal-embeddings
Eval Results (legacy)
Instructions to use MC7ever/pe-single-L14 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MC7ever/pe-single-L14 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="MC7ever/pe-single-L14")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModel processor = AutoProcessor.from_pretrained("MC7ever/pe-single-L14") model = AutoModel.from_pretrained("MC7ever/pe-single-L14", device_map="auto") - PerceptionEncoder
How to use MC7ever/pe-single-L14 with PerceptionEncoder:
# Use any PE model as a vision encoder import core.vision_encoder.pe as pe model = pe.VisionTransformer.from_config("MC7ever/pe-single-L14", pretrained=True) - Notebooks
- Google Colab
- Kaggle
|
Download README.md from MC7ever/pe-single-L14: direct link, hf CLI and curl.
- Browser
- Download file 6.48 kB
-
https://huggingface.co/MC7ever/pe-single-L14/resolve/main/README.md
- Command line
-
hf download hf://MC7ever/pe-single-L14/README.md
-
curl -L -o README.md https://huggingface.co/MC7ever/pe-single-L14/resolve/main/README.md
6.48 kB
| license: apache-2.0 | |
| library_name: transformers | |
| pipeline_tag: feature-extraction | |
| tags: | |
| - pe_audio_video | |
| - perception-encoder | |
| - audio-text-retrieval | |
| - image-text-retrieval | |
| - zero-shot-audio-classification | |
| - multimodal-embeddings | |
| language: | |
| - en | |
| metrics: | |
| - recall@1 | |
| - recall@5 | |
| - accuracy | |
| model-index: | |
| - name: pe-single-L14-a025 | |
| results: | |
| - task: | |
| type: text-to-image-retrieval | |
| dataset: | |
| name: COCO slice (150, MMEB MSCOCO_i2t rows) | |
| metrics: | |
| - type: recall@1 | |
| value: 0.766 | |
| name: t2i R@1 (search slice n=64) | |
| - type: recall@5 | |
| value: 0.95 | |
| name: t2i R@5 (approx, search slice) | |
| - task: | |
| type: zero-shot-audio-classification | |
| dataset: | |
| name: ESC-50 slice (100 clips) | |
| metrics: | |
| - type: accuracy | |
| value: 0.96 | |
| name: ESC-50 accuracy (search slice) | |
| # pe-single-L14 — one fused Meta Perception Encoder, no router | |
| Single-file joint embedding model (audio + image/video + text, 1024-d shared | |
| space, 1.03B params) fusing four Meta PE family members: | |
| | source | HF id | what was fused | | |
| |---|---|---| | |
| | PE-AV base (anchor) | `facebook/pe-av-base` | full joint model kept as scaffold | | |
| | PE A-Frame base | `facebook/pe-a-frame-base` | audio tower, weight 0.25 | | |
| | PE-Core-L14-336 | `facebook/PE-Core-L14-336` | vision trunk (near-identical to anchor video tower — init copy, ~no-op) | | |
| | PE-Spatial-L14-448 | `facebook/PE-Spatial-L14-448` | vision trunk, weight 0.0 in this tag (hurt COCO alignment; see below) | | |
| Why L-scale: PE-Core/Spatial-G14 are width-1536 × 50 layers while PE-AV/A-Frame | |
| are width-1024 — cross-width tensors cannot be weight-averaged. L/14 | |
| (1024-wide, 24 layers, patch 14) is the scale where Core + Spatial + AV-video | |
| share width/depth/patch, so true weight fusion is possible. | |
| ## Status (2026-10-04): v3 TUNED — local MPS fine-tune, new SOTA | |
| Heads-only contrastive tune (COCO-train 4000 pairs x 2 epochs, frozen towers, | |
| local Apple-Silicon MPS, fp32). Full-slice eval (COCO n=150, ESC-50 n=200): | |
| | model | COCO t2i R@1 / R@5 | COCO i2t R@1 / R@5 | ESC-50 acc | | |
| |---|---|---|---| | |
| | `facebook/pe-av-base` (anchor) | 0.720 / 0.953 | 0.013 / 0.073 | 0.885 | | |
| | frozen fusion (this repo, previous tag) | 0.713 / 0.953 | 0.013 / 0.047 | 0.895 | | |
| | **v3 tuned (this revision)** | **0.827 / 0.980** | **0.727 / 0.973** | 0.895 | | |
| The tune unlocked the text direction (i2t 0.013 -> 0.727) that no weight | |
| blend touched, and lifted t2i +0.11. ESC unchanged (v1 trained video<->text | |
| only; audio<->text tuning is next). Previous certification notes below remain | |
| the record for the frozen init. | |
| Full-slice validation (COCO n=150, ESC-50 n=200) vs anchor: | |
| | model | COCO t2i R@1 / R@5 | COCO i2t R@1 | ESC-50 acc | | |
| |---|---|---|---| | |
| | `facebook/pe-av-base` (anchor) | 0.720 / 0.953 | 0.013 | 0.885 | | |
| | this tag | 0.713 / 0.953 | 0.013 | 0.895 | | |
| Verdict: statistical tie (diffs within SE). Weight fusion preserves anchor | |
| quality while carrying A-Frame/Core/Spatial weights in one file — it does not | |
| beat the anchor on these benches by itself. Measured gains are deferred to | |
| alignment fine-tuning, for which this tag is the frozen init. Certification: | |
| 9-build grid + 18-genome evolution (fixed machinery) + full-slice validation, | |
| all agreeing (audio light, vision ~zero). | |
| 15-metric suite (COCO n=300, ESC-50 5x120, MPS fp32) vs anchor: | |
| vision tied (V1 t2i R@1 0.613 vs 0.617, V5 0.637 vs 0.643, V4 min 0.527 vs | |
| 0.547; V2/V3 at chance for both — text embeds are dummy-video dominated in | |
| this harness, a comparative-only handicap); audio 4/5 folds +0.017..+0.025 | |
| (fold3 -0.017), prompt-mean +0.008; joint diagnostics mixed tiny deltas | |
| (J5 +0.033). No regression anywhere; consistent small audio edge. Full table: | |
| suite.json in the GitHub repo. | |
| - Evolution rerun in progress after two voided attempts (documented below) — | |
| if it certifies a better genome, that becomes the next revision. | |
| - Layer-depth probe: naive mid-network readout does NOT beat the joint output | |
| (t2i R@1 ≤ 0.06 vs 0.79) — mid-network features need trained per-layer | |
| pooling (fine-tune phase), not a free readout change. | |
| - Voided runs: (1) blend-of-blend contamination via shared-storage state_dict; | |
| (2) silent no-op `load_state_dict` on this model's dual-prefix key layout; | |
| (3) joint-path audio mismatch that faked ESC collapse at high audio weight. | |
| All fixed (pristine clones, `copy_` by param name, standalone-tower-only | |
| audio mapping) before trusting any evolution number. | |
| ## How it was built | |
| 1. Audio tower: exact-shape suffix-mapped average of AV-audio and A-Frame-audio | |
| (both `pe_audio_encoder`, 1024/16L/8H — 422/422 tensors matched). | |
| 2. Video trunk: explicit native-PE → timm key map (patch_embed, cls_token, | |
| pos_embed@336px, norms, QKV, MLP ×24 blocks). Pool/head/proj kept from anchor. | |
| 3. Blend weights chosen by measured 3×3 grid search over COCO-retrieval + | |
| ESC-50 slices (script `search_pe_weights.py`), not by vibes. Vision donor | |
| weight degraded COCO t2i monotonically (0.766 → 0.64), so the winning tag | |
| keeps anchor vision and blends audio at 0.25. | |
| ## Usage | |
| ```python | |
| from transformers import AutoModel, AutoProcessor | |
| m = AutoModel.from_pretrained("MC7ever/pe-single-L14", trust_remote_code=True) | |
| p = AutoProcessor.from_pretrained("MC7ever/pe-single-L14", trust_remote_code=True) | |
| inputs = p(videos=[frames], audio=[waveform], text=["a cat playing in rain"], | |
| return_tensors="pt") | |
| out = m(**inputs) # out.video_embeds, out.audio_embeds, out.text_video_embeds, ... | |
| ``` | |
| Retrieval: dot-product `video_embeds @ text_video_embeds.T`. | |
| Zero-shot audio: argmax over `audio_embeds @ text_audio_embeds.T` with | |
| `"the sound of {label}"` prompts. | |
| ## Honest limits | |
| - Numbers above are CPU-run slices (COCO n=64, ESC n=100), comparative | |
| anchor-vs-variant only — not official MMEB/PE-paper figures. | |
| - i2t (text→image caption ranking over 150 similar COCO captions) sits near | |
| floor for anchor and variants alike; t2i and audio carry the signal. | |
| - Encoder only: no generation, detection, or segmentation heads. | |
| - Next: evolution-based per-tower weights, intermediate-layer readout | |
| (PE paper: best features are mid-network), GPU alignment fine-tune. | |
| ## Reproduce | |
| `fuse_pe_single.py` (fusion) · `eval_pe_bench.py` (COCO+ESC-50 bench) · | |
| `search_pe_weights.py` (grid) · `evo_fuse.py` (evolution) · `layer_sweep.py` | |
| (depth probe). Benchmark slice metadata: MMEB `MSCOCO_i2t` rows + COCO val2014 | |
| images + `ashraq/esc50` streaming clips. | |