Meridian / docs /assets /research /README.md
yycc's picture
Meridian
9f57754
|
Raw History Blame Contribute Delete
9.72 kB

Research figure assets

← Research article · Technical method

Illustrated method

The article uses one integrated illustration: editable SVG · PNG · provenance.

  • Video strips: actual matched source / warp / output samples from meridian_longtake_l150_nba3_apex_right14_175. The front card is output index 95; the partially visible back cards are indices 40 and 150, representing video rather than a single-image input. Aspect ratios are preserved; the front frames are uncropped.
  • 3D illustration: a procedural, colored basketball point cloud and camera frustums, explicitly labeled schematic. These are not saved VGGT-Omega points, estimated poses, or the measured target path from this take. No reconstruction was run to make the figure.
  • Data flow: VGGT-Omega estimates depth and source cameras; source RGB and depth are unprojected into per-frame colored points. User-specified target cameras produce the warp. The time-aligned source video bypasses geometry and joins the warp as the model's other video input. Blue frustums denote estimated source cameras; gold denotes authored cameras.

This depicts the released reconstruct, unproject, and warp pipeline, not a fused persistent world or direct point-cloud conditioning of the video model. The model consumes two videos, not the plotted points or camera icons. As in the compact earlier figures, noise, VAE/token packing, and the discarded audio branch are omitted.

The front samples map to prepared-input frame 70, original movie frame 85, PTS 2.836167 s. The other frame mappings and file hashes are in the provenance JSON. NBA source-use clearance remains pending; no public promotional permission or endorsement is implied. The Spring figure's CC BY license does not apply to the NBA samples.

Rebuild from the repository root:

/home/chenyun/miniforge3/envs/wan_new/bin/python docs/assets/research/build_illustrated_method.py

This reads retained videos and writes only this illustration's SVG, PNG and provenance. It uses CPU decoding and vector rasterization; no Studio requests, model runs, production video edits, or generative image replacements. Earlier figures are retained below.

NBA method figures

The previous two-figure version is retained for reference:

  1. Camera poses and warp: editable SVG · PNG. VGGT-Omega estimates source poses and depth; colored points are reprojected with manually specified target-camera poses. This is a schematic, not a measured camera-path plot.
  2. Video + warp → output: editable SVG · PNG. The three NBA panels are actual decoded frames from one take, with no generated replacements, retouching, or cropping.

Frame provenance and media hashes · Source, warp, output, and recorded controls

All three panels use output index 95 (zero-based) of meridian_longtake_l150_nba3_apex_right14_175. This is inside the requested hold: prepared-input frame 70, original movie frame 85, original PTS 2.836167 s. The source panel comes from the take's time-aligned source.mp4, not frame 95 of the original movie. The output uses Full200 / LoRA150 / CLI4 / shift3 / seed1234. This single-frame illustration does not certify exact pose locking or continuous-motion quality.

NBA footage was supplied for local research. Public promotional permission and endorsement are not established. The Spring figure's CC BY license below does not apply to the NBA panels.

The SVGs are the editable sources. PNGs are direct rasterizations. For either figure, run from the repository root, replacing meridian_poses with meridian_nba_generation for the second figure:

ffmpeg -v error -threads 2 -i docs/assets/research/meridian_poses.svg \
  -frames:v 1 -threads 2 -y docs/assets/research/meridian_poses.png

Architecture overview

Editable, full-size SVG · PNG · Frame provenance

The previous single-diagram version separates estimated source geometry, user-authored target cameras and time, and the two video-model inputs. Its data flow was checked against the released code:

Diagram element Implementation
VGGT-Omega depth and source-camera estimates; unprojection with source RGB reconstruct, unproject, warp
Camera position, look-at point, focal scale, and integer source-frame map plan_path
The same frame map selects both source images and geometry; the target camera projects the points geo
Two VAE-encoded video references, packed as tokens; target denoising and decoding pack, denoise, decode_video, do_render

Target-camera control is explicit, but does not guarantee pixel-perfect generated frames. Studio orientation is derived from position and look-at with zero roll; focal scale multiplies the estimated source focal lengths, rather than specifying an arbitrary intrinsic matrix. Time selection repeats or skips supplied frames, without interpolating new motion. The diagram omits target noise, reference noise augmentation, and the discarded audio branch; the technical method covers those details.

The three photographic panels reuse the same embedded JPEGs, unchanged, from the original figure below. The frame indices, source attribution, transformations, and provenance below apply to both figures. The schematic video-strip icon is not a data sample. The SVG is the editable source; its PNG is a direct rasterization, not an AI-generated or retouched image.

To refresh the PNG after editing the SVG, run from the repository root:

ffmpeg -v error -threads 2 -i docs/assets/research/meridian_architecture.svg \
  -frames:v 1 -threads 2 -y docs/assets/research/meridian_architecture.png

Two controls, two references, one new shot

Full-size SVG · PNG · Machine-readable provenance

The upper diagram follows the released inference implementation, not a proposed architecture:

  • recam/geometry.py: joint source reconstruction, filtered per-frame colored points, z-buffered projection and grey uncovered pixels. No fused persistent 4D scene or geometric inpainting.
  • recam/path.py: source-frame selection s(t) and target camera C(t).
  • recam/h3.py and service/app.py: both video references are VAE encoded and packed as reference tokens; the model denoises the target, then the VAE decodes it. Coverage is diagnostic, not a separate transformer mask input. The diagram omits target noise and the discarded audio branch; see the technical method for the full layout.

The lower panels are actual decoded source, geometric-reference and generated frames from meridian_grand_l150_flowers_forward70_baseaim243, all at output index 121 (zero-based). This maps to prepared-input frame 121 and original movie frame 9057 / PTS 377.382 s. The geometric panel is taken from the saved conditioning-resolution preview, not a cleaned-up render. Aspect ratios are preserved; small black margins are layout padding. This is a single-frame illustration, not evidence of continuous motion quality or recovered ground truth.

The source/output exports are 1920 × 800; the geometric export is 960 × 416. These are this take's production settings, not the default quickstart buckets. The generated frame uses Full200 + LoRA150, CLI --steps 4 --flow-shift 3 --seed 1234; it is not a teacher-30 result.

Attribution and transformations

Spring (2019), © Blender Foundation | project. Retained source: Spring — Blender Open Movie, identified as CC BY 4.0 in the retained source records.

The input was slowed by repeating source frames before inference. The geometry projection and generated view are transformations of that material. Figure preparation extracts one matched frame, downsamples source/output thumbnails, JPEG-encodes the panels and fits them without cropping. No generative image editing, enhancement, surface repair or color treatment is used for the figure. Retain the attribution and transformation notice when reusing the visual.

Original input preparation and exact map · Recorded recipe and raw audit · Source, projection and output in motion

Rebuild

From the repository root, with FFmpeg's librsvg decoder available:

/home/chenyun/miniforge3/envs/wan_new/bin/python docs/assets/research/build_method_figure.py

This reads the retained MP4s and writes only the SVG, PNG and provenance beside this file. It uses CPU decoding and SVG rasterization; it does not invoke the model, contact the Studio service, or change production videos.