OVIE-512

In-the-Wild Monocular Pretraining for Novel View Generation

Paper GitHub Collection

Part of the OVIE collection.

OVIE is resolution-agnostic by construction: the encoder and decoder are fully convolutional and the ViT bottleneck regenerates its positional encodings for the new grid. This checkpoint is the base model retrained at 512Γ—512 under the identical recipe (same in-the-wild mix with MoGe-2 pseudo-pairs), for 250K steps at global batch size 256, changing only the resolution.

At a like-for-like 256Γ—256 evaluation it matches or improves on the base model on PSNR, SSIM, LPIPS and MEt3R on both benchmarks, at a cost of roughly one FID point, which the paper attributes to the training mix β€” part of it sits below 512 and is upscaled, carrying little true high-frequency detail. Inference stays a single feed-forward pass with no per-scene construction: 41.6 ms per view on an H100 against 8.6 ms at 256Γ—256.

Metrics

eval@256 brings every output to 256Γ—256 before scoring (like-for-like with the base model); eval@512 scores at native resolution and is given for reference only, since SSIM and LPIPS are resolution-dependent.

Model Benchmark PSNR ↑ SSIM ↑ LPIPS ↓ FID ↓ MEt3R ↓
OVIE-512, eval@256 RealEstate10K 19.1 0.611 0.283 7.62 0.034
OVIE-512, eval@256 DL3DV 15.2 0.389 0.464 14.8 0.077
OVIE-512, eval@512 RealEstate10K 18.9 0.658 0.344 β€” β€”
OVIE-512, eval@512 DL3DV 15.1 0.475 0.509 β€” β€”

For reference, the base OVIE (256Γ—256) scores 18.8 / 0.602 / 0.279 / 6.74 / 0.035 on RealEstate10K and 14.8 / 0.369 / 0.464 / 13.6 / 0.078 on DL3DV.

Usage

import torch
from models.models import OVIEModel
from utils.pose_enc import extri_intri_to_pose_encoding
from torchvision.transforms import ToTensor
from PIL import Image

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

model = OVIEModel.from_pretrained("kyutai/ovie-512").to(device)
model.eval()
image_size = model.image_size  # 512, read from the saved config

img_pil = Image.open("image.jpg").convert("RGB").resize((image_size, image_size))
img_tensor = ToTensor()(img_pil).unsqueeze(0).to(device)

extrinsics = torch.tensor([[[1.0, 0.0, 0.0, -1.25],
                            [0.0, 1.0, 0.0,  0.5],
                            [0.0, 0.0, 1.0, -2.0]]], device=device)
dummy_intrinsics = torch.zeros(1, 1, 3, 3, device=device)

camera = extri_intri_to_pose_encoding(
    extrinsics=extrinsics.unsqueeze(0),
    intrinsics=dummy_intrinsics,
    image_size_hw=(image_size, image_size),
)
cam_token = camera[..., :7].squeeze(0)

with torch.no_grad():
    pred = model(x=img_tensor, cam_params=cam_token)  # (1, 3, 512, 512) in [0, 1]

Note the output resolution is fixed at training time: unlike methods that render from an explicit 3D representation, OVIE-512 cannot produce arbitrary resolutions at inference.

To evaluate with the repository's script, pass --image_size 512:

uv run python evaluate.py \
    --dataset_path /PATH/TO/RE10K/TEST \
    --config_path configs/config_ovie.yaml \
    --from_pretrained kyutai/ovie-512 \
    --image_size 512 --stride 3 --num_target_frames 14

See the repository for installation, data preprocessing, and the full benchmark table.

Citation

@misc{ovie2026,
      title={One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation},
      author={Adrien Ramanana Rahary and Nicolas Dufour and Patrick Perez and David Picard},
      year={2026},
      eprint={2603.23488},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2603.23488},
}
Downloads last month
62
Safetensors
Model size
0.1B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Collection including kyutai/ovie-512

Paper for kyutai/ovie-512