OVIE-512
In-the-Wild Monocular Pretraining for Novel View Generation
Part of the OVIE collection.
OVIE is resolution-agnostic by construction: the encoder and decoder are fully convolutional and the ViT bottleneck regenerates its positional encodings for the new grid. This checkpoint is the base model retrained at 512Γ512 under the identical recipe (same in-the-wild mix with MoGe-2 pseudo-pairs), for 250K steps at global batch size 256, changing only the resolution.
At a like-for-like 256Γ256 evaluation it matches or improves on the base model on PSNR, SSIM, LPIPS and MEt3R on both benchmarks, at a cost of roughly one FID point, which the paper attributes to the training mix β part of it sits below 512 and is upscaled, carrying little true high-frequency detail. Inference stays a single feed-forward pass with no per-scene construction: 41.6 ms per view on an H100 against 8.6 ms at 256Γ256.
Metrics
eval@256 brings every output to 256Γ256 before scoring (like-for-like with the base model); eval@512 scores at native resolution and is given for reference only, since SSIM and LPIPS are resolution-dependent.
| Model | Benchmark | PSNR β | SSIM β | LPIPS β | FID β | MEt3R β |
|---|---|---|---|---|---|---|
OVIE-512, eval@256 |
RealEstate10K | 19.1 | 0.611 | 0.283 | 7.62 | 0.034 |
OVIE-512, eval@256 |
DL3DV | 15.2 | 0.389 | 0.464 | 14.8 | 0.077 |
OVIE-512, eval@512 |
RealEstate10K | 18.9 | 0.658 | 0.344 | β | β |
OVIE-512, eval@512 |
DL3DV | 15.1 | 0.475 | 0.509 | β | β |
For reference, the base OVIE (256Γ256) scores 18.8 / 0.602 / 0.279 / 6.74 / 0.035 on RealEstate10K and 14.8 / 0.369 / 0.464 / 13.6 / 0.078 on DL3DV.
Usage
import torch
from models.models import OVIEModel
from utils.pose_enc import extri_intri_to_pose_encoding
from torchvision.transforms import ToTensor
from PIL import Image
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = OVIEModel.from_pretrained("kyutai/ovie-512").to(device)
model.eval()
image_size = model.image_size # 512, read from the saved config
img_pil = Image.open("image.jpg").convert("RGB").resize((image_size, image_size))
img_tensor = ToTensor()(img_pil).unsqueeze(0).to(device)
extrinsics = torch.tensor([[[1.0, 0.0, 0.0, -1.25],
[0.0, 1.0, 0.0, 0.5],
[0.0, 0.0, 1.0, -2.0]]], device=device)
dummy_intrinsics = torch.zeros(1, 1, 3, 3, device=device)
camera = extri_intri_to_pose_encoding(
extrinsics=extrinsics.unsqueeze(0),
intrinsics=dummy_intrinsics,
image_size_hw=(image_size, image_size),
)
cam_token = camera[..., :7].squeeze(0)
with torch.no_grad():
pred = model(x=img_tensor, cam_params=cam_token) # (1, 3, 512, 512) in [0, 1]
Note the output resolution is fixed at training time: unlike methods that render from an explicit 3D representation, OVIE-512 cannot produce arbitrary resolutions at inference.
To evaluate with the repository's script, pass --image_size 512:
uv run python evaluate.py \
--dataset_path /PATH/TO/RE10K/TEST \
--config_path configs/config_ovie.yaml \
--from_pretrained kyutai/ovie-512 \
--image_size 512 --stride 3 --num_target_frames 14
See the repository for installation, data preprocessing, and the full benchmark table.
Citation
@misc{ovie2026,
title={One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation},
author={Adrien Ramanana Rahary and Nicolas Dufour and Patrick Perez and David Picard},
year={2026},
eprint={2603.23488},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2603.23488},
}
- Downloads last month
- 62