--- license: mit tags: - novel-view-synthesis - image-to-image - computer-vision - pytorch - model_hub_mixin - pytorch_model_hub_mixin pipeline_tag: image-to-image arxiv: 2603.23488 library_name: pytorch --- # OVIE-512 **In-the-Wild Monocular Pretraining for Novel View Generation** [![Paper](https://img.shields.io/badge/arXiv-2603.23488-red?logo=arxiv)](https://arxiv.org/abs/2603.23488) [![GitHub](https://img.shields.io/badge/GitHub-kyutai--labs%2Fovie-black?logo=github)](https://github.com/kyutai-labs/ovie) [![Collection](https://img.shields.io/badge/🤗-OVIE%20collection-yellow)](https://huggingface.co/collections/kyutai/ovie) Part of the [OVIE collection](https://huggingface.co/collections/kyutai/ovie). OVIE is resolution-agnostic by construction: the encoder and decoder are fully convolutional and the ViT bottleneck regenerates its positional encodings for the new grid. This checkpoint is the base model retrained at **512×512** under the identical recipe (same in-the-wild mix with MoGe-2 pseudo-pairs), for 250K steps at global batch size 256, changing only the resolution. At a like-for-like 256×256 evaluation it matches or improves on the base model on PSNR, SSIM, LPIPS and MEt3R on both benchmarks, at a cost of roughly one FID point, which the paper attributes to the training mix — part of it sits below 512 and is upscaled, carrying little true high-frequency detail. Inference stays a single feed-forward pass with no per-scene construction: 41.6 ms per view on an H100 against 8.6 ms at 256×256. ## Metrics `eval@256` brings every output to 256×256 before scoring (like-for-like with the base model); `eval@512` scores at native resolution and is given for reference only, since SSIM and LPIPS are resolution-dependent. | Model | Benchmark | PSNR ↑ | SSIM ↑ | LPIPS ↓ | FID ↓ | MEt3R ↓ | |---|---|---|---|---|---|---| | OVIE-512, `eval@256` | RealEstate10K | 19.1 | 0.611 | 0.283 | 7.62 | 0.034 | | OVIE-512, `eval@256` | DL3DV | 15.2 | 0.389 | 0.464 | 14.8 | 0.077 | | OVIE-512, `eval@512` | RealEstate10K | 18.9 | 0.658 | 0.344 | — | — | | OVIE-512, `eval@512` | DL3DV | 15.1 | 0.475 | 0.509 | — | — | For reference, the base [OVIE](https://huggingface.co/kyutai/ovie) (256×256) scores 18.8 / 0.602 / 0.279 / 6.74 / 0.035 on RealEstate10K and 14.8 / 0.369 / 0.464 / 13.6 / 0.078 on DL3DV. ## Usage ```python import torch from models.models import OVIEModel from utils.pose_enc import extri_intri_to_pose_encoding from torchvision.transforms import ToTensor from PIL import Image device = torch.device("cuda" if torch.cuda.is_available() else "cpu") model = OVIEModel.from_pretrained("kyutai/ovie-512").to(device) model.eval() image_size = model.image_size # 512, read from the saved config img_pil = Image.open("image.jpg").convert("RGB").resize((image_size, image_size)) img_tensor = ToTensor()(img_pil).unsqueeze(0).to(device) extrinsics = torch.tensor([[[1.0, 0.0, 0.0, -1.25], [0.0, 1.0, 0.0, 0.5], [0.0, 0.0, 1.0, -2.0]]], device=device) dummy_intrinsics = torch.zeros(1, 1, 3, 3, device=device) camera = extri_intri_to_pose_encoding( extrinsics=extrinsics.unsqueeze(0), intrinsics=dummy_intrinsics, image_size_hw=(image_size, image_size), ) cam_token = camera[..., :7].squeeze(0) with torch.no_grad(): pred = model(x=img_tensor, cam_params=cam_token) # (1, 3, 512, 512) in [0, 1] ``` Note the output resolution is fixed at training time: unlike methods that render from an explicit 3D representation, OVIE-512 cannot produce arbitrary resolutions at inference. To evaluate with the repository's script, pass `--image_size 512`: ```sh uv run python evaluate.py \ --dataset_path /PATH/TO/RE10K/TEST \ --config_path configs/config_ovie.yaml \ --from_pretrained kyutai/ovie-512 \ --image_size 512 --stride 3 --num_target_frames 14 ``` See the [repository](https://github.com/kyutai-labs/ovie) for installation, data preprocessing, and the full benchmark table. ## Citation ```bibtex @misc{ovie2026, title={One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation}, author={Adrien Ramanana Rahary and Nicolas Dufour and Patrick Perez and David Picard}, year={2026}, eprint={2603.23488}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2603.23488}, } ```