Conventional novel view synthesis methods often pair high-capacity decoders with pixel-level reconstruction objectives that focus optimization on low-level photometric details. SNAP explores an alternative design with two components: predicting pre-trained feature tokens (from models like DINO, MAE, or SAM) instead of raw pixels, and using a lightweight cross-attention decoder with localized receptive fields. Under this design, the learned representations exhibit improved transfer across downstream 3D geometric benchmarks.
This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS requires reasoning about 3D scene structure, potentially enabling transferable multi-view geometric representations. However, representations learned from existing encoder-based NVS methods often show limited transfer to downstream geometric tasks. We investigate two contributing factors: spatially expressive decoders that can absorb multi-view spatial reasoning away from the scene encoder, and low-level pixel-space targets that emphasize high-frequency photometric details over geometric structure.
We present SNAP, a self-supervised encoder-decoder transformer designed with a pose-conditioned local decoder and a latent-space reconstruction objective. SNAP is general-purpose and evaluates across five geometric tasks: visual localization, pose estimation, point correspondence, depth estimation, and robot manipulation.
The resulting patch tokens exhibit viewpoint invariance that is competitive with 3D-supervised models. Under camera shifts where standard 2D representations experience performance drops, SNAP retains spatial correspondence across diverse evaluation scenes.
Predicting frozen pre-trained patch features (from models like DINO, MAE, or SAM) rather than raw RGB pixels reduces sensitivity to low-level illumination and texture variations, encouraging representations aligned with geometric structure.
Rather than relying on spatially expressive decoders with global receptive fields, SNAP uses a lightweight cross-attention decoder with localized interactions, allowing multi-view spatial aggregation to reside primarily within the scene encoder representations.
Without explicit 3D annotations or multi-view robotic demonstrations, the learned representations exhibit viewpoint invariance across camera shifts ($0.2$ to $0.6$ rad), maintaining policy performance across out-of-distribution viewpoints.
SNAP improves geometric representation quality when paired with different teacher backbones (DINOv3, SAM, MAE) across dense correspondence, depth, and pose benchmarks.
An asymmetric encoder-decoder design that decouples multi-view scene aggregation from target view decoding.
Processes $N$ unposed reference views through a frozen ViT featurizer $\phi$ (e.g. DINOv3 ViT-B/16) to produce patch features $F_k = \phi(I_k) \in \mathbb{R}^{L \times d}$. Notably, no camera parameters are provided to the encoder at any stage.
A 16-layer transformer with full cross-view self-attention aggregates patch tokens across all views with 2D rotary position embeddings (RoPE2D), outputting a multi-view scene latent $Z \in \mathbb{R}^{NL \times E}$.
Target patch queries attend to $Z$ using PRoPE-conditioned cross-attention based on context cameras $\{P_k\}$ and target camera $\{P_j^*\}$:
Following cross-attention, lateral communication between target queries is constrained via a BlockNN mask with $RF = 1\times 1$, where self-attention operates as a pointwise residual MLP across individual query locations.
PRoPE conditions cross-attention on extrinsic camera poses $T_k \in \mathrm{SE}(3)$ and intrinsics $K_k$ by projecting normalized 3D viewing rays without breaking geometric equivariance:
Geometric Intuition: Transforming queries via $(P_j^*)^\top q$ and keys via $P_k^{-1} k$ co-aligns rays in world space. Attention scores are maximized when target query rays and reference context rays intersect along epipolar lines, executing a continuous, differentiable soft-epipolar search.
In the target decoder, self-attention among target queries is partitioned into non-overlapping spatial windows of size $k \times k$ via an attention mask $\mathbf{M} \in \{0, -\infty\}^{L \times L}$:
The Critical Crossover ($k=1$): At $RF = 1 \times 1$, $\mathbf{M}_{u, v} = -\infty$ for all $u \neq v$, reducing self-attention to an identity operation. This removes lateral spatial mixing among target queries, ensuring target token synthesis relies directly on cross-attention queries into the multi-view scene latent $Z$.
Supervision is applied directly against normalized teacher features rather than RGB pixels, combining an L1 feature loss with a spatial finite-difference gradient penalty:
The spatial gradient penalty $\mathcal{L}_{\text{grad}}$ promotes sharp spatial fidelity and feature consistency at physical object boundaries.
PCA projections of learned features across novel camera trajectories and dynamic video sequences.
Select a sequence below to inspect synthesized feature PCA video synchronized with 2D camera trajectory and conditioning frames.
SNAP operates across diverse visual latent distributions without modification. Below, compare ground truth RGB video with feature PCAs trained on three distinct teacher objectives: DINOv3 (self-distillation), MAE (masked autoencoding), and SAM (segmentation supervision).
Evaluation of frozen encoder representations across five geometric and robotics tasks without fine-tuning.
Evaluating frozen representations via lightweight task heads across point correspondence, visual localization, relative pose estimation, and relative depth.
Pixel-level matching precision across RealEstate10K, DL3DV, and synthetic Hypersim evaluated against RoMa v2 pseudo-ground truth correspondences. Note that models like VGGT and LagerNVS have seen Hypersim during pre-training. Higher PCK and lower APE indicate more accurate point matching across novel camera viewpoints.
| Method | Supervision | RealEstate10K | DL3DV | Hypersim | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| @5px ↑ | @10px ↑ | APE ↓ | @5px ↑ | @10px ↑ | APE ↓ | @5px ↑ | @10px ↑ | APE ↓ | ||
| LVSM | NVS Supervised | 5.51 | 14.76 | 41.04 | 3.53 | 11.05 | 45.75 | 5.62 | 13.41 | 42.15 |
| RayZer | NVS Supervised | 3.81 | 13.53 | 39.63 | 2.73 | 13.53 | 47.03 | 4.21 | 11.21 | 44.82 |
| ERayZer | NVS Supervised | 4.59 | 12.47 | 39.67 | 3.55 | 11.26 | 45.30 | 4.31 | 14.89 | 40.54 |
| LagerNVS | NVS Supervised | 9.61 | 27.36 | 26.54 | 3.81 | 11.31 | 42.41 | 8.94 | 24.30 | 31.02 |
| Muskie | Masked Reconstruction | 12.09 | 34.43 | 20.90 | 3.58 | 11.39 | 41.78 | 12.71 | 34.15 | 21.15 |
| VGGT | Geometry Supervised | 12.85 | 32.78 | 21.79 | 4.15 | 13.65 | 39.34 | 14.92 | 35.38 | 19.14 |
| SNAP (Ours) | NVS Supervised | 15.48 | 36.29 | 18.41 | 4.20 | 13.33 | 39.26 | 15.01 | 34.87 | 19.82 |
Relative depth estimation reporting absolute relative error (AbsRel ↓) and threshold accuracy $\delta_1$ ↑ ($\max(d/\hat{d}, \hat{d}/d) < 1.25$). Models are evaluated on Hypersim and the held-out real-world 7Scenes dataset; notably, models like VGGT and LagerNVS have seen Hypersim during pre-training, whereas 7Scenes provides a held-out real-world benchmark.
| Method | Supervision | Hypersim | 7Scenes (Held-out Real-World) | ||
|---|---|---|---|---|---|
| AbsRel ↓ | $\delta_1$ (%) ↑ | AbsRel ↓ | $\delta_1$ (%) ↑ | ||
| VGGT (Geometry supervised; Hypersim seen in pre-training) | Geometry Supervised | 0.2080 | 94.60% | 1.9138 | 70.10% |
| LVSM | NVS Supervised | 2.0648 | 38.51% | 1.9474 | 32.57% |
| LagerNVS (Hypersim seen in training) | NVS Supervised | 0.7441 | 73.11% | 2.3858 | 49.27% |
| ERayZer | NVS Supervised | 1.1321 | 50.61% | 1.3320 | 35.57% |
| SNAP (Ours) | NVS Supervised | 0.7093 | 64.68% | 0.9257 | 52.02% |
Evaluating policy success rate (%) on Franka Panda arm tasks (Mimicgen benchmark) under systematic out-of-distribution camera shifts from 0.2 to 0.6 radians.
Dense point correspondence (PCK %) evaluated on a subset of Hypersim correspondences comparing the SNAP encoder against its teacher across 3 foundation encoders (DINOv3, SAM, MAE) and 3 distance thresholds ($\alpha = 0.1, 0.2, 0.3$).
Subsampled scaling analysis characterizing model parameters, peak GPU memory footprint (GB), and forward-pass runtime (ms) across view milestones ($V \in \{1, 10, 50, 100\}$). Benchmarked under identical conditions on an NVIDIA RTX PRO 6000 Blackwell GPU (bfloat16).
| Method | Supervision | Params | Metric | Input Views ($V$) | |||
|---|---|---|---|---|---|---|---|
| $V=1$ | $V=10$ | $V=50$ | $V=100$ | ||||
| Muskie | Masked Reconstruction | 303.1 M | VRAM (GB) | 0.59 | 0.70 | 1.16 | 1.73 |
| Latency (ms) | 12 | 16 | 67 | 180 | |||
| VGGT | Geometry Supervised | 909.1 M | VRAM (GB) | 2.18 | 2.29 | 2.79 | 3.41 |
| Latency (ms) | 60 | 89 | 294 | 712 | |||
| SNAP (Ours) | NVS Supervised | 233.4 M | VRAM (GB) | 0.45 | 0.51 | 0.74 | 1.03 |
| Latency (ms) | 20 | 22 | 85 | 245 | |||
Efficiency Summary: Parameter count, peak GPU memory footprint, and forward-pass runtime measured across context view counts $V \in \{1, 10, 50, 100\}$. At $V=100$, SNAP operates with 233.38M parameters, 1.03 GB peak VRAM, and 245 ms runtime.
@misc{snap2026,
title = {Less Decoder is More Encoder: Geometric Representation
Learning from Novel View Synthesis},
author = {Keerthi Kaashyap and Dennis Anthony and Akshay Krishnan and
Nhi Ngoc Nguyen and Jeremy Collins and James Hays and
Shreyas Kousik and Animesh Garg},
year = {2026},
note = {NeurIPS 2026}
}