Less Decoder is More Encoder:
Geometric Representation Learning from Novel View Synthesis

Keerthi Kaashyap Dennis Anthony Akshay Krishnan Nhi Ngoc Nguyen Jeremy Collins James Hays Shreyas Kousik Animesh Garg
Georgia Institute of Technology
NeurIPS 2026
SNAP Teaser: Paradigm comparison between Traditional novel view synthesis and SNAP, alongside multi-view geometric correspondence tracking across wide baseline camera shifts.

Figure 1: Left: Traditional novel view synthesis (NVS) encodes context views and renders novel views with a pose-conditioned renderer trained with an RGB reconstruction loss. SNAP instead replaces the renderer with a pose-conditioned local decoder trained with a loss in semantic DINO feature space, encouraging the encoder to learn richer visual features. Right: Point trajectories across three scenes, computed from frozen Multi-View Scene Encoder features, with predictions in red and ground truth in green. SNAP's tracks stay close to ground truth, while those from traditional RGB-supervised NVS pipelines drift.

Core Finding

Reconstruction fidelity alone is a deceptive proxy for representation quality.

Conventional novel view synthesis methods often pair high-capacity decoders with pixel-level reconstruction objectives that focus optimization on low-level photometric details. SNAP explores an alternative design with two components: predicting pre-trained feature tokens (from models like DINO, MAE, or SAM) instead of raw pixels, and using a lightweight cross-attention decoder with localized receptive fields. Under this design, the learned representations exhibit improved transfer across downstream 3D geometric benchmarks.

Abstract

This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS requires reasoning about 3D scene structure, potentially enabling transferable multi-view geometric representations. However, representations learned from existing encoder-based NVS methods often show limited transfer to downstream geometric tasks. We investigate two contributing factors: spatially expressive decoders that can absorb multi-view spatial reasoning away from the scene encoder, and low-level pixel-space targets that emphasize high-frequency photometric details over geometric structure.

We present SNAP, a self-supervised encoder-decoder transformer designed with a pose-conditioned local decoder and a latent-space reconstruction objective. SNAP is general-purpose and evaluates across five geometric tasks: visual localization, pose estimation, point correspondence, depth estimation, and robot manipulation.

The resulting patch tokens exhibit viewpoint invariance that is competitive with 3D-supervised models. Under camera shifts where standard 2D representations experience performance drops, SNAP retains spatial correspondence across diverse evaluation scenes.

01

Feature-Space Prediction

Predicting frozen pre-trained patch features (from models like DINO, MAE, or SAM) rather than raw RGB pixels reduces sensitivity to low-level illumination and texture variations, encouraging representations aligned with geometric structure.

02

Cross Attention Decoder

Rather than relying on spatially expressive decoders with global receptive fields, SNAP uses a lightweight cross-attention decoder with localized interactions, allowing multi-view spatial aggregation to reside primarily within the scene encoder representations.

03

Emergent Viewpoint Invariance

Without explicit 3D annotations or multi-view robotic demonstrations, the learned representations exhibit viewpoint invariance across camera shifts ($0.2$ to $0.6$ rad), maintaining policy performance across out-of-distribution viewpoints.

04

Generality Across Diverse Teachers

SNAP improves geometric representation quality when paired with different teacher backbones (DINOv3, SAM, MAE) across dense correspondence, depth, and pose benchmarks.

Method

An asymmetric encoder-decoder design that decouples multi-view scene aggregation from target view decoding.

SNAP Architecture Diagram

1. Pose-Free Multi-View Scene Encoder

Processes $N$ unposed reference views through a frozen ViT featurizer $\phi$ (e.g. DINOv3 ViT-B/16) to produce patch features $F_k = \phi(I_k) \in \mathbb{R}^{L \times d}$. Notably, no camera parameters are provided to the encoder at any stage.

A 16-layer transformer with full cross-view self-attention aggregates patch tokens across all views with 2D rotary position embeddings (RoPE2D), outputting a multi-view scene latent $Z \in \mathbb{R}^{NL \times E}$.

2. Pose-Conditioned Local Decoder ($RF = 1\times 1$)

Target patch queries attend to $Z$ using PRoPE-conditioned cross-attention based on context cameras $\{P_k\}$ and target camera $\{P_j^*\}$:

$$\mathrm{Attn}(q, k, v) = \mathrm{softmax}\left( \frac{(P_j^{*\top} q)(P_k^{-1} k)^\top}{\sqrt{E/H}} \right) v$$

Following cross-attention, lateral communication between target queries is constrained via a BlockNN mask with $RF = 1\times 1$, where self-attention operates as a pointwise residual MLP across individual query locations.

Technical Deep Dive

Mathematical Formalization: PRoPE Mechanism & BlockNN Bottleneck Mask

1. Perspective RoPE (PRoPE) & Soft-Epipolar Search

PRoPE conditions cross-attention on extrinsic camera poses $T_k \in \mathrm{SE}(3)$ and intrinsics $K_k$ by projecting normalized 3D viewing rays without breaking geometric equivariance:

$$P_k = \hat{K}_k T_k^{-1}, \quad P_j^* = \hat{K}_j^* (T_j^*)^{-1}$$ $$\mathrm{Attn}(q, k, v) = \mathrm{softmax}\left( \frac{\left( (P_j^*)^\top q \right) \left( P_k^{-1} k \right)^\top}{\sqrt{E/H}} \right) v$$

Geometric Intuition: Transforming queries via $(P_j^*)^\top q$ and keys via $P_k^{-1} k$ co-aligns rays in world space. Attention scores are maximized when target query rays and reference context rays intersect along epipolar lines, executing a continuous, differentiable soft-epipolar search.

2. BlockNN Mask & The $RF = 1\times 1$ Bottleneck

In the target decoder, self-attention among target queries is partitioned into non-overlapping spatial windows of size $k \times k$ via an attention mask $\mathbf{M} \in \{0, -\infty\}^{L \times L}$:

$$\mathbf{M}_{u, v} = \begin{cases} 0 & \text{if } \lfloor \mathbf{x}_u/k \rfloor = \lfloor \mathbf{x}_v/k \rfloor \text{ and } \lfloor \mathbf{y}_u/k \rfloor = \lfloor \mathbf{y}_v/k \rfloor \\ -\infty & \text{otherwise} \end{cases}$$

The Critical Crossover ($k=1$): At $RF = 1 \times 1$, $\mathbf{M}_{u, v} = -\infty$ for all $u \neq v$, reducing self-attention to an identity operation. This removes lateral spatial mixing among target queries, ensuring target token synthesis relies directly on cross-attention queries into the multi-view scene latent $Z$.

Latent Reconstruction Objective

Supervision is applied directly against normalized teacher features rather than RGB pixels, combining an L1 feature loss with a spatial finite-difference gradient penalty:

$$\mathcal{L}_{\text{recon}} = \frac{1}{Ld} \sum_{p,d} \left| \hat{F}_{p,d} - F^*_{p,d} \right| + \frac{1}{2} \left( \|\nabla_x \hat{F} - \nabla_x F^*\|_1 + \|\nabla_y \hat{F} - \nabla_y F^*\|_1 \right)$$

The spatial gradient penalty $\mathcal{L}_{\text{grad}}$ promotes sharp spatial fidelity and feature consistency at physical object boundaries.

Qualitative Feature Map Visualizations

PCA projections of learned features across novel camera trajectories and dynamic video sequences.

Novel View Feature PCA & Camera Trajectory

Select a sequence below to inspect synthesized feature PCA video synchronized with 2D camera trajectory and conditioning frames.

Breakfast Bar Novel View Feature PCA
0%
Top-Down Camera Trajectory Click / Drag to Scrub
Context Cameras (C1–C4)
Rendered Pathway Frustum
Conditioning Context Views (1–4): Click to seek

Consistent Feature Quality Across Alternative Teachers

SNAP operates across diverse visual latent distributions without modification. Below, compare ground truth RGB video with feature PCAs trained on three distinct teacher objectives: DINOv3 (self-distillation), MAE (masked autoencoding), and SAM (segmentation supervision).

Input Video Ground Truth RGB
SNAP Teacher: DINOv3
SNAP Teacher: MAE
SNAP Teacher: SAM
Select Evaluation Sequence (10 Sequences • Click to switch, hover to preview)
Synchronized 4-Way Stream

Downstream Benchmark Evaluations

Evaluation of frozen encoder representations across five geometric and robotics tasks without fine-tuning.

1. Dense 3D Geometry Benchmarks

Evaluating frozen representations via lightweight task heads across point correspondence, visual localization, relative pose estimation, and relative depth.

Table 1: Dense Point Correspondence

PCK ↑ (Percentage of Correct Keypoints) & APE ↓ (Average Position Error)

Pixel-level matching precision across RealEstate10K, DL3DV, and synthetic Hypersim evaluated against RoMa v2 pseudo-ground truth correspondences. Note that models like VGGT and LagerNVS have seen Hypersim during pre-training. Higher PCK and lower APE indicate more accurate point matching across novel camera viewpoints.

Method Supervision RealEstate10K DL3DV Hypersim
@5px â†‘ @10px â†‘ APE â†“ @5px â†‘ @10px â†‘ APE â†“ @5px â†‘ @10px â†‘ APE â†“
LVSM NVS Supervised 5.51 14.76 41.04 3.53 11.05 45.75 5.62 13.41 42.15
RayZer NVS Supervised 3.81 13.53 39.63 2.73 13.53 47.03 4.21 11.21 44.82
ERayZer NVS Supervised 4.59 12.47 39.67 3.55 11.26 45.30 4.31 14.89 40.54
LagerNVS NVS Supervised 9.61 27.36 26.54 3.81 11.31 42.41 8.94 24.30 31.02
Muskie Masked Reconstruction 12.09 34.43 20.90 3.58 11.39 41.78 12.71 34.15 21.15
VGGT Geometry Supervised 12.85 32.78 21.79 4.15 13.65 39.34 14.92 35.38 19.14
SNAP (Ours) NVS Supervised 15.48 36.29 18.41 4.20 13.33 39.26 15.01 34.87 19.82

Table 2: Relative Depth Estimation

Evaluated on Hypersim validation split and held-out real-world 7Scenes

Relative depth estimation reporting absolute relative error (AbsRel ↓) and threshold accuracy $\delta_1$ ↑ ($\max(d/\hat{d}, \hat{d}/d) < 1.25$). Models are evaluated on Hypersim and the held-out real-world 7Scenes dataset; notably, models like VGGT and LagerNVS have seen Hypersim during pre-training, whereas 7Scenes provides a held-out real-world benchmark.

Method Supervision Hypersim 7Scenes (Held-out Real-World)
AbsRel ↓ $\delta_1$ (%) ↑ AbsRel ↓ $\delta_1$ (%) ↑
VGGT (Geometry supervised; Hypersim seen in pre-training) Geometry Supervised 0.2080 94.60% 1.9138 70.10%
LVSM NVS Supervised 2.0648 38.51% 1.9474 32.57%
LagerNVS (Hypersim seen in training) NVS Supervised 0.7441 73.11% 2.3858 49.27%
ERayZer NVS Supervised 1.1321 50.61% 1.3320 35.57%
SNAP (Ours) NVS Supervised 0.7093 64.68% 0.9257 52.02%

2. Emergent Viewpoint Invariance in Robot Manipulation

Evaluating policy success rate (%) on Franka Panda arm tasks (Mimicgen benchmark) under systematic out-of-distribution camera shifts from 0.2 to 0.6 radians.

Select Evaluation Task:
Observation: Policy performance using 2D representations (DINOv3) decreases under increasing camera shifts. In contrast, SNAP retains policy performance across viewpoint variations of 0.2–0.6 rad, performing competitively with 3D-supervised models (VGGT) and multi-view baselines without pre-training on robotics demonstrations.

3. Consistent Improvement Across Diverse Teachers

Dense point correspondence (PCK %) evaluated on a subset of Hypersim correspondences comparing the SNAP encoder against its teacher across 3 foundation encoders (DINOv3, SAM, MAE) and 3 distance thresholds ($\alpha = 0.1, 0.2, 0.3$).

Group Layout:
Frozen Teacher
SNAP Gain (+Δ%)
Observation (Hypersim): Across all three PCK thresholds on Hypersim, SNAP improves dense point correspondence compared to the frozen teacher representations (+4.1% to +26.1% on DINOv3, +4.9% to +11.5% on SAM, and +23.4% to +39.7% on MAE).

4. Model Efficiency & Scaling Analysis

Subsampled scaling analysis characterizing model parameters, peak GPU memory footprint (GB), and forward-pass runtime (ms) across view milestones ($V \in \{1, 10, 50, 100\}$). Benchmarked under identical conditions on an NVIDIA RTX PRO 6000 Blackwell GPU (bfloat16).

Method Supervision Params Metric Input Views ($V$)
$V=1$ $V=10$ $V=50$ $V=100$
Muskie Masked Reconstruction 303.1 M VRAM (GB) 0.59 0.70 1.16 1.73
Latency (ms) 12 16 67 180
VGGT Geometry Supervised 909.1 M VRAM (GB) 2.18 2.29 2.79 3.41
Latency (ms) 60 89 294 712
SNAP (Ours) NVS Supervised 233.4 M VRAM (GB) 0.45 0.51 0.74 1.03
Latency (ms) 20 22 85 245

Efficiency Summary: Parameter count, peak GPU memory footprint, and forward-pass runtime measured across context view counts $V \in \{1, 10, 50, 100\}$. At $V=100$, SNAP operates with 233.38M parameters, 1.03 GB peak VRAM, and 245 ms runtime.

BibTeX

@misc{snap2026,
  title  = {Less Decoder is More Encoder: Geometric Representation
             Learning from Novel View Synthesis},
  author = {Keerthi Kaashyap and Dennis Anthony and Akshay Krishnan and
             Nhi Ngoc Nguyen and Jeremy Collins and James Hays and
             Shreyas Kousik and Animesh Garg},
  year   = {2026},
  note   = {NeurIPS 2026}
}