Resolving Representation Ambiguity in Feedforward Novel View Synthesis Transformer via Semantic-Spatial Decoupling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yihang, Sun, Yihang, Zhang, Shaofeng, Wu, Zuxuan, Yan, Junchi, Jia, Xiaosong, Jiang, Yu-gang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917508249812992
author Wu, Yihang
Sun, Yihang
Zhang, Shaofeng
Wu, Zuxuan
Yan, Junchi
Jia, Xiaosong
Jiang, Yu-gang
author_facet Wu, Yihang
Sun, Yihang
Zhang, Shaofeng
Wu, Zuxuan
Yan, Junchi
Jia, Xiaosong
Jiang, Yu-gang
contents Transformer-based models have advanced feedforward novel view synthesis (NVS). Current architectures such as GS-LRM and LVSM mix semantic information (e.g., RGB) and spatial information (e.g., Plücker rays) into a shared feature space. Since Plücker rays naturally carry lattice-like spatial structure, these designs can make the spatial bias interfere with appearance representation and degrade rendering fidelity. To this end, we propose to decouple the representation of feedforward NVS transformers into separate semantic and spatial tokens. The decoupled design keeps semantic and spatial information explicit in their branches while preserving cross-branch interaction through shared attention routing. Built on this design, we introduce optional categorized supervision and bidirectional modulation: the former provides branch-specific training signals, while the latter improves interaction between the two branches. Notably, the base decoupled design introduces virtually zero additional inference latency due to its architectural design. The proposed designs achieve consistent improvements, demonstrating effectiveness across decoder-only and encoder-decoder feedforward NVS models.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18599
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Resolving Representation Ambiguity in Feedforward Novel View Synthesis Transformer via Semantic-Spatial Decoupling
Wu, Yihang
Sun, Yihang
Zhang, Shaofeng
Wu, Zuxuan
Yan, Junchi
Jia, Xiaosong
Jiang, Yu-gang
Computer Vision and Pattern Recognition
Transformer-based models have advanced feedforward novel view synthesis (NVS). Current architectures such as GS-LRM and LVSM mix semantic information (e.g., RGB) and spatial information (e.g., Plücker rays) into a shared feature space. Since Plücker rays naturally carry lattice-like spatial structure, these designs can make the spatial bias interfere with appearance representation and degrade rendering fidelity. To this end, we propose to decouple the representation of feedforward NVS transformers into separate semantic and spatial tokens. The decoupled design keeps semantic and spatial information explicit in their branches while preserving cross-branch interaction through shared attention routing. Built on this design, we introduce optional categorized supervision and bidirectional modulation: the former provides branch-specific training signals, while the latter improves interaction between the two branches. Notably, the base decoupled design introduces virtually zero additional inference latency due to its architectural design. The proposed designs achieve consistent improvements, demonstrating effectiveness across decoder-only and encoder-decoder feedforward NVS models.
title Resolving Representation Ambiguity in Feedforward Novel View Synthesis Transformer via Semantic-Spatial Decoupling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.18599