Resolving Representation Ambiguity in Feedforward Novel View Synthesis Transformer via Semantic-Spatial Decoupling
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Yihang, Sun, Yihang, Zhang, Shaofeng, Wu, Zuxuan, Yan, Junchi, Jia, Xiaosong, Jiang, Yu-gang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient-LVSM: Faster, Cheaper, and Better Large View Synthesis Model via Decoupled Co-Refinement Attention
by: Jia, Xiaosong, et al.
Published: (2026)
by: Jia, Xiaosong, et al.
Published: (2026)
Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
by: Zhang, Xiangdong, et al.
Published: (2025)
by: Zhang, Xiangdong, et al.
Published: (2025)
FRUC: Feedforward Dynamic Scene Reconstruction from Uncalibrated Collaborative Driving Views
by: Tao, Yihang, et al.
Published: (2026)
by: Tao, Yihang, et al.
Published: (2026)
Attention Itself Could Retrieve.RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval
by: Zou, Zichen, et al.
Published: (2026)
by: Zou, Zichen, et al.
Published: (2026)
Fast Feedforward 3D Gaussian Splatting Compression
by: Chen, Yihang, et al.
Published: (2024)
by: Chen, Yihang, et al.
Published: (2024)
DriveTransformer: Unified Transformer for Scalable End-to-End Autonomous Driving
by: Jia, Xiaosong, et al.
Published: (2025)
by: Jia, Xiaosong, et al.
Published: (2025)
DTGBrepGen: A Novel B-rep Generative Model through Decoupling Topology and Geometry
by: Li, Jing, et al.
Published: (2025)
by: Li, Jing, et al.
Published: (2025)
Spatial Retrieval Augmented Autonomous Driving
by: Jia, Xiaosong, et al.
Published: (2025)
by: Jia, Xiaosong, et al.
Published: (2025)
PromptFusion: Decoupling Stability and Plasticity for Continual Learning
by: Chen, Haoran, et al.
Published: (2023)
by: Chen, Haoran, et al.
Published: (2023)
FlatFusion: Delving into Details of Sparse Transformer-based Camera-LiDAR Fusion for Autonomous Driving
by: Zhu, Yutao, et al.
Published: (2024)
by: Zhu, Yutao, et al.
Published: (2024)
Impact of domain adaptation in deep learning for medical image classifications
by: Wu, Yihang, et al.
Published: (2026)
by: Wu, Yihang, et al.
Published: (2026)
Federated CLIP for Resource-Efficient Heterogeneous Medical Image Classification
by: Wu, Yihang, et al.
Published: (2025)
by: Wu, Yihang, et al.
Published: (2025)
BevSplat: Resolving Height Ambiguity via Feature-Based Gaussian Primitives for Weakly-Supervised Cross-View Localization
by: Wang, Qiwei, et al.
Published: (2025)
by: Wang, Qiwei, et al.
Published: (2025)
DeRA: Decoupled Representation Alignment for Video Tokenization
by: Guo, Pengbo, et al.
Published: (2025)
by: Guo, Pengbo, et al.
Published: (2025)
PCP-MAE: Learning to Predict Centers for Point Masked Autoencoders
by: Zhang, Xiangdong, et al.
Published: (2024)
by: Zhang, Xiangdong, et al.
Published: (2024)
Can Users Specify Driving Speed? Bench2Drive-Speed: Benchmark and Baselines for Desired-Speed Conditioned Autonomous Driving
by: Shao, Yuqian, et al.
Published: (2026)
by: Shao, Yuqian, et al.
Published: (2026)
Image Synthesis with Class-Aware Semantic Diffusion Models for Surgical Scene Segmentation
by: Zhou, Yihang, et al.
Published: (2024)
by: Zhou, Yihang, et al.
Published: (2024)
DriveVGGT: Calibration-Constrained Visual Geometry Transformers for Multi-Camera Autonomous Driving
by: Jia, Xiaosong, et al.
Published: (2025)
by: Jia, Xiaosong, et al.
Published: (2025)
Dual-Branch Center-Surrounding Contrast: Rethinking Contrastive Learning for 3D Point Clouds
by: Zhang, Shaofeng, et al.
Published: (2025)
by: Zhang, Shaofeng, et al.
Published: (2025)
Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank
by: Zhang, Shaofeng, et al.
Published: (2025)
by: Zhang, Shaofeng, et al.
Published: (2025)
Exploration of Reproducible Generated Image Detection
by: Duan, Yihang
Published: (2025)
by: Duan, Yihang
Published: (2025)
Semi-Supervised Medical Image Segmentation via Dual Networks
by: Lu, Yunyao, et al.
Published: (2025)
by: Lu, Yunyao, et al.
Published: (2025)
Towards Ambiguity-Free Spatial Foundation Model: Rethinking and Decoupling Depth Ambiguity
by: Xu, Xiaohao, et al.
Published: (2025)
by: Xu, Xiaohao, et al.
Published: (2025)
Deep Modeling and Interpretation for Bladder Cancer Classification
by: Chaddad, Ahmad, et al.
Published: (2026)
by: Chaddad, Ahmad, et al.
Published: (2026)
Federated Vision Transformer with Adaptive Focal Loss for Medical Image Classification
by: Zhao, Xinyuan, et al.
Published: (2026)
by: Zhao, Xinyuan, et al.
Published: (2026)
PointAlign: Feature-Level Alignment Regularization for 3D Vision-Language Models
by: Su, Yuanhao, et al.
Published: (2026)
by: Su, Yuanhao, et al.
Published: (2026)
Classification based deep learning models for lung cancer and disease using medical images
by: Chaddad, Ahmad, et al.
Published: (2025)
by: Chaddad, Ahmad, et al.
Published: (2025)
Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving
by: Jia, Xiaosong, et al.
Published: (2024)
by: Jia, Xiaosong, et al.
Published: (2024)
FACMIC: Federated Adaptative CLIP Model for Medical Image Classification
by: Wu, Yihang, et al.
Published: (2024)
by: Wu, Yihang, et al.
Published: (2024)
GMGaze: MoE-Based Context-Aware Gaze Estimation with CLIP and Multiscale Transformer
by: Zhao, Xinyuan, et al.
Published: (2026)
by: Zhao, Xinyuan, et al.
Published: (2026)
Learning Accurate Segmentation Purely from Self-Supervision
by: You, Zuyao, et al.
Published: (2026)
by: You, Zuyao, et al.
Published: (2026)
Feedforward 3D Editing Learns from Semantic-Part Transformation
by: Weng, Jiawei, et al.
Published: (2026)
by: Weng, Jiawei, et al.
Published: (2026)
Bench2Drive-R: Turning Real World Data into Reactive Closed-Loop Autonomous Driving Benchmark by Generative Model
by: You, Junqi, et al.
Published: (2024)
by: You, Junqi, et al.
Published: (2024)
Identifiable Object Representations under Spatial Ambiguities
by: Kori, Avinash, et al.
Published: (2025)
by: Kori, Avinash, et al.
Published: (2025)
DLEN: Dual Branch of Transformer for Low-Light Image Enhancement in Dual Domains
by: Xia, Junyu, et al.
Published: (2025)
by: Xia, Junyu, et al.
Published: (2025)
GeoGS3D: Single-view 3D Reconstruction via Geometric-aware Diffusion Model and Gaussian Splatting
by: Feng, Qijun, et al.
Published: (2024)
by: Feng, Qijun, et al.
Published: (2024)
FreqPDE: Rethinking Positional Depth Embedding for Multi-View 3D Object Detection Transformers
by: Su, Haisheng, et al.
Published: (2025)
by: Su, Haisheng, et al.
Published: (2025)
Novel View Synthesis as Video Completion
by: Wu, Qi, et al.
Published: (2026)
by: Wu, Qi, et al.
Published: (2026)
Evading Visual Aphasia: Contrastive Adaptive Semantic Token Pruning for Vision-Language Models
by: Ma, Jie, et al.
Published: (2026)
by: Ma, Jie, et al.
Published: (2026)
Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue
by: Lin, Xingyao, et al.
Published: (2025)
by: Lin, Xingyao, et al.
Published: (2025)
Similar Items
-
Efficient-LVSM: Faster, Cheaper, and Better Large View Synthesis Model via Decoupled Co-Refinement Attention
by: Jia, Xiaosong, et al.
Published: (2026) -
Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
by: Zhang, Xiangdong, et al.
Published: (2025) -
FRUC: Feedforward Dynamic Scene Reconstruction from Uncalibrated Collaborative Driving Views
by: Tao, Yihang, et al.
Published: (2026) -
Attention Itself Could Retrieve.RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval
by: Zou, Zichen, et al.
Published: (2026) -
Fast Feedforward 3D Gaussian Splatting Compression
by: Chen, Yihang, et al.
Published: (2024)