Scaling View Synthesis Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Evan, Ryu, Hyunwoo, Mitchel, Thomas W., Sitzmann, Vincent |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
True Self-Supervised Novel View Synthesis is Transferable
by: Mitchel, Thomas W., et al.
Published: (2025)
by: Mitchel, Thomas W., et al.
Published: (2025)
Neural Isometries: Taming Transformations for Equivariant ML
by: Mitchel, Thomas W., et al.
Published: (2024)
by: Mitchel, Thomas W., et al.
Published: (2024)
Understanding Multi-View Transformers
by: Stary, Michal, et al.
Published: (2025)
by: Stary, Michal, et al.
Published: (2025)
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
by: Cazenavette, George, et al.
Published: (2025)
by: Cazenavette, George, et al.
Published: (2025)
Bridging Implicit and Explicit Geometric Transformation for Single-Image View Synthesis
by: Park, Byeongjun, et al.
Published: (2022)
by: Park, Byeongjun, et al.
Published: (2022)
Turning Video Models into Generalist Robot Policies
by: Li, Sizhe Lester, et al.
Published: (2026)
by: Li, Sizhe Lester, et al.
Published: (2026)
DT-NVS: Diffusion Transformers for Novel View Synthesis
by: Jang, Wonbong, et al.
Published: (2025)
by: Jang, Wonbong, et al.
Published: (2025)
Cross-aware Early Fusion with Stage-divided Vision and Language Transformer Encoders for Referring Image Segmentation
by: Cho, Yubin, et al.
Published: (2024)
by: Cho, Yubin, et al.
Published: (2024)
DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO
by: Park, Jinyoung, et al.
Published: (2025)
by: Park, Jinyoung, et al.
Published: (2025)
Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis
by: Menapace, Willi, et al.
Published: (2024)
by: Menapace, Willi, et al.
Published: (2024)
DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis
by: Lu, Cheng-You, et al.
Published: (2026)
by: Lu, Cheng-You, et al.
Published: (2026)
3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View Transformation
by: Oh, Gyeongrok, et al.
Published: (2025)
by: Oh, Gyeongrok, et al.
Published: (2025)
Intrinsic Image Diffusion for Indoor Single-view Material Estimation
by: Kocsis, Peter, et al.
Published: (2023)
by: Kocsis, Peter, et al.
Published: (2023)
Diversity-Driven View Subset Selection for Indoor Novel View Synthesis
by: Wang, Zehao, et al.
Published: (2024)
by: Wang, Zehao, et al.
Published: (2024)
Dynamic View Synthesis as an Inverse Problem
by: Yesiltepe, Hidir, et al.
Published: (2025)
by: Yesiltepe, Hidir, et al.
Published: (2025)
WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
by: Yu, Seungjun, et al.
Published: (2025)
by: Yu, Seungjun, et al.
Published: (2025)
Error as Signal: Stiffness-Aware Diffusion Sampling via Embedded Runge-Kutta Guidance
by: Kong, Inho, et al.
Published: (2026)
by: Kong, Inho, et al.
Published: (2026)
Learning Neural Exposure Fields for View Synthesis
by: Niemeyer, Michael, et al.
Published: (2025)
by: Niemeyer, Michael, et al.
Published: (2025)
View-Invariant Pixelwise Anomaly Detection in Multi-object Scenes with Adaptive View Synthesis
by: Varghese, Subin, et al.
Published: (2024)
by: Varghese, Subin, et al.
Published: (2024)
CycleBEV: Regularizing View Transformation Networks via View Cycle Consistency for Bird's-Eye-View Semantic Segmentation
by: Hong, Jeongbin, et al.
Published: (2026)
by: Hong, Jeongbin, et al.
Published: (2026)
OmniView: An All-Seeing Diffusion Model for 3D and 4D View Synthesis
by: Fan, Xiang, et al.
Published: (2025)
by: Fan, Xiang, et al.
Published: (2025)
An Initial Study of Bird's-Eye View Generation for Autonomous Vehicles using Cross-View Transformers
by: Santos, Felipe Carlos dos, et al.
Published: (2025)
by: Santos, Felipe Carlos dos, et al.
Published: (2025)
360 in the Wild: Dataset for Depth Prediction and View Synthesis
by: Park, Kibaek, et al.
Published: (2024)
by: Park, Kibaek, et al.
Published: (2024)
Aria-NeRF: Multimodal Egocentric View Synthesis
by: Sun, Jiankai, et al.
Published: (2023)
by: Sun, Jiankai, et al.
Published: (2023)
Human Multi-View Synthesis from a Single-View Model:Transferred Body and Face Representations
by: Feng, Yu, et al.
Published: (2024)
by: Feng, Yu, et al.
Published: (2024)
PhysGaia: A Physics-Aware Benchmark with Multi-Body Interactions for Dynamic Novel View Synthesis
by: Kim, Mijeong, et al.
Published: (2025)
by: Kim, Mijeong, et al.
Published: (2025)
Fast and Lightweight Novel View Synthesis with Differentiable Multiplane Image
by: Zhang, Kaidi, et al.
Published: (2026)
by: Zhang, Kaidi, et al.
Published: (2026)
Memory Efficient Full-gradient Attacks (MEFA) Framework for Adversarial Defense Evaluations
by: Du, Yuan, et al.
Published: (2026)
by: Du, Yuan, et al.
Published: (2026)
Vision Bridge Transformer at Scale
by: Tan, Zhenxiong, et al.
Published: (2025)
by: Tan, Zhenxiong, et al.
Published: (2025)
Enhancing Close-up Novel View Synthesis via Pseudo-labeling
by: Xia, Jiatong, et al.
Published: (2025)
by: Xia, Jiatong, et al.
Published: (2025)
Learning Novel View Synthesis from Heterogeneous Low-light Captures
by: Zheng, Quan, et al.
Published: (2024)
by: Zheng, Quan, et al.
Published: (2024)
Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
by: Kim, Sohee, et al.
Published: (2025)
by: Kim, Sohee, et al.
Published: (2025)
Semantic Smoothing via Novel View Synthesis for Robust SAR Image Classification
by: Brignac, Daniel, et al.
Published: (2026)
by: Brignac, Daniel, et al.
Published: (2026)
Simultaneous Dual-View Mammogram Synthesis Using Denoising Diffusion Probabilistic Models
by: Garza-Abdala, Jorge Alberto, et al.
Published: (2026)
by: Garza-Abdala, Jorge Alberto, et al.
Published: (2026)
Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
by: Acuna, David, et al.
Published: (2025)
by: Acuna, David, et al.
Published: (2025)
Scaling Non-Parametric Sampling with Representation
by: Lu, Vincent, et al.
Published: (2025)
by: Lu, Vincent, et al.
Published: (2025)
Compositional Image Synthesis with Inference-Time Scaling
by: Ji, Minsuk, et al.
Published: (2025)
by: Ji, Minsuk, et al.
Published: (2025)
Addressing Diverging Training Costs using BEVRestore for High-resolution Bird's Eye View Map Construction
by: Kim, Minsu, et al.
Published: (2024)
by: Kim, Minsu, et al.
Published: (2024)
Summer-22B: A Systematic Approach to Dataset Engineering and Training at Scale for Video Foundation Model
by: Ryu, Simo, et al.
Published: (2026)
by: Ryu, Simo, et al.
Published: (2026)
MammoRGB: Dual-View Mammogram Synthesis Using Denoising Diffusion Probabilistic Models
by: Garza-Abdala, Jorge Alberto, et al.
Published: (2025)
by: Garza-Abdala, Jorge Alberto, et al.
Published: (2025)
Similar Items
-
True Self-Supervised Novel View Synthesis is Transferable
by: Mitchel, Thomas W., et al.
Published: (2025) -
Neural Isometries: Taming Transformations for Equivariant ML
by: Mitchel, Thomas W., et al.
Published: (2024) -
Understanding Multi-View Transformers
by: Stary, Michal, et al.
Published: (2025) -
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
by: Cazenavette, George, et al.
Published: (2025) -
Bridging Implicit and Explicit Geometric Transformation for Single-Image View Synthesis
by: Park, Byeongjun, et al.
Published: (2022)