Where Am I and What Will I See: An Auto-Regressive Model for Spatial Localization and View Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Junyi, Huang, Di, Ye, Weicai, Ouyang, Wanli, He, Tong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GigaGS: Scaling up Planar-Based 3D Gaussians for Large Scene Surface Reconstruction
by: Chen, Junyi, et al.
Published: (2024)
by: Chen, Junyi, et al.
Published: (2024)
NeuRodin: A Two-stage Framework for High-Fidelity Neural Surface Reconstruction
by: Wang, Yifan, et al.
Published: (2024)
by: Wang, Yifan, et al.
Published: (2024)
NeRF-Det++: Incorporating Semantic Cues and Perspective-aware Depth Supervision for Indoor Multi-View 3D Detection
by: Huang, Chenxi, et al.
Published: (2024)
by: Huang, Chenxi, et al.
Published: (2024)
Point Cloud Matters: Rethinking the Impact of Different Observation Spaces on Robot Learning
by: Zhu, Haoyi, et al.
Published: (2024)
by: Zhu, Haoyi, et al.
Published: (2024)
HiSplat: Hierarchical 3D Gaussian Splatting for Generalizable Sparse-View Reconstruction
by: Tang, Shengji, et al.
Published: (2024)
by: Tang, Shengji, et al.
Published: (2024)
DATAP-SfM: Dynamic-Aware Tracking Any Point for Robust Structure from Motion in the Wild
by: Ye, Weicai, et al.
Published: (2024)
by: Ye, Weicai, et al.
Published: (2024)
Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering
by: Choi, Yura, et al.
Published: (2026)
by: Choi, Yura, et al.
Published: (2026)
VINO: A Unified Visual Generator with Interleaved OmniModal Context
by: Chen, Junyi, et al.
Published: (2026)
by: Chen, Junyi, et al.
Published: (2026)
Where am I? Cross-View Geo-localization with Natural Language Descriptions
by: Ye, Junyan, et al.
Published: (2024)
by: Ye, Junyan, et al.
Published: (2024)
GVGEN: Text-to-3D Generation with Volumetric Representation
by: He, Xianglong, et al.
Published: (2024)
by: He, Xianglong, et al.
Published: (2024)
Progressive Cross-Stream Cooperation in Spatial and Temporal Domain for Action Localization
by: Su, Rui, et al.
Published: (2019)
by: Su, Rui, et al.
Published: (2019)
EMR-Merging: Tuning-Free High-Performance Model Merging
by: Huang, Chenyu, et al.
Published: (2024)
by: Huang, Chenyu, et al.
Published: (2024)
MeshCraft: Exploring Efficient and Controllable Mesh Generation with Flow-based DiTs
by: He, Xianglong, et al.
Published: (2025)
by: He, Xianglong, et al.
Published: (2025)
I Am Big, You Are Little; I Am Right, You Are Wrong
by: Kelly, David A., et al.
Published: (2025)
by: Kelly, David A., et al.
Published: (2025)
DynaSurfGS: Dynamic Surface Reconstruction with Planar-based Gaussian Splatting
by: Cai, Weiwei, et al.
Published: (2024)
by: Cai, Weiwei, et al.
Published: (2024)
Auto-Regressively Generating Multi-View Consistent Images
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
Decoupling What to Count and Where to See for Referring Expression Counting
by: Zou, Yuda, et al.
Published: (2025)
by: Zou, Yuda, et al.
Published: (2025)
PredBench: Benchmarking Spatio-Temporal Prediction across Diverse Disciplines
by: Wang, ZiDong, et al.
Published: (2024)
by: Wang, ZiDong, et al.
Published: (2024)
LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
by: Fuller, Anthony, et al.
Published: (2025)
by: Fuller, Anthony, et al.
Published: (2025)
Agent3D-Zero: An Agent for Zero-shot 3D Understanding
by: Zhang, Sha, et al.
Published: (2024)
by: Zhang, Sha, et al.
Published: (2024)
DiffPano: Scalable and Consistent Text to Panorama Generation with Spherical Epipolar-Aware Diffusion
by: Ye, Weicai, et al.
Published: (2024)
by: Ye, Weicai, et al.
Published: (2024)
A Watermark for Auto-Regressive Image Generation Models
by: Wu, Yihan, et al.
Published: (2025)
by: Wu, Yihan, et al.
Published: (2025)
CoSurfGS:Collaborative 3D Surface Gaussian Splatting with Distributed Learning for Large Scene Reconstruction
by: Gao, Yuanyuan, et al.
Published: (2024)
by: Gao, Yuanyuan, et al.
Published: (2024)
Adapter-X: A Novel General Parameter-Efficient Fine-Tuning Framework for Vision
by: Li, Minglei, et al.
Published: (2024)
by: Li, Minglei, et al.
Published: (2024)
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
by: Feng, Zhiyuan, et al.
Published: (2025)
by: Feng, Zhiyuan, et al.
Published: (2025)
ND-SDF: Learning Normal Deflection Fields for High-Fidelity Indoor Reconstruction
by: Tang, Ziyu, et al.
Published: (2024)
by: Tang, Ziyu, et al.
Published: (2024)
PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation
by: Liu, Zhuoman, et al.
Published: (2024)
by: Liu, Zhuoman, et al.
Published: (2024)
"Where am I?" Scene Retrieval with Language
by: Chen, Jiaqi, et al.
Published: (2024)
by: Chen, Jiaqi, et al.
Published: (2024)
MRGeo: Robust Cross-View Geo-Localization of Corrupted Images via Spatial and Channel Feature Enhancement
by: Wu, Le, et al.
Published: (2026)
by: Wu, Le, et al.
Published: (2026)
PoI: A Filter to Extract Pixel of Interest from Novel Views for Scene Coordinate Regression
by: Li, Feifei, et al.
Published: (2025)
by: Li, Feifei, et al.
Published: (2025)
See4D: Pose-Free 4D Generation via Auto-Regressive Video Inpainting
by: Lu, Dongyue, et al.
Published: (2025)
by: Lu, Dongyue, et al.
Published: (2025)
ZipAR: Parallel Auto-regressive Image Generation through Spatial Locality
by: He, Yefei, et al.
Published: (2024)
by: He, Yefei, et al.
Published: (2024)
Quality-Aware Language-Conditioned Local Auto-Regressive Anomaly Synthesis and Detection
by: Qian, Long, et al.
Published: (2025)
by: Qian, Long, et al.
Published: (2025)
Improving Weakly Supervised Temporal Action Localization by Exploiting Multi-resolution Information in Temporal Domain
by: Su, Rui, et al.
Published: (2025)
by: Su, Rui, et al.
Published: (2025)
PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm
by: Zhu, Haoyi, et al.
Published: (2023)
by: Zhu, Haoyi, et al.
Published: (2023)
Do You See What I Say? Generalizable Deepfake Detection based on Visual Speech Recognition
by: Bora, Maheswar, et al.
Published: (2025)
by: Bora, Maheswar, et al.
Published: (2025)
Seeing Delta Parameters as JPEG Images: Data-Free Delta Compression with Discrete Cosine Transform
by: Huang, Chenyu, et al.
Published: (2025)
by: Huang, Chenyu, et al.
Published: (2025)
Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient
by: Chen, Zigeng, et al.
Published: (2024)
by: Chen, Zigeng, et al.
Published: (2024)
M2ORT: Many-To-One Regression Transformer for Spatial Transcriptomics Prediction from Histopathology Images
by: Wang, Hongyi, et al.
Published: (2024)
by: Wang, Hongyi, et al.
Published: (2024)
FiT: Flexible Vision Transformer for Diffusion Model
by: Lu, Zeyu, et al.
Published: (2024)
by: Lu, Zeyu, et al.
Published: (2024)
Similar Items
-
GigaGS: Scaling up Planar-Based 3D Gaussians for Large Scene Surface Reconstruction
by: Chen, Junyi, et al.
Published: (2024) -
NeuRodin: A Two-stage Framework for High-Fidelity Neural Surface Reconstruction
by: Wang, Yifan, et al.
Published: (2024) -
NeRF-Det++: Incorporating Semantic Cues and Perspective-aware Depth Supervision for Indoor Multi-View 3D Detection
by: Huang, Chenxi, et al.
Published: (2024) -
Point Cloud Matters: Rethinking the Impact of Different Observation Spaces on Robot Learning
by: Zhu, Haoyi, et al.
Published: (2024) -
HiSplat: Hierarchical 3D Gaussian Splatting for Generalizable Sparse-View Reconstruction
by: Tang, Shengji, et al.
Published: (2024)