Predicting Long-horizon Futures by Conditioning on Geometry and Time
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Khurana, Tarasha, Ramanan, Deva |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reconstruct, Inpaint, Test-Time Finetune: Dynamic Novel-view Synthesis from Monocular Videos
von: Chen, Kaihua, et al.
Veröffentlicht: (2025)
von: Chen, Kaihua, et al.
Veröffentlicht: (2025)
Using Diffusion Priors for Video Amodal Segmentation
von: Chen, Kaihua, et al.
Veröffentlicht: (2024)
von: Chen, Kaihua, et al.
Veröffentlicht: (2024)
MonoFusion: Sparse-View 4D Reconstruction via Monocular Fusion
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
TAO-Amodal: A Benchmark for Tracking Any Object Amodally
von: Hsieh, Cheng-Yen, et al.
Veröffentlicht: (2023)
von: Hsieh, Cheng-Yen, et al.
Veröffentlicht: (2023)
Shelf-Supervised Cross-Modal Pre-Training for 3D Object Detection
von: Khurana, Mehar, et al.
Veröffentlicht: (2024)
von: Khurana, Mehar, et al.
Veröffentlicht: (2024)
RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
von: Robinson, Isaac, et al.
Veröffentlicht: (2025)
von: Robinson, Isaac, et al.
Veröffentlicht: (2025)
RefAV: Towards Planning-Centric Scenario Mining
von: Davidson, Cainan, et al.
Veröffentlicht: (2025)
von: Davidson, Cainan, et al.
Veröffentlicht: (2025)
SMORE: Simultaneous Map and Object REconstruction
von: Chodosh, Nathaniel, et al.
Veröffentlicht: (2024)
von: Chodosh, Nathaniel, et al.
Veröffentlicht: (2024)
Revisiting Few-Shot Object Detection with Vision-Language Models
von: Madan, Anish, et al.
Veröffentlicht: (2023)
von: Madan, Anish, et al.
Veröffentlicht: (2023)
Long-Tailed 3D Detection via Multi-Modal Fusion
von: Ma, Yechi, et al.
Veröffentlicht: (2023)
von: Ma, Yechi, et al.
Veröffentlicht: (2023)
DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion
von: Zhao, Qitao, et al.
Veröffentlicht: (2025)
von: Zhao, Qitao, et al.
Veröffentlicht: (2025)
RaySt3R: Predicting Novel Depth Maps for Zero-Shot Object Completion
von: Duisterhof, Bardienus P., et al.
Veröffentlicht: (2025)
von: Duisterhof, Bardienus P., et al.
Veröffentlicht: (2025)
I Can't Believe It's Not Scene Flow!
von: Khatri, Ishan, et al.
Veröffentlicht: (2024)
von: Khatri, Ishan, et al.
Veröffentlicht: (2024)
DressRecon: Freeform 4D Human Reconstruction from Monocular Video
von: Tan, Jeff, et al.
Veröffentlicht: (2024)
von: Tan, Jeff, et al.
Veröffentlicht: (2024)
PAI-Bench: A Comprehensive Benchmark For Physical AI
von: Zhou, Fengzhe, et al.
Veröffentlicht: (2025)
von: Zhou, Fengzhe, et al.
Veröffentlicht: (2025)
AerialMegaDepth: Learning Aerial-Ground Reconstruction and View Synthesis
von: Vuong, Khiem, et al.
Veröffentlicht: (2025)
von: Vuong, Khiem, et al.
Veröffentlicht: (2025)
Novel View Synthesis as Video Completion
von: Wu, Qi, et al.
Veröffentlicht: (2026)
von: Wu, Qi, et al.
Veröffentlicht: (2026)
DynOMo: Online Point Tracking by Dynamic Online Monocular Gaussian Reconstruction
von: Seidenschwarz, Jenny, et al.
Veröffentlicht: (2024)
von: Seidenschwarz, Jenny, et al.
Veröffentlicht: (2024)
Soft Augmentation for Image Classification
von: Liu, Yang, et al.
Veröffentlicht: (2022)
von: Liu, Yang, et al.
Veröffentlicht: (2022)
Generating Physically Stable and Buildable Brick Structures from Text
von: Pun, Ava, et al.
Veröffentlicht: (2025)
von: Pun, Ava, et al.
Veröffentlicht: (2025)
DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection
von: Gare, Gautam Rajendrakumar, et al.
Veröffentlicht: (2026)
von: Gare, Gautam Rajendrakumar, et al.
Veröffentlicht: (2026)
Depth-supervised NeRF: Fewer Views and Faster Training for Free
von: Deng, Kangle, et al.
Veröffentlicht: (2021)
von: Deng, Kangle, et al.
Veröffentlicht: (2021)
Revisiting the Role of Language Priors in Vision-Language Models
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2023)
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2023)
UniFlow: Zero-Shot LiDAR Scene Flow for Autonomous Vehicles
von: Li, Siyi, et al.
Veröffentlicht: (2025)
von: Li, Siyi, et al.
Veröffentlicht: (2025)
Steerable Visual Representations
von: Ruthardt, Jona, et al.
Veröffentlicht: (2026)
von: Ruthardt, Jona, et al.
Veröffentlicht: (2026)
Better Call SAL: Towards Learning to Segment Anything in Lidar
von: Ošep, Aljoša, et al.
Veröffentlicht: (2024)
von: Ošep, Aljoša, et al.
Veröffentlicht: (2024)
Towards Foundational Models for Single-Chip Radar
von: Huang, Tianshu, et al.
Veröffentlicht: (2025)
von: Huang, Tianshu, et al.
Veröffentlicht: (2025)
Just Add Geometry: Gradient-Free Open-Vocabulary 3D Detection Without Human-in-the-Loop
von: Goel, Atharv, et al.
Veröffentlicht: (2025)
von: Goel, Atharv, et al.
Veröffentlicht: (2025)
Lidar Panoptic Segmentation in an Open World
von: Chakravarthy, Anirudh S, et al.
Veröffentlicht: (2024)
von: Chakravarthy, Anirudh S, et al.
Veröffentlicht: (2024)
WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning
von: Chen, Delong, et al.
Veröffentlicht: (2025)
von: Chen, Delong, et al.
Veröffentlicht: (2025)
LIVE: Long-horizon Interactive Video World Modeling
von: Huang, Junchao, et al.
Veröffentlicht: (2026)
von: Huang, Junchao, et al.
Veröffentlicht: (2026)
RadarSim: Simulating Single-Chip Radar via Multimodal Neural Fields
von: Chen, Chuhan, et al.
Veröffentlicht: (2026)
von: Chen, Chuhan, et al.
Veröffentlicht: (2026)
Cameras as Rays: Pose Estimation via Ray Diffusion
von: Zhang, Jason Y., et al.
Veröffentlicht: (2024)
von: Zhang, Jason Y., et al.
Veröffentlicht: (2024)
CRISP: Contact-Guided Real2Sim from Monocular Video with Planar Scene Primitives
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
Geometry-aided Vision-based Localization of Future Mars Helicopters in Challenging Illumination Conditions
von: Pisanti, Dario, et al.
Veröffentlicht: (2025)
von: Pisanti, Dario, et al.
Veröffentlicht: (2025)
Roboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language Models
von: Robicheaux, Peter, et al.
Veröffentlicht: (2025)
von: Robicheaux, Peter, et al.
Veröffentlicht: (2025)
Neural Eulerian Scene Flow Fields
von: Vedder, Kyle, et al.
Veröffentlicht: (2024)
von: Vedder, Kyle, et al.
Veröffentlicht: (2024)
InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning
von: Wan, Zifu, et al.
Veröffentlicht: (2025)
von: Wan, Zifu, et al.
Veröffentlicht: (2025)
Towards Long-horizon Agentic Multimodal Search
von: Du, Yifan, et al.
Veröffentlicht: (2026)
von: Du, Yifan, et al.
Veröffentlicht: (2026)
Probability-Flow Distillation: Exact Wasserstein Gradient Flow for High-Fidelity 3D Generation
von: Ramanan, Rohith, et al.
Veröffentlicht: (2026)
von: Ramanan, Rohith, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Reconstruct, Inpaint, Test-Time Finetune: Dynamic Novel-view Synthesis from Monocular Videos
von: Chen, Kaihua, et al.
Veröffentlicht: (2025) -
Using Diffusion Priors for Video Amodal Segmentation
von: Chen, Kaihua, et al.
Veröffentlicht: (2024) -
MonoFusion: Sparse-View 4D Reconstruction via Monocular Fusion
von: Wang, Zihan, et al.
Veröffentlicht: (2025) -
TAO-Amodal: A Benchmark for Tracking Any Object Amodally
von: Hsieh, Cheng-Yen, et al.
Veröffentlicht: (2023) -
Shelf-Supervised Cross-Modal Pre-Training for 3D Object Detection
von: Khurana, Mehar, et al.
Veröffentlicht: (2024)