Can Generative Video Models Help Pose Estimation?
Fuente:
arXiv
Salvato in:
| Autori principali: | Cai, Ruojin, Zhang, Jason Y., Henzler, Philipp, Li, Zhengqi, Snavely, Noah, Martin-Brualla, Ricardo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ArchSym: Detecting 3D-Grounded Architectural Symmetries in the Wild
di: Chen, Hanyu, et al.
Pubblicazione: (2026)
di: Chen, Hanyu, et al.
Pubblicazione: (2026)
CityRAG: Stepping Into a City via Spatially-Grounded Video Generation
di: Chou, Gene, et al.
Pubblicazione: (2026)
di: Chou, Gene, et al.
Pubblicazione: (2026)
Generative Image Dynamics
di: Li, Zhengqi, et al.
Pubblicazione: (2023)
di: Li, Zhengqi, et al.
Pubblicazione: (2023)
Doppelgangers++: Improved Visual Disambiguation with Geometric 3D Features
di: Xiangli, Yuanbo, et al.
Pubblicazione: (2024)
di: Xiangli, Yuanbo, et al.
Pubblicazione: (2024)
Long-tail Internet photo reconstruction
di: Li, Yuan, et al.
Pubblicazione: (2026)
di: Li, Yuan, et al.
Pubblicazione: (2026)
Wide-Baseline Relative Camera Pose Estimation with Directional Learning
di: Chen, Kefan, et al.
Pubblicazione: (2021)
di: Chen, Kefan, et al.
Pubblicazione: (2021)
Bolt3D: Generating 3D Scenes in Seconds
di: Szymanowicz, Stanislaw, et al.
Pubblicazione: (2025)
di: Szymanowicz, Stanislaw, et al.
Pubblicazione: (2025)
Streetscapes: Large-scale Consistent Street View Generation Using Autoregressive Video Diffusion
di: Deng, Boyang, et al.
Pubblicazione: (2024)
di: Deng, Boyang, et al.
Pubblicazione: (2024)
CAT3D: Create Anything in 3D with Multi-View Diffusion Models
di: Gao, Ruiqi, et al.
Pubblicazione: (2024)
di: Gao, Ruiqi, et al.
Pubblicazione: (2024)
Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos
di: Jin, Linyi, et al.
Pubblicazione: (2024)
di: Jin, Linyi, et al.
Pubblicazione: (2024)
MegaScenes: Scene-Level View Synthesis at Scale
di: Tung, Joseph, et al.
Pubblicazione: (2024)
di: Tung, Joseph, et al.
Pubblicazione: (2024)
IllumiNeRF: 3D Relighting Without Inverse Rendering
di: Zhao, Xiaoming, et al.
Pubblicazione: (2024)
di: Zhao, Xiaoming, et al.
Pubblicazione: (2024)
ROGR: Relightable 3D Objects using Generative Relighting
di: Tang, Jiapeng, et al.
Pubblicazione: (2025)
di: Tang, Jiapeng, et al.
Pubblicazione: (2025)
Learning Feature Descriptors using Camera Pose Supervision
di: Wang, Qianqian, et al.
Pubblicazione: (2020)
di: Wang, Qianqian, et al.
Pubblicazione: (2020)
Beyond the Frame: Generating 360 Panoramic Videos from Perspective Videos
di: Luo, Rundong, et al.
Pubblicazione: (2025)
di: Luo, Rundong, et al.
Pubblicazione: (2025)
MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos
di: Li, Zhengqi, et al.
Pubblicazione: (2024)
di: Li, Zhengqi, et al.
Pubblicazione: (2024)
Extreme Rotation Estimation in the Wild
di: Bezalel, Hana, et al.
Pubblicazione: (2024)
di: Bezalel, Hana, et al.
Pubblicazione: (2024)
Perturb-and-Revise: Flexible 3D Editing with Generative Trajectories
di: Hong, Susung, et al.
Pubblicazione: (2024)
di: Hong, Susung, et al.
Pubblicazione: (2024)
G3T Up! Gravity Aligned Coordinate Frames Simplify Pointmap Processing
di: Kani, Bharath Raj Nagoor, et al.
Pubblicazione: (2026)
di: Kani, Bharath Raj Nagoor, et al.
Pubblicazione: (2026)
FlashDepth: Real-time Streaming Video Depth Estimation at 2K Resolution
di: Chou, Gene, et al.
Pubblicazione: (2025)
di: Chou, Gene, et al.
Pubblicazione: (2025)
Seeing a Rose in Five Thousand Ways
di: Zhang, Yunzhi, et al.
Pubblicazione: (2022)
di: Zhang, Yunzhi, et al.
Pubblicazione: (2022)
KFC-W: Generating 3D-Consistent Videos from Unposed Internet Photos
di: Chou, Gene, et al.
Pubblicazione: (2024)
di: Chou, Gene, et al.
Pubblicazione: (2024)
End-to-End Multi-Person Pose Estimation with Pose-Aware Video Transformer
di: Yu, Yonghui, et al.
Pubblicazione: (2025)
di: Yu, Yonghui, et al.
Pubblicazione: (2025)
MoMaps: Semantics-Aware Scene Motion Generation with Motion Maps
di: Lei, Jiahui, et al.
Pubblicazione: (2025)
di: Lei, Jiahui, et al.
Pubblicazione: (2025)
Can Large Pretrained Depth Estimation Models Help With Image Dehazing?
di: Zhang, Hongfei, et al.
Pubblicazione: (2025)
di: Zhang, Hongfei, et al.
Pubblicazione: (2025)
C3Po: Cross-View Cross-Modality Correspondence by Pointmap Prediction
di: Huang, Kuan Wei, et al.
Pubblicazione: (2025)
di: Huang, Kuan Wei, et al.
Pubblicazione: (2025)
Honey, I Shrunk the Arc de Triomphe!
di: Xiangli, Yuanbo, et al.
Pubblicazione: (2026)
di: Xiangli, Yuanbo, et al.
Pubblicazione: (2026)
VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression
di: Sargent, Kyle, et al.
Pubblicazione: (2025)
di: Sargent, Kyle, et al.
Pubblicazione: (2025)
PhysDreamer: Physics-Based Interaction with 3D Objects via Video Generation
di: Zhang, Tianyuan, et al.
Pubblicazione: (2024)
di: Zhang, Tianyuan, et al.
Pubblicazione: (2024)
Emergent Extreme-View Geometry in 3D Foundation Models
di: Zhang, Yiwen, et al.
Pubblicazione: (2025)
di: Zhang, Yiwen, et al.
Pubblicazione: (2025)
PoseCrafter: Extreme Pose Estimation with Hybrid Video Synthesis
di: Mao, Qing, et al.
Pubblicazione: (2025)
di: Mao, Qing, et al.
Pubblicazione: (2025)
LFM-3D: Learnable Feature Matching Across Wide Baselines Using 3D Signals
di: Karpur, Arjun, et al.
Pubblicazione: (2023)
di: Karpur, Arjun, et al.
Pubblicazione: (2023)
Kinematics Modeling Network for Video-based Human Pose Estimation
di: Dang, Yonghao, et al.
Pubblicazione: (2022)
di: Dang, Yonghao, et al.
Pubblicazione: (2022)
Linear Relative Pose Estimation Founded on Pose-only Imaging Geometry
di: Cai, Qi, et al.
Pubblicazione: (2024)
di: Cai, Qi, et al.
Pubblicazione: (2024)
Eye2Eye: A Simple Approach for Monocular-to-Stereo Video Synthesis
di: Geyer, Michal, et al.
Pubblicazione: (2025)
di: Geyer, Michal, et al.
Pubblicazione: (2025)
Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos
di: Ma, Yue, et al.
Pubblicazione: (2023)
di: Ma, Yue, et al.
Pubblicazione: (2023)
DBINDS -- Can Initial Noise from Diffusion Model Inversion Help Reveal AI-Generated Videos?
di: Wu, Yanlin, et al.
Pubblicazione: (2025)
di: Wu, Yanlin, et al.
Pubblicazione: (2025)
CRAG: Can 3D Generative Models Help 3D Assembly?
di: Jiang, Zeyu, et al.
Pubblicazione: (2026)
di: Jiang, Zeyu, et al.
Pubblicazione: (2026)
How Can Objects Help Video-Language Understanding?
di: Tang, Zitian, et al.
Pubblicazione: (2025)
di: Tang, Zitian, et al.
Pubblicazione: (2025)
CubeDiff: Repurposing Diffusion-Based Image Models for Panorama Generation
di: Kalischek, Nikolai, et al.
Pubblicazione: (2025)
di: Kalischek, Nikolai, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ArchSym: Detecting 3D-Grounded Architectural Symmetries in the Wild
di: Chen, Hanyu, et al.
Pubblicazione: (2026) -
CityRAG: Stepping Into a City via Spatially-Grounded Video Generation
di: Chou, Gene, et al.
Pubblicazione: (2026) -
Generative Image Dynamics
di: Li, Zhengqi, et al.
Pubblicazione: (2023) -
Doppelgangers++: Improved Visual Disambiguation with Geometric 3D Features
di: Xiangli, Yuanbo, et al.
Pubblicazione: (2024) -
Long-tail Internet photo reconstruction
di: Li, Yuan, et al.
Pubblicazione: (2026)