Floating No More: Object-Ground Reconstruction from a Single Image
Fuente:
arXiv
Saved in:
| Main Authors: | Man, Yunze, Sheng, Yichen, Zhang, Jianming, Gui, Liang-Yan, Wang, Yu-Xiong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DualCross: Cross-Modality Cross-Domain Adaptation for Monocular BEV Perception
by: Man, Yunze, et al.
Published: (2023)
by: Man, Yunze, et al.
Published: (2023)
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
by: Man, Yunze, et al.
Published: (2025)
by: Man, Yunze, et al.
Published: (2025)
Situational Awareness Matters in 3D Vision Language Reasoning
by: Man, Yunze, et al.
Published: (2024)
by: Man, Yunze, et al.
Published: (2024)
PaintScene4D: Consistent 4D Scene Generation from Text Prompts
by: Gupta, Vinayak, et al.
Published: (2024)
by: Gupta, Vinayak, et al.
Published: (2024)
Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision
by: Cao, Shengcao, et al.
Published: (2024)
by: Cao, Shengcao, et al.
Published: (2024)
LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
by: Man, Yunze, et al.
Published: (2025)
by: Man, Yunze, et al.
Published: (2025)
SceneCraft: Layout-Guided 3D Scene Generation
by: Yang, Xiuyu, et al.
Published: (2024)
by: Yang, Xiuyu, et al.
Published: (2024)
TripoSR: Fast 3D Object Reconstruction from a Single Image
by: Tochilkin, Dmitry, et al.
Published: (2024)
by: Tochilkin, Dmitry, et al.
Published: (2024)
Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
by: Man, Yunze, et al.
Published: (2024)
by: Man, Yunze, et al.
Published: (2024)
REACTO: Reconstructing Articulated Objects from a Single Video
by: Song, Chaoyue, et al.
Published: (2024)
by: Song, Chaoyue, et al.
Published: (2024)
PPTArena: A Benchmark for Agentic PowerPoint Editing
by: Ofengenden, Michael, et al.
Published: (2025)
by: Ofengenden, Michael, et al.
Published: (2025)
InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object Interaction
by: Xu, Sirui, et al.
Published: (2024)
by: Xu, Sirui, et al.
Published: (2024)
InterMimic: Towards Universal Whole-Body Control for Physics-Based Human-Object Interactions
by: Xu, Sirui, et al.
Published: (2025)
by: Xu, Sirui, et al.
Published: (2025)
Hyperbolic-constraint Point Cloud Reconstruction from Single RGB-D Images
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
Realistic Clothed Human and Object Joint Reconstruction from a Single Image
by: Dutta, Ayushi, et al.
Published: (2025)
by: Dutta, Ayushi, et al.
Published: (2025)
Research on Detection of Floating Objects in River and Lake Based on AI Intelligent Image Recognition
by: Zhang, Jingyu, et al.
Published: (2024)
by: Zhang, Jingyu, et al.
Published: (2024)
AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark
by: Gauba, Aruna, et al.
Published: (2025)
by: Gauba, Aruna, et al.
Published: (2025)
MoReact: Generating Reactive Motion from Textual Descriptions
by: Xu, Xiyan, et al.
Published: (2025)
by: Xu, Xiyan, et al.
Published: (2025)
Unleashing Guidance Without Classifiers for Human-Object Interaction Animation
by: Wang, Ziyin, et al.
Published: (2026)
by: Wang, Ziyin, et al.
Published: (2026)
HASSOD: Hierarchical Adaptive Self-Supervised Object Detection
by: Cao, Shengcao, et al.
Published: (2024)
by: Cao, Shengcao, et al.
Published: (2024)
GroundingBooth: Grounding Text-to-Image Customization
by: Xiong, Zhexiao, et al.
Published: (2024)
by: Xiong, Zhexiao, et al.
Published: (2024)
PixelDiT: Pixel Diffusion Transformers for Image Generation
by: Yu, Yongsheng, et al.
Published: (2025)
by: Yu, Yongsheng, et al.
Published: (2025)
Single-Slice-to-3D Reconstruction in Medical Imaging and Natural Objects: A Comparative Benchmark with SAM 3D
by: Luo, Yan, et al.
Published: (2026)
by: Luo, Yan, et al.
Published: (2026)
Detecting Human Artifacts from Text-to-Image Models
by: Wang, Kaihong, et al.
Published: (2024)
by: Wang, Kaihong, et al.
Published: (2024)
Frozen Transformers in Language Models Are Effective Visual Encoder Layers
by: Pang, Ziqi, et al.
Published: (2023)
by: Pang, Ziqi, et al.
Published: (2023)
Detection Based Part-level Articulated Object Reconstruction from Single RGBD Image
by: Kawana, Yuki, et al.
Published: (2025)
by: Kawana, Yuki, et al.
Published: (2025)
MTFusion: Reconstructing Any 3D Object from Single Image Using Multi-word Textual Inversion
by: Liu, Yu, et al.
Published: (2024)
by: Liu, Yu, et al.
Published: (2024)
Holo-Relighting: Controllable Volumetric Portrait Relighting from a Single Image
by: Mei, Yiqun, et al.
Published: (2024)
by: Mei, Yiqun, et al.
Published: (2024)
Effective Gaussian Management for High-fidelity Object Reconstruction
by: Liu, Jiateng, et al.
Published: (2025)
by: Liu, Jiateng, et al.
Published: (2025)
Dual-Integrated Low-Latency Single-Lens Infrared Computational Imaging for Object Detection
by: Wang, Xuquan, et al.
Published: (2026)
by: Wang, Xuquan, et al.
Published: (2026)
Variable Radiance Field for Real-World Category-Specific Reconstruction from Single Image
by: Wang, Kun, et al.
Published: (2023)
by: Wang, Kun, et al.
Published: (2023)
Calibration & Reconstruction: Deep Integrated Language for Referring Image Segmentation
by: Yan, Yichen, et al.
Published: (2024)
by: Yan, Yichen, et al.
Published: (2024)
SR3D: Unleashing Single-view 3D Reconstruction for Transparent and Specular Object Grasping
by: Zhang, Mingxu, et al.
Published: (2025)
by: Zhang, Mingxu, et al.
Published: (2025)
Robust Tiny Object Detection in Aerial Images amidst Label Noise
by: Zhu, Haoran, et al.
Published: (2024)
by: Zhu, Haoran, et al.
Published: (2024)
High-fidelity 3D Object Generation from Single Image with RGBN-Volume Gaussian Reconstruction Model
by: Shen, Yiyang, et al.
Published: (2025)
by: Shen, Yiyang, et al.
Published: (2025)
Tiny Object Detection with Single Point Supervision
by: Zhu, Haoran, et al.
Published: (2024)
by: Zhu, Haoran, et al.
Published: (2024)
LiCAF: LiDAR-Camera Asymmetric Fusion for Gait Recognition
by: Deng, Yunze, et al.
Published: (2024)
by: Deng, Yunze, et al.
Published: (2024)
TelePhysics: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interaction
by: Zhang, Xin, et al.
Published: (2026)
by: Zhang, Xin, et al.
Published: (2026)
Visual Grounding with Multi-modal Conditional Adaptation
by: Yao, Ruilin, et al.
Published: (2024)
by: Yao, Ruilin, et al.
Published: (2024)
CADDreamer: CAD Object Generation from Single-view Images
by: Li, Yuan, et al.
Published: (2025)
by: Li, Yuan, et al.
Published: (2025)
Similar Items
-
DualCross: Cross-Modality Cross-Domain Adaptation for Monocular BEV Perception
by: Man, Yunze, et al.
Published: (2023) -
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
by: Man, Yunze, et al.
Published: (2025) -
Situational Awareness Matters in 3D Vision Language Reasoning
by: Man, Yunze, et al.
Published: (2024) -
PaintScene4D: Consistent 4D Scene Generation from Text Prompts
by: Gupta, Vinayak, et al.
Published: (2024) -
Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision
by: Cao, Shengcao, et al.
Published: (2024)