Towards Unified 3D Object Detection via Algorithm and Data Unification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zhuoling, Xu, Xiaogang, Lim, SerNam, Zhao, Hengshuang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence
von: Li, Zhuoling, et al.
Veröffentlicht: (2024)
von: Li, Zhuoling, et al.
Veröffentlicht: (2024)
DreamMask: Boosting Open-vocabulary Panoptic Segmentation with Synthetic Data
von: Tu, Yuanpeng, et al.
Veröffentlicht: (2025)
von: Tu, Yuanpeng, et al.
Veröffentlicht: (2025)
OV-Uni3DETR: Towards Unified Open-Vocabulary 3D Object Detection via Cycle-Modality Propagation
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024)
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024)
Towards Chunk-Wise Generation for Long Videos
von: Zhang, Siyang, et al.
Veröffentlicht: (2024)
von: Zhang, Siyang, et al.
Veröffentlicht: (2024)
VideoMerge: Towards Training-free Long Video Generation
von: Zhang, Siyang, et al.
Veröffentlicht: (2025)
von: Zhang, Siyang, et al.
Veröffentlicht: (2025)
FocalClick-XL: Towards Unified and High-quality Interactive Segmentation
von: Chen, Xi, et al.
Veröffentlicht: (2025)
von: Chen, Xi, et al.
Veröffentlicht: (2025)
One for All: Multi-Domain Joint Training for Point Cloud Based 3D Object Detection
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024)
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024)
Composing Object Relations and Attributes for Image-Text Matching
von: Pham, Khoi, et al.
Veröffentlicht: (2024)
von: Pham, Khoi, et al.
Veröffentlicht: (2024)
FSViewFusion: Few-Shots View Generation of Novel Objects
von: Hussain, Rukhshanda, et al.
Veröffentlicht: (2024)
von: Hussain, Rukhshanda, et al.
Veröffentlicht: (2024)
Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds
von: Fan, Xianzhe, et al.
Veröffentlicht: (2026)
von: Fan, Xianzhe, et al.
Veröffentlicht: (2026)
Object Recognition as Next Token Prediction
von: Yue, Kaiyu, et al.
Veröffentlicht: (2023)
von: Yue, Kaiyu, et al.
Veröffentlicht: (2023)
Towards Category Unification of 3D Single Object Tracking on Point Clouds
von: Nie, Jiahao, et al.
Veröffentlicht: (2024)
von: Nie, Jiahao, et al.
Veröffentlicht: (2024)
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
von: Yang, Lihe, et al.
Veröffentlicht: (2024)
von: Yang, Lihe, et al.
Veröffentlicht: (2024)
Mitigating Dialogue Hallucination for Large Vision Language Models via Adversarial Instruction Tuning
von: Park, Dongmin, et al.
Veröffentlicht: (2024)
von: Park, Dongmin, et al.
Veröffentlicht: (2024)
LION: Linear Group RNN for 3D Object Detection in Point Clouds
von: Liu, Zhe, et al.
Veröffentlicht: (2024)
von: Liu, Zhe, et al.
Veröffentlicht: (2024)
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
von: Zhou, Xin, et al.
Veröffentlicht: (2026)
von: Zhou, Xin, et al.
Veröffentlicht: (2026)
Enhancing Diffusion-based Restoration Models via Difficulty-Adaptive Reinforcement Learning with IQA Reward
von: Xu, Xiaogang, et al.
Veröffentlicht: (2025)
von: Xu, Xiaogang, et al.
Veröffentlicht: (2025)
DiffCamera: Arbitrary Refocusing on Images
von: Wang, Yiyang, et al.
Veröffentlicht: (2025)
von: Wang, Yiyang, et al.
Veröffentlicht: (2025)
Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2024)
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2024)
BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration
von: Gao, Bo, et al.
Veröffentlicht: (2026)
von: Gao, Bo, et al.
Veröffentlicht: (2026)
Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models
von: Wang, Zehan, et al.
Veröffentlicht: (2024)
von: Wang, Zehan, et al.
Veröffentlicht: (2024)
Fast Encoding and Decoding for Implicit Video Representation
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
Scene Co-pilot: Procedural Text to Video Generation with Human in the Loop
von: Qian, Zhaofang, et al.
Veröffentlicht: (2024)
von: Qian, Zhaofang, et al.
Veröffentlicht: (2024)
Video Decomposition Prior: A Methodology to Decompose Videos into Layers
von: Shrivastava, Gaurav, et al.
Veröffentlicht: (2024)
von: Shrivastava, Gaurav, et al.
Veröffentlicht: (2024)
FashionComposer: Compositional Fashion Image Generation
von: Ji, Sihui, et al.
Veröffentlicht: (2024)
von: Ji, Sihui, et al.
Veröffentlicht: (2024)
UniLION: Towards Unified Autonomous Driving Model with Linear Group RNNs
von: Liu, Zhe, et al.
Veröffentlicht: (2025)
von: Liu, Zhe, et al.
Veröffentlicht: (2025)
CoopDETR: A Unified Cooperative Perception Framework for 3D Detection via Object Query
von: Wang, Zhe, et al.
Veröffentlicht: (2025)
von: Wang, Zhe, et al.
Veröffentlicht: (2025)
Seamless Detection: Unifying Salient Object Detection and Camouflaged Object Detection
von: Liu, Yi, et al.
Veröffentlicht: (2024)
von: Liu, Yi, et al.
Veröffentlicht: (2024)
GDRO: Group-level Reward Post-training Suitable for Diffusion Models
von: Wang, Yiyang, et al.
Veröffentlicht: (2026)
von: Wang, Yiyang, et al.
Veröffentlicht: (2026)
RoboFusion: Towards Robust Multi-Modal 3D Object Detection via SAM
von: Song, Ziying, et al.
Veröffentlicht: (2024)
von: Song, Ziying, et al.
Veröffentlicht: (2024)
DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses
von: Pang, Yatian, et al.
Veröffentlicht: (2024)
von: Pang, Yatian, et al.
Veröffentlicht: (2024)
Depth Anything V2
von: Yang, Lihe, et al.
Veröffentlicht: (2024)
von: Yang, Lihe, et al.
Veröffentlicht: (2024)
An Open and Comprehensive Pipeline for Unified Object Grounding and Detection
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2024)
Language-free Compositional Action Generation via Decoupling Refinement
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
Towards Cross-modal Backward-compatible Representation Learning for Vision-Language Models
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
Towards Stable 3D Object Detection
von: Wang, Jiabao, et al.
Veröffentlicht: (2024)
von: Wang, Jiabao, et al.
Veröffentlicht: (2024)
DiffDoctor: Diagnosing Image Diffusion Models Before Treating
von: Wang, Yiyang, et al.
Veröffentlicht: (2025)
von: Wang, Yiyang, et al.
Veröffentlicht: (2025)
TranSplat: Generalizable 3D Gaussian Splatting from Sparse Multi-View Images with Transformers
von: Zhang, Chuanrui, et al.
Veröffentlicht: (2024)
von: Zhang, Chuanrui, et al.
Veröffentlicht: (2024)
SUP-NeRF: A Streamlined Unification of Pose Estimation and NeRF for Monocular 3D Object Reconstruction
von: Guo, Yuliang, et al.
Veröffentlicht: (2024)
von: Guo, Yuliang, et al.
Veröffentlicht: (2024)
FaceCom: Towards High-fidelity 3D Facial Shape Completion via Optimization and Inpainting Guidance
von: Li, Yinglong, et al.
Veröffentlicht: (2024)
von: Li, Yinglong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence
von: Li, Zhuoling, et al.
Veröffentlicht: (2024) -
DreamMask: Boosting Open-vocabulary Panoptic Segmentation with Synthetic Data
von: Tu, Yuanpeng, et al.
Veröffentlicht: (2025) -
OV-Uni3DETR: Towards Unified Open-Vocabulary 3D Object Detection via Cycle-Modality Propagation
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024) -
Towards Chunk-Wise Generation for Long Videos
von: Zhang, Siyang, et al.
Veröffentlicht: (2024) -
VideoMerge: Towards Training-free Long Video Generation
von: Zhang, Siyang, et al.
Veröffentlicht: (2025)