Hierarchical Point Attention for Indoor 3D Object Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shu, Manli, Xue, Le, Yu, Ning, Martín-Martín, Roberto, Xiong, Caiming, Goldstein, Tom, Niebles, Juan Carlos, Xu, Ran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ULIP-2: Towards Scalable Multimodal Pre-training for 3D Understanding
von: Xue, Le, et al.
Veröffentlicht: (2023)
von: Xue, Le, et al.
Veröffentlicht: (2023)
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
von: Ryoo, Michael S., et al.
Veröffentlicht: (2024)
von: Ryoo, Michael S., et al.
Veröffentlicht: (2024)
Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2025)
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2025)
X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2023)
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2023)
LayoutDETR: Detection Transformer Is a Good Multimodal Layout Designer
von: Yu, Ning, et al.
Veröffentlicht: (2022)
von: Yu, Ning, et al.
Veröffentlicht: (2022)
ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models
von: Zhang, Jieyu, et al.
Veröffentlicht: (2024)
von: Zhang, Jieyu, et al.
Veröffentlicht: (2024)
Future Optical Flow Prediction Improves Robot Control & Video Generation
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2026)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2026)
Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models
von: Xu, Yuancheng, et al.
Veröffentlicht: (2024)
von: Xu, Yuancheng, et al.
Veröffentlicht: (2024)
LATTE: Learning to Think with Vision Specialists
von: Ma, Zixian, et al.
Veröffentlicht: (2024)
von: Ma, Zixian, et al.
Veröffentlicht: (2024)
Investigating Domain Gaps for Indoor 3D Object Detection
von: Zhao, Zijing, et al.
Veröffentlicht: (2025)
von: Zhao, Zijing, et al.
Veröffentlicht: (2025)
Coercing LLMs to do and reveal (almost) anything
von: Geiping, Jonas, et al.
Veröffentlicht: (2024)
von: Geiping, Jonas, et al.
Veröffentlicht: (2024)
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
von: Chen, Jiuhai, et al.
Veröffentlicht: (2025)
von: Chen, Jiuhai, et al.
Veröffentlicht: (2025)
Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
von: Fang, Yu, et al.
Veröffentlicht: (2025)
von: Fang, Yu, et al.
Veröffentlicht: (2025)
xGen-VideoSyn-1: High-fidelity Text-to-Video Synthesis with Compressed Representations
von: Qin, Can, et al.
Veröffentlicht: (2024)
von: Qin, Can, et al.
Veröffentlicht: (2024)
Anyview: Generalizable Indoor 3D Object Detection with Variable Frames
von: Wu, Zhenyu, et al.
Veröffentlicht: (2023)
von: Wu, Zhenyu, et al.
Veröffentlicht: (2023)
Open-Vocabulary Indoor Object Grounding with 3D Hierarchical Scene Graph
von: Linok, Sergey, et al.
Veröffentlicht: (2025)
von: Linok, Sergey, et al.
Veröffentlicht: (2025)
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data
von: Zhou, Honglu, et al.
Veröffentlicht: (2025)
von: Zhou, Honglu, et al.
Veröffentlicht: (2025)
Cubify Anything: Scaling Indoor 3D Object Detection
von: Lazarow, Justin, et al.
Veröffentlicht: (2024)
von: Lazarow, Justin, et al.
Veröffentlicht: (2024)
BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions
von: Awadalla, Anas, et al.
Veröffentlicht: (2024)
von: Awadalla, Anas, et al.
Veröffentlicht: (2024)
ViUniT: Visual Unit Tests for More Robust Visual Programming
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2024)
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2024)
PVAFN: Point-Voxel Attention Fusion Network with Multi-Pooling Enhancing for 3D Object Detection
von: Li, Yidi, et al.
Veröffentlicht: (2024)
von: Li, Yidi, et al.
Veröffentlicht: (2024)
UniDet3D: Multi-dataset Indoor 3D Object Detection
von: Kolodiazhnyi, Maksim, et al.
Veröffentlicht: (2024)
von: Kolodiazhnyi, Maksim, et al.
Veröffentlicht: (2024)
MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
von: Awadalla, Anas, et al.
Veröffentlicht: (2024)
von: Awadalla, Anas, et al.
Veröffentlicht: (2024)
Style-Consistent 3D Indoor Scene Synthesis with Decoupled Objects
von: Zhang, Yunfan, et al.
Veröffentlicht: (2024)
von: Zhang, Yunfan, et al.
Veröffentlicht: (2024)
MVSDet: Multi-View Indoor 3D Object Detection via Efficient Plane Sweeps
von: Xu, Yating, et al.
Veröffentlicht: (2024)
von: Xu, Yating, et al.
Veröffentlicht: (2024)
Causal Layering via Conditional Entropy
von: Feigenbaum, Itai, et al.
Veröffentlicht: (2024)
von: Feigenbaum, Itai, et al.
Veröffentlicht: (2024)
Editing Arbitrary Propositions in LLMs without Subject Labels
von: Feigenbaum, Itai, et al.
Veröffentlicht: (2024)
von: Feigenbaum, Itai, et al.
Veröffentlicht: (2024)
Boosting Multi-View Indoor 3D Object Detection via Adaptive 3D Volume Construction
von: Zhang, Runmin, et al.
Veröffentlicht: (2025)
von: Zhang, Runmin, et al.
Veröffentlicht: (2025)
HASSOD: Hierarchical Adaptive Self-Supervised Object Detection
von: Cao, Shengcao, et al.
Veröffentlicht: (2024)
von: Cao, Shengcao, et al.
Veröffentlicht: (2024)
Syn-to-Real Unsupervised Domain Adaptation for Indoor 3D Object Detection
von: Wang, Yunsong, et al.
Veröffentlicht: (2024)
von: Wang, Yunsong, et al.
Veröffentlicht: (2024)
Few-Shot Incremental 3D Object Detection in Dynamic Indoor Environments
von: Zhu, Yun, et al.
Veröffentlicht: (2026)
von: Zhu, Yun, et al.
Veröffentlicht: (2026)
Retroformer: Retrospective Large Language Agents with Policy Gradient Optimization
von: Yao, Weiran, et al.
Veröffentlicht: (2023)
von: Yao, Weiran, et al.
Veröffentlicht: (2023)
REX: Rapid Exploration and eXploitation for AI Agents
von: Murthy, Rithesh, et al.
Veröffentlicht: (2023)
von: Murthy, Rithesh, et al.
Veröffentlicht: (2023)
Trust but Verify: Programmatic VLM Evaluation in the Wild
von: Prabhu, Viraj, et al.
Veröffentlicht: (2024)
von: Prabhu, Viraj, et al.
Veröffentlicht: (2024)
3DGS-DET: Empower 3D Gaussian Splatting with Boundary Guidance and Box-Focused Sampling for Indoor 3D Object Detection
von: Cao, Yang, et al.
Veröffentlicht: (2024)
von: Cao, Yang, et al.
Veröffentlicht: (2024)
Direct Numerical Layout Generation for 3D Indoor Scene Synthesis via Spatial Reasoning
von: Ran, Xingjian, et al.
Veröffentlicht: (2025)
von: Ran, Xingjian, et al.
Veröffentlicht: (2025)
StripDet: Strip Attention-Based Lightweight 3D Object Detection from Point Cloud
von: Wang, Weichao, et al.
Veröffentlicht: (2025)
von: Wang, Weichao, et al.
Veröffentlicht: (2025)
Weakly Supervised Point Clouds Transformer for 3D Object Detection
von: Tang, Zuojin, et al.
Veröffentlicht: (2023)
von: Tang, Zuojin, et al.
Veröffentlicht: (2023)
Are Dense Labels Always Necessary for 3D Object Detection from Point Cloud?
von: Gao, Chenqiang, et al.
Veröffentlicht: (2024)
von: Gao, Chenqiang, et al.
Veröffentlicht: (2024)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ULIP-2: Towards Scalable Multimodal Pre-training for 3D Understanding
von: Xue, Le, et al.
Veröffentlicht: (2023) -
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
von: Ryoo, Michael S., et al.
Veröffentlicht: (2024) -
Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2025) -
X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2023) -
LayoutDETR: Detection Transformer Is a Good Multimodal Layout Designer
von: Yu, Ning, et al.
Veröffentlicht: (2022)