Joint Top-Down and Bottom-Up Frameworks for 3D Visual Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yang, Liu, Daizong, Hu, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Survey on Text-guided 3D Visual Grounding: Elements, Recent Advances, and Future Directions
by: Liu, Daizong, et al.
Published: (2024)
by: Liu, Daizong, et al.
Published: (2024)
Leveraging Bottom-Up and Top-Down Attention for Few-Shot Object Detection
by: Chen, Xianyu, et al.
Published: (2020)
by: Chen, Xianyu, et al.
Published: (2020)
Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding
by: Huang, Wencan, et al.
Published: (2025)
by: Huang, Wencan, et al.
Published: (2025)
Improving the Transferability of 3D Point Cloud Attack via Spectral-aware Admix and Optimization Designs
by: Hu, Shiyu, et al.
Published: (2024)
by: Hu, Shiyu, et al.
Published: (2024)
From Part to Whole: 3D Generative World Model with an Adaptive Structural Hierarchy
by: Du, Bi'an, et al.
Published: (2026)
by: Du, Bi'an, et al.
Published: (2026)
Bodhi VLM: Privacy-Alignment Modeling for Hierarchical Visual Representations in Vision Backbones and VLM Encoders via Bottom-Up and Top-Down Feature Search
by: Ma, Bo, et al.
Published: (2026)
by: Ma, Bo, et al.
Published: (2026)
Top-Down Framework for Weakly-supervised Grounded Image Captioning
by: Cai, Chen, et al.
Published: (2023)
by: Cai, Chen, et al.
Published: (2023)
You Only Look Bottom-Up for Monocular 3D Object Detection
by: Xiong, Kaixin, et al.
Published: (2024)
by: Xiong, Kaixin, et al.
Published: (2024)
Talk2PC: Enhancing 3D Visual Grounding through LiDAR and Radar Point Clouds Fusion for Autonomous Driving
by: Guan, Runwei, et al.
Published: (2025)
by: Guan, Runwei, et al.
Published: (2025)
BUOL: A Bottom-Up Framework with Occupancy-aware Lifting for Panoptic 3D Scene Reconstruction From A Single Image
by: Chu, Tao, et al.
Published: (2023)
by: Chu, Tao, et al.
Published: (2023)
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning
by: li, Bonan, et al.
Published: (2025)
by: li, Bonan, et al.
Published: (2025)
Improving Generalized Visual Grounding with Instance-aware Joint Learning
by: Dai, Ming, et al.
Published: (2025)
by: Dai, Ming, et al.
Published: (2025)
Hard-Label Black-Box Attacks on 3D Point Clouds
by: Liu, Daizong, et al.
Published: (2024)
by: Liu, Daizong, et al.
Published: (2024)
Rethinking Layered Graphic Design Generation with a Top-Down Approach
by: Chen, Jingye, et al.
Published: (2025)
by: Chen, Jingye, et al.
Published: (2025)
A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends
by: Liu, Daizong, et al.
Published: (2024)
by: Liu, Daizong, et al.
Published: (2024)
ChangingGrounding: 3D Visual Grounding in Changing Scenes
by: Hu, Miao, et al.
Published: (2025)
by: Hu, Miao, et al.
Published: (2025)
DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework
by: Zhang, Yani, et al.
Published: (2025)
by: Zhang, Yani, et al.
Published: (2025)
Tracking Skiers from the Top to the Bottom
by: Dunnhofer, Matteo, et al.
Published: (2023)
by: Dunnhofer, Matteo, et al.
Published: (2023)
Towards Top-Down Reasoning: An Explainable Multi-Agent Approach for Visual Question Answering
by: Wang, Zeqing, et al.
Published: (2023)
by: Wang, Zeqing, et al.
Published: (2023)
3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment
by: Li, Xiaoqi, et al.
Published: (2025)
by: Li, Xiaoqi, et al.
Published: (2025)
Joint Fusion and Encoding: Advancing Multimodal Retrieval from the Ground Up
by: Huang, Lang, et al.
Published: (2025)
by: Huang, Lang, et al.
Published: (2025)
LeafTrackNet: A Deep Learning Framework for Robust Leaf Tracking in Top-Down Plant Phenotyping
by: Liu, Shanghua, et al.
Published: (2025)
by: Liu, Shanghua, et al.
Published: (2025)
NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving
by: Li, Fuhao, et al.
Published: (2025)
by: Li, Fuhao, et al.
Published: (2025)
Unified Representation Space for 3D Visual Grounding
by: Zheng, Yinuo, et al.
Published: (2025)
by: Zheng, Yinuo, et al.
Published: (2025)
Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning
by: Wu, Shengqiong, et al.
Published: (2024)
by: Wu, Shengqiong, et al.
Published: (2024)
Top2Pano: Learning to Generate Indoor Panoramas from Top-Down View
by: Zhang, Zitong, et al.
Published: (2025)
by: Zhang, Zitong, et al.
Published: (2025)
Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments
by: Liu, Zhan, et al.
Published: (2026)
by: Liu, Zhan, et al.
Published: (2026)
Top-Down Guidance for Learning Object-Centric Representations
by: Zou, Junhong, et al.
Published: (2024)
by: Zou, Junhong, et al.
Published: (2024)
TD-RD: A Top-Down Benchmark with Real-Time Framework for Road Damage Detection
by: Xiao, Xi, et al.
Published: (2025)
by: Xiao, Xi, et al.
Published: (2025)
Hierarchical Local-Global Transformer for Temporal Sentence Grounding
by: Fang, Xiang, et al.
Published: (2022)
by: Fang, Xiang, et al.
Published: (2022)
An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMs
by: Luo, Zhi, et al.
Published: (2025)
by: Luo, Zhi, et al.
Published: (2025)
Top-Down Semantic Refinement for Image Captioning
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
Open-Text Aerial Detection: A Unified Framework For Aerial Visual Grounding And Detection
by: Wei, Guoting, et al.
Published: (2026)
by: Wei, Guoting, et al.
Published: (2026)
Multi-Pair Temporal Sentence Grounding via Multi-Thread Knowledge Transfer Network
by: Fang, Xiang, et al.
Published: (2024)
by: Fang, Xiang, et al.
Published: (2024)
Reasoning Matters for 3D Visual Grounding
by: Huang, Hsiang-Wei, et al.
Published: (2026)
by: Huang, Hsiang-Wei, et al.
Published: (2026)
A Bottom-Up Approach to Class-Agnostic Image Segmentation
by: Dille, Sebastian, et al.
Published: (2024)
by: Dille, Sebastian, et al.
Published: (2024)
AutoVQA-G: Self-Improving Agentic Framework for Automated Visual Question Answering and Grounding Annotation
by: Hu, Rongsheng, et al.
Published: (2026)
by: Hu, Rongsheng, et al.
Published: (2026)
Unifying Top-down and Bottom-up Scanpath Prediction Using Transformers
by: Yang, Zhibo, et al.
Published: (2023)
by: Yang, Zhibo, et al.
Published: (2023)
Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesis
by: Yang, Yixuan, et al.
Published: (2026)
by: Yang, Yixuan, et al.
Published: (2026)
Similar Items
-
A Survey on Text-guided 3D Visual Grounding: Elements, Recent Advances, and Future Directions
by: Liu, Daizong, et al.
Published: (2024) -
Leveraging Bottom-Up and Top-Down Attention for Few-Shot Object Detection
by: Chen, Xianyu, et al.
Published: (2020) -
Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding
by: Huang, Wencan, et al.
Published: (2025) -
Improving the Transferability of 3D Point Cloud Attack via Spectral-aware Admix and Optimization Designs
by: Hu, Shiyu, et al.
Published: (2024) -
From Part to Whole: 3D Generative World Model with an Adaptive Structural Hierarchy
by: Du, Bi'an, et al.
Published: (2026)