Four Ways to Improve Verbo-visual Fusion for Dense 3D Visual Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Unal, Ozan, Sakaridis, Christos, Saha, Suman, Van Gool, Luc |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fine-Grained Spatial and Verbal Losses for 3D Visual Grounding
by: Dey, Sombit, et al.
Published: (2024)
by: Dey, Sombit, et al.
Published: (2024)
Bayesian Self-Training for Semi-Supervised 3D Segmentation
by: Unal, Ozan, et al.
Published: (2024)
by: Unal, Ozan, et al.
Published: (2024)
CAFuser: Condition-Aware Multimodal Fusion for Robust Semantic Perception of Driving Scenes
by: Broedermann, Tim, et al.
Published: (2024)
by: Broedermann, Tim, et al.
Published: (2024)
Optimizing against Infeasible Inclusions from Data for Semantic Segmentation through Morphology
by: Basu, Shamik, et al.
Published: (2024)
by: Basu, Shamik, et al.
Published: (2024)
TrafficBots V1.5: Traffic Simulation via Conditional VAEs and Transformers with Relative Pose Encoding
by: Zhang, Zhejun, et al.
Published: (2024)
by: Zhang, Zhejun, et al.
Published: (2024)
Language-Guided Instance-Aware Domain-Adaptive Panoptic Segmentation
by: Mansour, Elham Amin, et al.
Published: (2024)
by: Mansour, Elham Amin, et al.
Published: (2024)
Condition-Invariant Semantic Segmentation
by: Sakaridis, Christos, et al.
Published: (2023)
by: Sakaridis, Christos, et al.
Published: (2023)
Sun Off, Lights On: Photorealistic Monocular Nighttime Simulation for Robust Semantic Perception
by: Tzevelekakis, Konstantinos, et al.
Published: (2024)
by: Tzevelekakis, Konstantinos, et al.
Published: (2024)
DGFusion: Depth-Guided Sensor Fusion for Robust Semantic Perception
by: Broedermannn, Tim, et al.
Published: (2025)
by: Broedermannn, Tim, et al.
Published: (2025)
PBR-NeRF: Inverse Rendering with Physics-Based Neural Fields
by: Wu, Sean, et al.
Published: (2024)
by: Wu, Sean, et al.
Published: (2024)
Video Depth Propagation
by: Piccinelli, Luigi, et al.
Published: (2025)
by: Piccinelli, Luigi, et al.
Published: (2025)
UniK3D: Universal Camera Monocular 3D Estimation
by: Piccinelli, Luigi, et al.
Published: (2025)
by: Piccinelli, Luigi, et al.
Published: (2025)
Vanishing-Point-Guided Video Semantic Segmentation of Driving Scenes
by: Guo, Diandian, et al.
Published: (2024)
by: Guo, Diandian, et al.
Published: (2024)
Taming CLIP for Fine-grained and Structured Visual Understanding of Museum Exhibits
by: Balauca, Ada-Astrid, et al.
Published: (2024)
by: Balauca, Ada-Astrid, et al.
Published: (2024)
MUSES: The Multi-Sensor Semantic Perception Dataset for Driving under Uncertainty
by: Brödermann, Tim, et al.
Published: (2024)
by: Brödermann, Tim, et al.
Published: (2024)
UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler
by: Piccinelli, Luigi, et al.
Published: (2025)
by: Piccinelli, Luigi, et al.
Published: (2025)
Advances in Deep Concealed Scene Understanding
by: Fan, Deng-Ping, et al.
Published: (2023)
by: Fan, Deng-Ping, et al.
Published: (2023)
UniDepth: Universal Monocular Metric Depth Estimation
by: Piccinelli, Luigi, et al.
Published: (2024)
by: Piccinelli, Luigi, et al.
Published: (2024)
You Only Train Once
by: Sakaridis, Christos
Published: (2025)
by: Sakaridis, Christos
Published: (2025)
Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models
by: Motamed, Saman, et al.
Published: (2023)
by: Motamed, Saman, et al.
Published: (2023)
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding
by: Zheng, Henry, et al.
Published: (2025)
by: Zheng, Henry, et al.
Published: (2025)
Weakly-Supervised 3D Visual Grounding based on Visual Language Alignment
by: Xu, Xiaoxu, et al.
Published: (2023)
by: Xu, Xiaoxu, et al.
Published: (2023)
ACDC: The Adverse Conditions Dataset with Correspondences for Robust Semantic Driving Scene Perception
by: Sakaridis, Christos, et al.
Published: (2021)
by: Sakaridis, Christos, et al.
Published: (2021)
Referencing Where to Focus: Improving VisualGrounding with Referential Query
by: Wang, Yabing, et al.
Published: (2024)
by: Wang, Yabing, et al.
Published: (2024)
Loopy-SLAM: Dense Neural SLAM with Loop Closures
by: Liso, Lorenzo, et al.
Published: (2024)
by: Liso, Lorenzo, et al.
Published: (2024)
You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
by: Lawrence, Logan, et al.
Published: (2025)
by: Lawrence, Logan, et al.
Published: (2025)
A Unified Framework for Event-based Frame Interpolation with Ad-hoc Deblurring in the Wild
by: Sun, Lei, et al.
Published: (2023)
by: Sun, Lei, et al.
Published: (2023)
Implicit-Zoo: A Large-Scale Dataset of Neural Implicit Functions for 2D Images and 3D Scenes
by: Ma, Qi, et al.
Published: (2024)
by: Ma, Qi, et al.
Published: (2024)
Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction Anticipation
by: Pasca, Razvan-George, et al.
Published: (2023)
by: Pasca, Razvan-George, et al.
Published: (2023)
DenseAnnotate: Enabling Scalable Dense Caption Collection for Images and 3D Scenes via Spoken Descriptions
by: Lin, Xiaoyu, et al.
Published: (2025)
by: Lin, Xiaoyu, et al.
Published: (2025)
HandDiff: 3D Hand Pose Estimation with Diffusion on Image-Point Cloud
by: Cheng, Wencan, et al.
Published: (2024)
by: Cheng, Wencan, et al.
Published: (2024)
See It All: Contextualized Late Aggregation for 3D Dense Captioning
by: Kim, Minjung, et al.
Published: (2024)
by: Kim, Minjung, et al.
Published: (2024)
ScanReason: Empowering 3D Visual Grounding with Reasoning Capabilities
by: Zhu, Chenming, et al.
Published: (2024)
by: Zhu, Chenming, et al.
Published: (2024)
Radar Fields: Frequency-Space Neural Scene Representations for FMCW Radar
by: Borts, David, et al.
Published: (2024)
by: Borts, David, et al.
Published: (2024)
Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
by: Yang, Ziyan, et al.
Published: (2022)
by: Yang, Ziyan, et al.
Published: (2022)
Which Way Does Time Flow? A Psychophysics-Grounded Evaluation for Vision-Language Models
by: Matta, Shiho, et al.
Published: (2025)
by: Matta, Shiho, et al.
Published: (2025)
Camera-Only 3D Panoptic Scene Completion for Autonomous Driving through Differentiable Object Shapes
by: Marinello, Nicola, et al.
Published: (2025)
by: Marinello, Nicola, et al.
Published: (2025)
ViGiL3D: A Linguistically Diverse Dataset for 3D Visual Grounding
by: Wang, Austin T., et al.
Published: (2025)
by: Wang, Austin T., et al.
Published: (2025)
Test-time Training for Hyperspectral Image Super-resolution
by: Li, Ke, et al.
Published: (2024)
by: Li, Ke, et al.
Published: (2024)
Grounding Language Models for Visual Entity Recognition
by: Xiao, Zilin, et al.
Published: (2024)
by: Xiao, Zilin, et al.
Published: (2024)
Similar Items
-
Fine-Grained Spatial and Verbal Losses for 3D Visual Grounding
by: Dey, Sombit, et al.
Published: (2024) -
Bayesian Self-Training for Semi-Supervised 3D Segmentation
by: Unal, Ozan, et al.
Published: (2024) -
CAFuser: Condition-Aware Multimodal Fusion for Robust Semantic Perception of Driving Scenes
by: Broedermann, Tim, et al.
Published: (2024) -
Optimizing against Infeasible Inclusions from Data for Semantic Segmentation through Morphology
by: Basu, Shamik, et al.
Published: (2024) -
TrafficBots V1.5: Traffic Simulation via Conditional VAEs and Transformers with Relative Pose Encoding
by: Zhang, Zhejun, et al.
Published: (2024)