Fine-Grained Spatial and Verbal Losses for 3D Visual Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Dey, Sombit, Unal, Ozan, Sakaridis, Christos, Van Gool, Luc |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Four Ways to Improve Verbo-visual Fusion for Dense 3D Visual Grounding
by: Unal, Ozan, et al.
Published: (2023)
by: Unal, Ozan, et al.
Published: (2023)
Bayesian Self-Training for Semi-Supervised 3D Segmentation
by: Unal, Ozan, et al.
Published: (2024)
by: Unal, Ozan, et al.
Published: (2024)
Optimizing against Infeasible Inclusions from Data for Semantic Segmentation through Morphology
by: Basu, Shamik, et al.
Published: (2024)
by: Basu, Shamik, et al.
Published: (2024)
TrafficBots V1.5: Traffic Simulation via Conditional VAEs and Transformers with Relative Pose Encoding
by: Zhang, Zhejun, et al.
Published: (2024)
by: Zhang, Zhejun, et al.
Published: (2024)
CAFuser: Condition-Aware Multimodal Fusion for Robust Semantic Perception of Driving Scenes
by: Broedermann, Tim, et al.
Published: (2024)
by: Broedermann, Tim, et al.
Published: (2024)
Sun Off, Lights On: Photorealistic Monocular Nighttime Simulation for Robust Semantic Perception
by: Tzevelekakis, Konstantinos, et al.
Published: (2024)
by: Tzevelekakis, Konstantinos, et al.
Published: (2024)
Condition-Invariant Semantic Segmentation
by: Sakaridis, Christos, et al.
Published: (2023)
by: Sakaridis, Christos, et al.
Published: (2023)
ReVLA: Reverting Visual Domain Limitation of Robotic Foundation Models
by: Dey, Sombit, et al.
Published: (2024)
by: Dey, Sombit, et al.
Published: (2024)
PBR-NeRF: Inverse Rendering with Physics-Based Neural Fields
by: Wu, Sean, et al.
Published: (2024)
by: Wu, Sean, et al.
Published: (2024)
Video Depth Propagation
by: Piccinelli, Luigi, et al.
Published: (2025)
by: Piccinelli, Luigi, et al.
Published: (2025)
UniK3D: Universal Camera Monocular 3D Estimation
by: Piccinelli, Luigi, et al.
Published: (2025)
by: Piccinelli, Luigi, et al.
Published: (2025)
Vanishing-Point-Guided Video Semantic Segmentation of Driving Scenes
by: Guo, Diandian, et al.
Published: (2024)
by: Guo, Diandian, et al.
Published: (2024)
DGFusion: Depth-Guided Sensor Fusion for Robust Semantic Perception
by: Broedermannn, Tim, et al.
Published: (2025)
by: Broedermannn, Tim, et al.
Published: (2025)
From Scan to Action: Leveraging Realistic Scans for Embodied Scene Understanding
by: Halacheva, Anna-Maria, et al.
Published: (2025)
by: Halacheva, Anna-Maria, et al.
Published: (2025)
MUSES: The Multi-Sensor Semantic Perception Dataset for Driving under Uncertainty
by: Brödermann, Tim, et al.
Published: (2024)
by: Brödermann, Tim, et al.
Published: (2024)
Language-Guided Instance-Aware Domain-Adaptive Panoptic Segmentation
by: Mansour, Elham Amin, et al.
Published: (2024)
by: Mansour, Elham Amin, et al.
Published: (2024)
UniDepth: Universal Monocular Metric Depth Estimation
by: Piccinelli, Luigi, et al.
Published: (2024)
by: Piccinelli, Luigi, et al.
Published: (2024)
UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler
by: Piccinelli, Luigi, et al.
Published: (2025)
by: Piccinelli, Luigi, et al.
Published: (2025)
Advances in Deep Concealed Scene Understanding
by: Fan, Deng-Ping, et al.
Published: (2023)
by: Fan, Deng-Ping, et al.
Published: (2023)
You Only Train Once
by: Sakaridis, Christos
Published: (2025)
by: Sakaridis, Christos
Published: (2025)
Taming CLIP for Fine-grained and Structured Visual Understanding of Museum Exhibits
by: Balauca, Ada-Astrid, et al.
Published: (2024)
by: Balauca, Ada-Astrid, et al.
Published: (2024)
ACDC: The Adverse Conditions Dataset with Correspondences for Robust Semantic Driving Scene Perception
by: Sakaridis, Christos, et al.
Published: (2021)
by: Sakaridis, Christos, et al.
Published: (2021)
A Unified Framework for Event-based Frame Interpolation with Ad-hoc Deblurring in the Wild
by: Sun, Lei, et al.
Published: (2023)
by: Sun, Lei, et al.
Published: (2023)
Implicit-Zoo: A Large-Scale Dataset of Neural Implicit Functions for 2D Images and 3D Scenes
by: Ma, Qi, et al.
Published: (2024)
by: Ma, Qi, et al.
Published: (2024)
HandDiff: 3D Hand Pose Estimation with Diffusion on Image-Point Cloud
by: Cheng, Wencan, et al.
Published: (2024)
by: Cheng, Wencan, et al.
Published: (2024)
Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Enhancing Fine-Grained Spatial Grounding in 3D CT Report Generation via Discriminative Guidance
by: Wang, Chenyu, et al.
Published: (2026)
by: Wang, Chenyu, et al.
Published: (2026)
Radar Fields: Frequency-Space Neural Scene Representations for FMCW Radar
by: Borts, David, et al.
Published: (2024)
by: Borts, David, et al.
Published: (2024)
Camera-Only 3D Panoptic Scene Completion for Autonomous Driving through Differentiable Object Shapes
by: Marinello, Nicola, et al.
Published: (2025)
by: Marinello, Nicola, et al.
Published: (2025)
Test-time Training for Hyperspectral Image Super-resolution
by: Li, Ke, et al.
Published: (2024)
by: Li, Ke, et al.
Published: (2024)
Burst Image Super-Resolution with Mamba
by: Unal, Ozan, et al.
Published: (2025)
by: Unal, Ozan, et al.
Published: (2025)
SeasonScapes: Learning Large-scale Re-lightable 3D Landscapes with Seasonal Variation from Sparse Webcams
by: Kleger, Timo, et al.
Published: (2026)
by: Kleger, Timo, et al.
Published: (2026)
Know Your Neighbors: Improving Single-View Reconstruction via Spatial Vision-Language Reasoning
by: Li, Rui, et al.
Published: (2024)
by: Li, Rui, et al.
Published: (2024)
LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part Segmentation
by: Miao, Yang, et al.
Published: (2025)
by: Miao, Yang, et al.
Published: (2025)
EvenNICER-SLAM: Event-based Neural Implicit Encoding SLAM
by: Chen, Shi, et al.
Published: (2024)
by: Chen, Shi, et al.
Published: (2024)
SIGHT: Synthesizing Image-Text Conditioned and Geometry-Guided 3D Hand-Object Trajectories
by: Gavryushin, Alexey, et al.
Published: (2025)
by: Gavryushin, Alexey, et al.
Published: (2025)
Inferring Compositional 4D Scenes without Ever Seeing One
by: Gokmen, Ahmet Berke, et al.
Published: (2025)
by: Gokmen, Ahmet Berke, et al.
Published: (2025)
Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding
by: Guo, Leilei, et al.
Published: (2025)
by: Guo, Leilei, et al.
Published: (2025)
SubGrapher: Visual Fingerprinting of Chemical Structures
by: Morin, Lucas, et al.
Published: (2025)
by: Morin, Lucas, et al.
Published: (2025)
GeoVLM-R1: Reinforcement Fine-Tuning for Improved Remote Sensing Reasoning
by: Fiaz, Mustansar, et al.
Published: (2025)
by: Fiaz, Mustansar, et al.
Published: (2025)
Similar Items
-
Four Ways to Improve Verbo-visual Fusion for Dense 3D Visual Grounding
by: Unal, Ozan, et al.
Published: (2023) -
Bayesian Self-Training for Semi-Supervised 3D Segmentation
by: Unal, Ozan, et al.
Published: (2024) -
Optimizing against Infeasible Inclusions from Data for Semantic Segmentation through Morphology
by: Basu, Shamik, et al.
Published: (2024) -
TrafficBots V1.5: Traffic Simulation via Conditional VAEs and Transformers with Relative Pose Encoding
by: Zhang, Zhejun, et al.
Published: (2024) -
CAFuser: Condition-Aware Multimodal Fusion for Robust Semantic Perception of Driving Scenes
by: Broedermann, Tim, et al.
Published: (2024)