CAFuser: Condition-Aware Multimodal Fusion for Robust Semantic Perception of Driving Scenes
Fuente:
arXiv
Saved in:
| Main Authors: | Broedermann, Tim, Sakaridis, Christos, Fu, Yuqian, Van Gool, Luc |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PBR-NeRF: Inverse Rendering with Physics-Based Neural Fields
by: Wu, Sean, et al.
Published: (2024)
by: Wu, Sean, et al.
Published: (2024)
DGFusion: Depth-Guided Sensor Fusion for Robust Semantic Perception
by: Broedermannn, Tim, et al.
Published: (2025)
by: Broedermannn, Tim, et al.
Published: (2025)
ACDC: The Adverse Conditions Dataset with Correspondences for Robust Semantic Driving Scene Perception
by: Sakaridis, Christos, et al.
Published: (2021)
by: Sakaridis, Christos, et al.
Published: (2021)
Sun Off, Lights On: Photorealistic Monocular Nighttime Simulation for Robust Semantic Perception
by: Tzevelekakis, Konstantinos, et al.
Published: (2024)
by: Tzevelekakis, Konstantinos, et al.
Published: (2024)
Vanishing-Point-Guided Video Semantic Segmentation of Driving Scenes
by: Guo, Diandian, et al.
Published: (2024)
by: Guo, Diandian, et al.
Published: (2024)
Condition-Invariant Semantic Segmentation
by: Sakaridis, Christos, et al.
Published: (2023)
by: Sakaridis, Christos, et al.
Published: (2023)
MUSES: The Multi-Sensor Semantic Perception Dataset for Driving under Uncertainty
by: Brödermann, Tim, et al.
Published: (2024)
by: Brödermann, Tim, et al.
Published: (2024)
Optimizing against Infeasible Inclusions from Data for Semantic Segmentation through Morphology
by: Basu, Shamik, et al.
Published: (2024)
by: Basu, Shamik, et al.
Published: (2024)
TrafficBots V1.5: Traffic Simulation via Conditional VAEs and Transformers with Relative Pose Encoding
by: Zhang, Zhejun, et al.
Published: (2024)
by: Zhang, Zhejun, et al.
Published: (2024)
Bayesian Self-Training for Semi-Supervised 3D Segmentation
by: Unal, Ozan, et al.
Published: (2024)
by: Unal, Ozan, et al.
Published: (2024)
Four Ways to Improve Verbo-visual Fusion for Dense 3D Visual Grounding
by: Unal, Ozan, et al.
Published: (2023)
by: Unal, Ozan, et al.
Published: (2023)
Fine-Grained Spatial and Verbal Losses for 3D Visual Grounding
by: Dey, Sombit, et al.
Published: (2024)
by: Dey, Sombit, et al.
Published: (2024)
Advances in Deep Concealed Scene Understanding
by: Fan, Deng-Ping, et al.
Published: (2023)
by: Fan, Deng-Ping, et al.
Published: (2023)
Video Depth Propagation
by: Piccinelli, Luigi, et al.
Published: (2025)
by: Piccinelli, Luigi, et al.
Published: (2025)
Radar Fields: Frequency-Space Neural Scene Representations for FMCW Radar
by: Borts, David, et al.
Published: (2024)
by: Borts, David, et al.
Published: (2024)
Vision encoders should be image size agnostic and task driven
by: Prisadnikov, Nedyalko, et al.
Published: (2025)
by: Prisadnikov, Nedyalko, et al.
Published: (2025)
Self-supervised pretraining for an iterative image size agnostic vision transformer
by: Prisadnikov, Nedyalko, et al.
Published: (2026)
by: Prisadnikov, Nedyalko, et al.
Published: (2026)
UniDepth: Universal Monocular Metric Depth Estimation
by: Piccinelli, Luigi, et al.
Published: (2024)
by: Piccinelli, Luigi, et al.
Published: (2024)
UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler
by: Piccinelli, Luigi, et al.
Published: (2025)
by: Piccinelli, Luigi, et al.
Published: (2025)
UniK3D: Universal Camera Monocular 3D Estimation
by: Piccinelli, Luigi, et al.
Published: (2025)
by: Piccinelli, Luigi, et al.
Published: (2025)
You Only Train Once
by: Sakaridis, Christos
Published: (2025)
by: Sakaridis, Christos
Published: (2025)
Camera-Only 3D Panoptic Scene Completion for Autonomous Driving through Differentiable Object Shapes
by: Marinello, Nicola, et al.
Published: (2025)
by: Marinello, Nicola, et al.
Published: (2025)
Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025
by: Fu, Yuqian, et al.
Published: (2025)
by: Fu, Yuqian, et al.
Published: (2025)
Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis
by: Mahdi, Mohammad, et al.
Published: (2025)
by: Mahdi, Mohammad, et al.
Published: (2025)
A Unified Framework for Event-based Frame Interpolation with Ad-hoc Deblurring in the Wild
by: Sun, Lei, et al.
Published: (2023)
by: Sun, Lei, et al.
Published: (2023)
EgoSpot:Egocentric Multimodal Control for Hands-Free Mobile Manipulation
by: Zhang, Ganlin, et al.
Published: (2023)
by: Zhang, Ganlin, et al.
Published: (2023)
SLAck: Semantic, Location, and Appearance Aware Open-Vocabulary Tracking
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
Inferring Compositional 4D Scenes without Ever Seeing One
by: Gokmen, Ahmet Berke, et al.
Published: (2025)
by: Gokmen, Ahmet Berke, et al.
Published: (2025)
BiXFormer: A Robust Framework for Maximizing Modality Effectiveness in Multi-Modal Semantic Segmentation
by: Chen, Jialei, et al.
Published: (2025)
by: Chen, Jialei, et al.
Published: (2025)
Implicit-Zoo: A Large-Scale Dataset of Neural Implicit Functions for 2D Images and 3D Scenes
by: Ma, Qi, et al.
Published: (2024)
by: Ma, Qi, et al.
Published: (2024)
XTrack: Multimodal Training Boosts RGB-X Video Object Trackers
by: Tan, Yuedong, et al.
Published: (2024)
by: Tan, Yuedong, et al.
Published: (2024)
Occluded nuScenes: A Multi-Sensor Dataset for Evaluating Perception Robustness in Automated Driving
by: Kumar, Sanjay, et al.
Published: (2025)
by: Kumar, Sanjay, et al.
Published: (2025)
Robust Promptable Video Object Segmentation
by: Lee, Sohyun, et al.
Published: (2026)
by: Lee, Sohyun, et al.
Published: (2026)
EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering
by: Li, Yanjun, et al.
Published: (2025)
by: Li, Yanjun, et al.
Published: (2025)
VF-NeRF: Learning Neural Vector Fields for Indoor Scene Reconstruction
by: Puigjaner, Albert Gassol, et al.
Published: (2024)
by: Puigjaner, Albert Gassol, et al.
Published: (2024)
Test-time Training for Hyperspectral Image Super-resolution
by: Li, Ke, et al.
Published: (2024)
by: Li, Ke, et al.
Published: (2024)
MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
by: Segu, Mattia, et al.
Published: (2025)
by: Segu, Mattia, et al.
Published: (2025)
Locate Anything on Earth: Advancing Open-Vocabulary Object Detection for Remote Sensing Community
by: Pan, Jiancheng, et al.
Published: (2024)
by: Pan, Jiancheng, et al.
Published: (2024)
Accelerating Vision Foundation Models with Drop-in Depthwise Convolution
by: Scribano, Carmelo, et al.
Published: (2026)
by: Scribano, Carmelo, et al.
Published: (2026)
DGInStyle: Domain-Generalizable Semantic Segmentation with Image Diffusion Models and Stylized Semantic Control
by: Jia, Yuru, et al.
Published: (2023)
by: Jia, Yuru, et al.
Published: (2023)
Similar Items
-
PBR-NeRF: Inverse Rendering with Physics-Based Neural Fields
by: Wu, Sean, et al.
Published: (2024) -
DGFusion: Depth-Guided Sensor Fusion for Robust Semantic Perception
by: Broedermannn, Tim, et al.
Published: (2025) -
ACDC: The Adverse Conditions Dataset with Correspondences for Robust Semantic Driving Scene Perception
by: Sakaridis, Christos, et al.
Published: (2021) -
Sun Off, Lights On: Photorealistic Monocular Nighttime Simulation for Robust Semantic Perception
by: Tzevelekakis, Konstantinos, et al.
Published: (2024) -
Vanishing-Point-Guided Video Semantic Segmentation of Driving Scenes
by: Guo, Diandian, et al.
Published: (2024)