UAV traffic scene understanding: A regulation embedded multi-modal network and a unified benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yu, Zhao, Zhicheng, Luo, Ze, Li, Chenglong, Tang, Jin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generating metamers of human scene understanding
by: Raina, Ritik, et al.
Published: (2026)
by: Raina, Ritik, et al.
Published: (2026)
Towards Robust Optical-SAR Object Detection under Missing Modalities: A Dynamic Quality-Aware Fusion Framework
by: Zhao, Zhicheng, et al.
Published: (2025)
by: Zhao, Zhicheng, et al.
Published: (2025)
DCG ReID: Disentangling Collaboration and Guidance Fusion Representations for Multi-modal Vehicle Re-Identification
by: Zheng, Aihua, et al.
Published: (2026)
by: Zheng, Aihua, et al.
Published: (2026)
Near, far: Patch-ordering enhances vision foundation models' scene understanding
by: Pariza, Valentinos, et al.
Published: (2024)
by: Pariza, Valentinos, et al.
Published: (2024)
Quantifying the synthetic and real domain gap in aerial scene understanding
by: Marcu, Alina
Published: (2024)
by: Marcu, Alina
Published: (2024)
Large Language Model Guided Progressive Feature Alignment for Multimodal UAV Object Detection
by: Wu, Wentao, et al.
Published: (2025)
by: Wu, Wentao, et al.
Published: (2025)
Open-set object detection: towards unified problem formulation and benchmarking
by: Ammar, Hejer, et al.
Published: (2024)
by: Ammar, Hejer, et al.
Published: (2024)
Omni-SILA: Towards Omni-scene Driven Visual Sentiment Identifying, Locating and Attributing in Videos
by: Luo, Jiamin, et al.
Published: (2025)
by: Luo, Jiamin, et al.
Published: (2025)
From Pixels to Predicates Structuring urban perception with scene graphs
by: Liu, Yunlong, et al.
Published: (2025)
by: Liu, Yunlong, et al.
Published: (2025)
STEI-PCN: an efficient pure convolutional network for traffic prediction via spatial-temporal encoding and inferring
by: Hu, Kai, et al.
Published: (2025)
by: Hu, Kai, et al.
Published: (2025)
DEF-oriCORN: efficient 3D scene understanding for robust language-directed manipulation without demonstrations
by: Son, Dongwon, et al.
Published: (2024)
by: Son, Dongwon, et al.
Published: (2024)
Assessing the generalization performance of SAM for ureteroscopy scene understanding
by: Villagrana, Martin, et al.
Published: (2025)
by: Villagrana, Martin, et al.
Published: (2025)
Pedestrian Attribute Recognition via CLIP based Prompt Vision-Language Fusion
by: Wang, Xiao, et al.
Published: (2023)
by: Wang, Xiao, et al.
Published: (2023)
PMMD: A pose-guided multi-view multi-modal diffusion for person generation
by: Shang, Ziyu, et al.
Published: (2025)
by: Shang, Ziyu, et al.
Published: (2025)
Sherlock: Towards Multi-scene Video Abnormal Event Extraction and Localization via a Global-local Spatial-sensitive LLM
by: Ma, Junxiao, et al.
Published: (2025)
by: Ma, Junxiao, et al.
Published: (2025)
EyeCLIP: A visual-language foundation model for multi-modal ophthalmic image analysis
by: Shi, Danli, et al.
Published: (2024)
by: Shi, Danli, et al.
Published: (2024)
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework
by: Wu, Wentao, et al.
Published: (2025)
by: Wu, Wentao, et al.
Published: (2025)
Efficient endometrial carcinoma screening via cross-modal synthesis and gradient distillation
by: Shan, Dongjing, et al.
Published: (2026)
by: Shan, Dongjing, et al.
Published: (2026)
Less yet robust: crucial region selection for scene recognition
by: Zhang, Jianqi, et al.
Published: (2024)
by: Zhang, Jianqi, et al.
Published: (2024)
MMM-RS: A Multi-modal, Multi-GSD, Multi-scene Remote Sensing Dataset and Benchmark for Text-to-Image Generation
by: Luo, Jialin, et al.
Published: (2024)
by: Luo, Jialin, et al.
Published: (2024)
Vehicle-centric Perception via Multimodal Structured Pre-training
by: Wu, Wentao, et al.
Published: (2025)
by: Wu, Wentao, et al.
Published: (2025)
ICFNet: Integrated Cross-modal Fusion Network for Survival Prediction
by: Zhang, Binyu, et al.
Published: (2025)
by: Zhang, Binyu, et al.
Published: (2025)
DialBench: Towards Accurate Reading Recognition of Pointer Meter using Large Foundation Models
by: Wang, Futian, et al.
Published: (2025)
by: Wang, Futian, et al.
Published: (2025)
Graph-based Semantic Calibration Network for Unaligned UAV RGBT Image Semantic Segmentation and A Large-scale Benchmark
by: Fan, Fangqiang, et al.
Published: (2026)
by: Fan, Fangqiang, et al.
Published: (2026)
e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings
by: Chen, Haonan, et al.
Published: (2026)
by: Chen, Haonan, et al.
Published: (2026)
Human-annotated label noise and their impact on ConvNets for remote sensing image scene classification
by: Peng, Longkang, et al.
Published: (2023)
by: Peng, Longkang, et al.
Published: (2023)
KAN-RCBEVDepth: A multi-modal fusion algorithm in object detection for autonomous driving
by: Lai, Zhihao, et al.
Published: (2024)
by: Lai, Zhihao, et al.
Published: (2024)
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
by: Liang, Yujia, et al.
Published: (2025)
by: Liang, Yujia, et al.
Published: (2025)
Swift4D:Adaptive divide-and-conquer Gaussian Splatting for compact and efficient reconstruction of dynamic scene
by: Wu, Jiahao, et al.
Published: (2025)
by: Wu, Jiahao, et al.
Published: (2025)
A benchmark multimodal oro-dental dataset for large vision-language models
by: Lv, Haoxin, et al.
Published: (2025)
by: Lv, Haoxin, et al.
Published: (2025)
mChartQA: A universal benchmark for multimodal Chart Question Answer based on Vision-Language Alignment and Reasoning
by: Wei, Jingxuan, et al.
Published: (2024)
by: Wei, Jingxuan, et al.
Published: (2024)
Fusion-Mamba for Cross-modality Object Detection
by: Dong, Wenhao, et al.
Published: (2024)
by: Dong, Wenhao, et al.
Published: (2024)
Physics-Constrained Cross-Resolution Enhancement Network for Optics-Guided Thermal UAV Image Super-Resolution
by: Zhao, Zhicheng, et al.
Published: (2026)
by: Zhao, Zhicheng, et al.
Published: (2026)
Zero-Shot Image Moderation in Google Ads with LLM-Assisted Textual Descriptions and Cross-modal Co-embeddings
by: Luo, Enming, et al.
Published: (2024)
by: Luo, Enming, et al.
Published: (2024)
An Empirical Study of Mamba-based Pedestrian Attribute Recognition
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
by: Liu, Chonghan, et al.
Published: (2025)
by: Liu, Chonghan, et al.
Published: (2025)
RAMer: Reconstruction-based Adversarial Model for Multi-party Multi-modal Multi-label Emotion Recognition
by: Yang, Xudong, et al.
Published: (2025)
by: Yang, Xudong, et al.
Published: (2025)
Compositional-Degradation UAV Image Restoration: Conditional Decoupled MoE Network and A Benchmark
by: Yan, Jinquan, et al.
Published: (2026)
by: Yan, Jinquan, et al.
Published: (2026)
Assessment of Sentinel-2 spatial and temporal coverage based on the scene classification layer
by: Sanchez, Cristhian, et al.
Published: (2024)
by: Sanchez, Cristhian, et al.
Published: (2024)
All-day Multi-scenes Lifelong Vision-and-Language Navigation with Tucker Adaptation
by: Wang, Xudong, et al.
Published: (2026)
by: Wang, Xudong, et al.
Published: (2026)
Similar Items
-
Generating metamers of human scene understanding
by: Raina, Ritik, et al.
Published: (2026) -
Towards Robust Optical-SAR Object Detection under Missing Modalities: A Dynamic Quality-Aware Fusion Framework
by: Zhao, Zhicheng, et al.
Published: (2025) -
DCG ReID: Disentangling Collaboration and Guidance Fusion Representations for Multi-modal Vehicle Re-Identification
by: Zheng, Aihua, et al.
Published: (2026) -
Near, far: Patch-ordering enhances vision foundation models' scene understanding
by: Pariza, Valentinos, et al.
Published: (2024) -
Quantifying the synthetic and real domain gap in aerial scene understanding
by: Marcu, Alina
Published: (2024)