Revisiting Misalignment in Multispectral Pedestrian Detection: A Language-Driven Approach for Cross-modal Alignment Fusion
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kim, Taeheon, Chung, Sangyun, Yu, Youngjoon, Ro, Yong Man |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
MSCoTDet: Language-driven Multi-modal Fusion for Improved Multispectral Pedestrian Detection
par: Kim, Taeheon, et autres
Publié: (2024)
par: Kim, Taeheon, et autres
Publié: (2024)
Causal Mode Multiplexer: A Novel Framework for Unbiased Multispectral Pedestrian Detection
par: Kim, Taeheon, et autres
Publié: (2024)
par: Kim, Taeheon, et autres
Publié: (2024)
SPARK: Multi-Vision Sensor Perception and Reasoning Benchmark for Large-scale Vision-Language Models
par: Yu, Youngjoon, et autres
Publié: (2024)
par: Yu, Youngjoon, et autres
Publié: (2024)
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking
par: Chung, Sangyun, et autres
Publié: (2024)
par: Chung, Sangyun, et autres
Publié: (2024)
Robust Pedestrian Detection via Constructing Versatile Pedestrian Knowledge Bank
par: Park, Sungjune, et autres
Publié: (2024)
par: Park, Sungjune, et autres
Publié: (2024)
Integrating Language-Derived Appearance Elements with Visual Cues in Pedestrian Detection
par: Park, Sungjune, et autres
Publié: (2023)
par: Park, Sungjune, et autres
Publié: (2023)
Phantom of Latent for Large Language and Vision Models
par: Lee, Byung-Kwan, et autres
Publié: (2024)
par: Lee, Byung-Kwan, et autres
Publié: (2024)
Strip-Fusion: Spatiotemporal Fusion for Multispectral Pedestrian Detection
par: Kanu-Asiegbu, Asiegbu Miracle, et autres
Publié: (2026)
par: Kanu-Asiegbu, Asiegbu Miracle, et autres
Publié: (2026)
TroL: Traversal of Layers for Large Language and Vision Models
par: Lee, Byung-Kwan, et autres
Publié: (2024)
par: Lee, Byung-Kwan, et autres
Publié: (2024)
Multispectral Pedestrian Detection with Sparsely Annotated Label
par: Lee, Chan, et autres
Publié: (2025)
par: Lee, Chan, et autres
Publié: (2025)
AMFD: Distillation via Adaptive Multimodal Fusion for Multispectral Pedestrian Detection
par: Chen, Zizhao, et autres
Publié: (2024)
par: Chen, Zizhao, et autres
Publié: (2024)
Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling
par: Park, Sungjune, et autres
Publié: (2025)
par: Park, Sungjune, et autres
Publié: (2025)
MS-DETR: Multispectral Pedestrian Detection Transformer with Loosely Coupled Fusion and Modality-Balanced Optimization
par: Xing, Yinghui, et autres
Publié: (2023)
par: Xing, Yinghui, et autres
Publié: (2023)
GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory
par: Yeo, Jeong Hun, et autres
Publié: (2025)
par: Yeo, Jeong Hun, et autres
Publié: (2025)
Language-guided Learning for Object Detection Tackling Multiple Variations in Aerial Images
par: Park, Sungjune, et autres
Publié: (2025)
par: Park, Sungjune, et autres
Publié: (2025)
CODE: Contrasting Self-generated Description to Combat Hallucination in Large Multi-modal Models
par: Kim, Junho, et autres
Publié: (2024)
par: Kim, Junho, et autres
Publié: (2024)
What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-modal Models
par: Kim, Junho, et autres
Publié: (2024)
par: Kim, Junho, et autres
Publié: (2024)
Deep Understanding of Sign Language for Sign to Subtitle Alignment
par: Jang, Youngjoon, et autres
Publié: (2025)
par: Jang, Youngjoon, et autres
Publié: (2025)
WCCNet: Wavelet-context Cooperative Network for Efficient Multispectral Pedestrian Detection
par: Wang, Xingjian, et autres
Publié: (2023)
par: Wang, Xingjian, et autres
Publié: (2023)
Cross-modal Offset-guided Dynamic Alignment and Fusion for Weakly Aligned UAV Object Detection
par: Zongzhen, Liu, et autres
Publié: (2025)
par: Zongzhen, Liu, et autres
Publié: (2025)
TFDet: Target-Aware Fusion for RGB-T Pedestrian Detection
par: Zhang, Xue, et autres
Publié: (2023)
par: Zhang, Xue, et autres
Publié: (2023)
Lost in Translation, Found in Embeddings: Sign Language Translation and Alignment
par: Jang, Youngjoon, et autres
Publié: (2025)
par: Jang, Youngjoon, et autres
Publié: (2025)
Multispectral State-Space Feature Fusion: Bridging Shared and Cross-Parametric Interactions for Object Detection
par: Shen, Jifeng, et autres
Publié: (2025)
par: Shen, Jifeng, et autres
Publié: (2025)
Multispectral Detection Transformer with Infrared-Centric Feature Fusion
par: Hwang, Seongmin, et autres
Publié: (2025)
par: Hwang, Seongmin, et autres
Publié: (2025)
Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models
par: Lee, Byung-Kwan, et autres
Publié: (2024)
par: Lee, Byung-Kwan, et autres
Publié: (2024)
CoLLaVO: Crayon Large Language and Vision mOdel
par: Lee, Byung-Kwan, et autres
Publié: (2024)
par: Lee, Byung-Kwan, et autres
Publié: (2024)
MoAI: Mixture of All Intelligence for Large Language and Vision Models
par: Lee, Byung-Kwan, et autres
Publié: (2024)
par: Lee, Byung-Kwan, et autres
Publié: (2024)
Robust Egocentric Visual Attention Prediction Through Language-guided Scene Context-aware Learning
par: Park, Sungjune, et autres
Publié: (2026)
par: Park, Sungjune, et autres
Publié: (2026)
Rethinking Early-Fusion Strategies for Improved Multispectral Object Detection
par: Zhang, Xue, et autres
Publié: (2024)
par: Zhang, Xue, et autres
Publié: (2024)
CSAKD: Knowledge Distillation with Cross Self-Attention for Hyperspectral and Multispectral Image Fusion
par: Hsu, Chih-Chung, et autres
Publié: (2024)
par: Hsu, Chih-Chung, et autres
Publié: (2024)
Pedestrian Crossing Intent Prediction via Psychological Features and Transformer Fusion
par: Ashayer, Sima, et autres
Publié: (2026)
par: Ashayer, Sima, et autres
Publié: (2026)
Cross-modal Full-mode Fine-grained Alignment for Text-to-Image Person Retrieval
par: Yin, Hao, et autres
Publié: (2025)
par: Yin, Hao, et autres
Publié: (2025)
DIP-R1: Deep Inspection and Perception with RL Looking Through and Understanding Complex Scenes
par: Park, Sungjune, et autres
Publié: (2025)
par: Park, Sungjune, et autres
Publié: (2025)
AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding
par: Jung, Chaeyoung, et autres
Publié: (2025)
par: Jung, Chaeyoung, et autres
Publié: (2025)
Fusion-Mamba for Cross-modality Object Detection
par: Dong, Wenhao, et autres
Publié: (2024)
par: Dong, Wenhao, et autres
Publié: (2024)
Fourier-enhanced Implicit Neural Fusion Network for Multispectral and Hyperspectral Image Fusion
par: Liang, Yu-Jie, et autres
Publié: (2024)
par: Liang, Yu-Jie, et autres
Publié: (2024)
Empathetic Response in Audio-Visual Conversations Using Emotion Preference Optimization and MambaCompressor
par: Kim, Yeonju, et autres
Publié: (2024)
par: Kim, Yeonju, et autres
Publié: (2024)
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis
par: Kim, Junho, et autres
Publié: (2024)
par: Kim, Junho, et autres
Publié: (2024)
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding
par: Lee, Hosu, et autres
Publié: (2025)
par: Lee, Hosu, et autres
Publié: (2025)
DLRMamba: Distilling Low-Rank Mamba for Edge Multispectral Fusion Object Detection
par: Zhang, Qianqian, et autres
Publié: (2026)
par: Zhang, Qianqian, et autres
Publié: (2026)
Documents similaires
-
MSCoTDet: Language-driven Multi-modal Fusion for Improved Multispectral Pedestrian Detection
par: Kim, Taeheon, et autres
Publié: (2024) -
Causal Mode Multiplexer: A Novel Framework for Unbiased Multispectral Pedestrian Detection
par: Kim, Taeheon, et autres
Publié: (2024) -
SPARK: Multi-Vision Sensor Perception and Reasoning Benchmark for Large-scale Vision-Language Models
par: Yu, Youngjoon, et autres
Publié: (2024) -
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking
par: Chung, Sangyun, et autres
Publié: (2024) -
Robust Pedestrian Detection via Constructing Versatile Pedestrian Knowledge Bank
par: Park, Sungjune, et autres
Publié: (2024)