Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Jeonghyeon, Hwang, Sangheum |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Localized Concept Erasure in Text-to-Image Diffusion Models via High-Level Representation Misdirection
by: Lee, Uichan, et al.
Published: (2026)
by: Lee, Uichan, et al.
Published: (2026)
Reflexive Guidance: Improving OoDD in Vision-Language Models via Self-Guided Image-Adaptive Concept Generation
by: Kim, Jihyo, et al.
Published: (2024)
by: Kim, Jihyo, et al.
Published: (2024)
Objectomaly: Objectness-Aware Refinement for OoD Segmentation with Structural Consistency and Boundary Precision
by: Song, Jeonghoon, et al.
Published: (2025)
by: Song, Jeonghoon, et al.
Published: (2025)
Logit Disagreement: OoD Detection with Bayesian Neural Networks
by: Raina, Kevin
Published: (2025)
by: Raina, Kevin
Published: (2025)
BAM: Box Abstraction Monitors for Real-time OoD Detection in Object Detection
by: Wu, Changshun, et al.
Published: (2024)
by: Wu, Changshun, et al.
Published: (2024)
Safe and Robust Watermark Injection with a Single OoD Image
by: Yu, Shuyang, et al.
Published: (2023)
by: Yu, Shuyang, et al.
Published: (2023)
Parallel Rescaling: Rebalancing Consistency Guidance for Personalized Diffusion Models
by: Chae, JungWoo, et al.
Published: (2025)
by: Chae, JungWoo, et al.
Published: (2025)
Enhancing Multimodal Unified Representations for Cross Modal Generalization
by: Huang, Hai, et al.
Published: (2024)
by: Huang, Hai, et al.
Published: (2024)
MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and Retention
by: Wang, Tianyi, et al.
Published: (2025)
by: Wang, Tianyi, et al.
Published: (2025)
Look Before You Fuse: 2D-Guided Cross-Modal Alignment for Robust 3D Detection
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models
by: Yan, Hanqi, et al.
Published: (2025)
by: Yan, Hanqi, et al.
Published: (2025)
Unifying Visual and Semantic Feature Spaces with Diffusion Models for Enhanced Cross-Modal Alignment
by: Zheng, Yuze, et al.
Published: (2024)
by: Zheng, Yuze, et al.
Published: (2024)
MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
by: Wang, Jiaze, et al.
Published: (2024)
by: Wang, Jiaze, et al.
Published: (2024)
MIAR: Modality Interaction and Alignment Representation Fuison for Multimodal Emotion
by: Zhu, Jichao, et al.
Published: (2026)
by: Zhu, Jichao, et al.
Published: (2026)
Beyond CLIP: Knowledge-Enhanced Multimodal Transformers for Cross-Modal Alignment in Diabetic Retinopathy Diagnosis
by: Samanta, Argha Kamal, et al.
Published: (2025)
by: Samanta, Argha Kamal, et al.
Published: (2025)
Modelling Visual Semantics via Image Captioning to extract Enhanced Multi-Level Cross-Modal Semantic Incongruity Representation with Attention for Multimodal Sarcasm Detection
by: Aggarwal, Sajal, et al.
Published: (2024)
by: Aggarwal, Sajal, et al.
Published: (2024)
PAN-Crafter: Learning Modality-Consistent Alignment for PAN-Sharpening
by: Do, Jeonghyeok, et al.
Published: (2025)
by: Do, Jeonghyeok, et al.
Published: (2025)
Leveraging Cross-Modal Neighbor Representation for Improved CLIP Classification
by: Yi, Chao, et al.
Published: (2024)
by: Yi, Chao, et al.
Published: (2024)
MMPB: It's Time for Multi-Modal Personalization
by: Kim, Jaeik, et al.
Published: (2025)
by: Kim, Jaeik, et al.
Published: (2025)
CrossModalityDiffusion: Multi-Modal Novel View Synthesis with Unified Intermediate Representation
by: Berian, Alex, et al.
Published: (2025)
by: Berian, Alex, et al.
Published: (2025)
DiffCrossGait: Trajectory-Level Alignment for 2D-3D Cross-Modal Gait Recognition via Latent Diffusion
by: Lu, Zhiyang, et al.
Published: (2026)
by: Lu, Zhiyang, et al.
Published: (2026)
M4-BLIP: Advancing Multi-Modal Media Manipulation Detection through Face-Enhanced Local Analysis
by: Wu, Hang, et al.
Published: (2025)
by: Wu, Hang, et al.
Published: (2025)
AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
by: Zhang, Junyang, et al.
Published: (2025)
by: Zhang, Junyang, et al.
Published: (2025)
Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering
by: Prakash, Nirmalendu, et al.
Published: (2026)
by: Prakash, Nirmalendu, et al.
Published: (2026)
Multi-Prompt with Depth Partitioned Cross-Modal Learning
by: Tian, Yingjie, et al.
Published: (2023)
by: Tian, Yingjie, et al.
Published: (2023)
AnalogRetriever: Learning Cross-Modal Representations for Analog Circuit Retrieval
by: Wang, Yihan, et al.
Published: (2026)
by: Wang, Yihan, et al.
Published: (2026)
Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025
by: Fu, Yuqian, et al.
Published: (2025)
by: Fu, Yuqian, et al.
Published: (2025)
Explicit Correlation Learning for Generalizable Cross-Modal Deepfake Detection
by: Yu, Cai, et al.
Published: (2024)
by: Yu, Cai, et al.
Published: (2024)
Geometry-Aware Cross Modal Alignment for Light Field-LiDAR Semantic Segmentation
by: Luo, Jie, et al.
Published: (2025)
by: Luo, Jie, et al.
Published: (2025)
Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model
by: Feng, Qianhan, et al.
Published: (2024)
by: Feng, Qianhan, et al.
Published: (2024)
Opportunistic Cardiac Health Assessment: Estimating Phenotypes from Localizer MRI through Multi-Modal Representations
by: Zeybek, Busra Nur, et al.
Published: (2026)
by: Zeybek, Busra Nur, et al.
Published: (2026)
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
by: Fang, Xiang, et al.
Published: (2022)
by: Fang, Xiang, et al.
Published: (2022)
Hyperbolic Distillation: Geometry-Guided Cross-Modal Transfer for Robust 3D Object Detection
by: Ning, Kanglin, et al.
Published: (2026)
by: Ning, Kanglin, et al.
Published: (2026)
Cross-Modal Purification and Fusion for Small-Object RGB-D Transmission-Line Defect Detection
by: Cui, Jiaming, et al.
Published: (2026)
by: Cui, Jiaming, et al.
Published: (2026)
Unsupervised Audio-Visual Segmentation with Modality Alignment
by: Bhosale, Swapnil, et al.
Published: (2024)
by: Bhosale, Swapnil, et al.
Published: (2024)
A Generalized Multi-Modal Fusion Detection Framework
by: Cui, Leichao, et al.
Published: (2023)
by: Cui, Leichao, et al.
Published: (2023)
All in One: Exploring Unified Vision-Language Tracking with Multi-Modal Alignment
by: Zhang, Chunhui, et al.
Published: (2023)
by: Zhang, Chunhui, et al.
Published: (2023)
DPG-CD: Depth-Prior-Guided Cross-Modal Joint 2D-3D Change Detection
by: Zhang, Luqi, et al.
Published: (2026)
by: Zhang, Luqi, et al.
Published: (2026)
Cross-Modal Mapping and Dual-Branch Reconstruction for 2D-3D Multimodal Industrial Anomaly Detection
by: Daci, Radia, et al.
Published: (2026)
by: Daci, Radia, et al.
Published: (2026)
Diffusion-Based Restoration for Multi-Modal 3D Object Detection in Adverse Weather
by: He, Zhijian, et al.
Published: (2025)
by: He, Zhijian, et al.
Published: (2025)
Similar Items
-
Localized Concept Erasure in Text-to-Image Diffusion Models via High-Level Representation Misdirection
by: Lee, Uichan, et al.
Published: (2026) -
Reflexive Guidance: Improving OoDD in Vision-Language Models via Self-Guided Image-Adaptive Concept Generation
by: Kim, Jihyo, et al.
Published: (2024) -
Objectomaly: Objectness-Aware Refinement for OoD Segmentation with Structural Consistency and Boundary Precision
by: Song, Jeonghoon, et al.
Published: (2025) -
Logit Disagreement: OoD Detection with Bayesian Neural Networks
by: Raina, Kevin
Published: (2025) -
BAM: Box Abstraction Monitors for Real-time OoD Detection in Object Detection
by: Wu, Changshun, et al.
Published: (2024)