Out-of-Distribution Detection with Attention Head Masking for Multimodal Document Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Constantinou, Christos, Ioannides, Georgios, Chadha, Aman, Elkins, Aaron, Simpson, Edwin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Density Adaptive Attention is All You Need: Robust Parameter-Efficient Fine-Tuning Across Multiple Modalities
by: Ioannides, Georgios, et al.
Published: (2024)
by: Ioannides, Georgios, et al.
Published: (2024)
MOD-X: A Modular Open Decentralized eXchange Framework proposal for Heterogeneous Interoperable Artificial Intelligence Agents
by: Ioannides, Georgios, et al.
Published: (2025)
by: Ioannides, Georgios, et al.
Published: (2025)
Enhancing Adverse Drug Event Detection with Multimodal Dataset: Corpus Creation and Model Development
by: Sahoo, Pranab, et al.
Published: (2024)
by: Sahoo, Pranab, et al.
Published: (2024)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
by: Sinha, Neelabh, et al.
Published: (2024)
by: Sinha, Neelabh, et al.
Published: (2024)
Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions
by: Ghosh, Akash, et al.
Published: (2024)
by: Ghosh, Akash, et al.
Published: (2024)
From Fog to Failure: The Unintended Consequences of Dehazing on Object Detection in Clear Images
by: Kumar, Ashutosh, et al.
Published: (2025)
by: Kumar, Ashutosh, et al.
Published: (2025)
The Evolution of Multimodal Model Architectures
by: Wadekar, Shakti N., et al.
Published: (2024)
by: Wadekar, Shakti N., et al.
Published: (2024)
Image-Text Out-Of-Context Detection Using Synthetic Multimodal Misinformation
by: Shalabi, Fatma, et al.
Published: (2024)
by: Shalabi, Fatma, et al.
Published: (2024)
SeTAR: Out-of-Distribution Detection with Selective Low-Rank Approximation
by: Li, Yixia, et al.
Published: (2024)
by: Li, Yixia, et al.
Published: (2024)
MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding
by: Deng, Jingyuan, et al.
Published: (2025)
by: Deng, Jingyuan, et al.
Published: (2025)
How Culturally Aware are Vision-Language Models?
by: Burda-Lassen, Olena, et al.
Published: (2024)
by: Burda-Lassen, Olena, et al.
Published: (2024)
$Δ$-AttnMask: Attention-Guided Masked Hidden States for Efficient Data Selection and Augmentation
by: Hu, Jucheng, et al.
Published: (2025)
by: Hu, Jucheng, et al.
Published: (2025)
Density Adaptive Attention-based Speech Network: Enhancing Feature Understanding for Mental Health Disorders
by: Ioannides, Georgios, et al.
Published: (2024)
by: Ioannides, Georgios, et al.
Published: (2024)
Solving Trojan Detection Competitions with Linear Weight Classification
by: Huster, Todd, et al.
Published: (2024)
by: Huster, Todd, et al.
Published: (2024)
SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection
by: Qi, Peng, et al.
Published: (2024)
by: Qi, Peng, et al.
Published: (2024)
Head Pursuit: Probing Attention Specialization in Multimodal Transformers
by: Basile, Lorenzo, et al.
Published: (2025)
by: Basile, Lorenzo, et al.
Published: (2025)
DPU: Dynamic Prototype Updating for Multimodal Out-of-Distribution Detection
by: Li, Shawn, et al.
Published: (2024)
by: Li, Shawn, et al.
Published: (2024)
Unsupervised Document and Template Clustering using Multimodal Embeddings
by: Sampaio, Phillipe R., et al.
Published: (2025)
by: Sampaio, Phillipe R., et al.
Published: (2025)
SciMDR: Advancing Scientific Multimodal Document Reasoning
by: Chen, Ziyu, et al.
Published: (2026)
by: Chen, Ziyu, et al.
Published: (2026)
CANAMRF: An Attention-Based Model for Multimodal Depression Detection
by: Wei, Yuntao, et al.
Published: (2024)
by: Wei, Yuntao, et al.
Published: (2024)
GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning
by: Ebouky, Brown, et al.
Published: (2026)
by: Ebouky, Brown, et al.
Published: (2026)
DBMF: A Dual-Branch Multimodal Framework for Out-of-Distribution Detection
by: Yue, Jiangbei, et al.
Published: (2026)
by: Yue, Jiangbei, et al.
Published: (2026)
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
by: Mitra, Chancharik, et al.
Published: (2024)
by: Mitra, Chancharik, et al.
Published: (2024)
YesBut: A High-Quality Annotated Multimodal Dataset for evaluating Satire Comprehension capability of Vision-Language Models
by: Nandy, Abhilash, et al.
Published: (2024)
by: Nandy, Abhilash, et al.
Published: (2024)
Confidence-Aware Document OCR Error Detection
by: Hemmer, Arthur, et al.
Published: (2024)
by: Hemmer, Arthur, et al.
Published: (2024)
EmoMeta: A Multimodal Dataset for Fine-grained Emotion Classification in Chinese Metaphors
by: Lu, Xingyuan, et al.
Published: (2025)
by: Lu, Xingyuan, et al.
Published: (2025)
ZINA: Multimodal Fine-grained Hallucination Detection and Editing
by: Wada, Yuiga, et al.
Published: (2025)
by: Wada, Yuiga, et al.
Published: (2025)
Vision Token Reduction via Attention-Driven Self-Compression for Efficient Multimodal Large Language Models
by: Deniz, Omer Faruk, et al.
Published: (2026)
by: Deniz, Omer Faruk, et al.
Published: (2026)
PP-DocBee: Improving Multimodal Document Understanding Through a Bag of Tricks
by: Ni, Feng, et al.
Published: (2025)
by: Ni, Feng, et al.
Published: (2025)
PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding
by: Huang, Kui, et al.
Published: (2025)
by: Huang, Kui, et al.
Published: (2025)
Distributional Uncertainty for Out-of-Distribution Detection
by: Kim, JinYoung, et al.
Published: (2025)
by: Kim, JinYoung, et al.
Published: (2025)
GDCNet: Generative Discrepancy Comparison Network for Multimodal Sarcasm Detection
by: Zhang, Shuguang, et al.
Published: (2026)
by: Zhang, Shuguang, et al.
Published: (2026)
Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads
by: Yeo, Wei Jie, et al.
Published: (2025)
by: Yeo, Wei Jie, et al.
Published: (2025)
Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation
by: Liang, Yupu, et al.
Published: (2025)
by: Liang, Yupu, et al.
Published: (2025)
Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Multimodal Detection of Fake Reviews using BERT and ResNet-50
by: Veluru, Suhasnadh Reddy, et al.
Published: (2025)
by: Veluru, Suhasnadh Reddy, et al.
Published: (2025)
Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models
by: Van, Minh-Hao, et al.
Published: (2025)
by: Van, Minh-Hao, et al.
Published: (2025)
Transcending Domains through Text-to-Image Diffusion: A Source-Free Approach to Domain Adaptation
by: Chopra, Shivang, et al.
Published: (2023)
by: Chopra, Shivang, et al.
Published: (2023)
Efficient Masked Attention Transformer for Few-Shot Classification and Segmentation
by: Carrión-Ojeda, Dustin, et al.
Published: (2025)
by: Carrión-Ojeda, Dustin, et al.
Published: (2025)
DeHate: A Stable Diffusion-based Multimodal Approach to Mitigate Hate Speech in Images
by: Dalal, Dwip, et al.
Published: (2025)
by: Dalal, Dwip, et al.
Published: (2025)
Similar Items
-
Density Adaptive Attention is All You Need: Robust Parameter-Efficient Fine-Tuning Across Multiple Modalities
by: Ioannides, Georgios, et al.
Published: (2024) -
MOD-X: A Modular Open Decentralized eXchange Framework proposal for Heterogeneous Interoperable Artificial Intelligence Agents
by: Ioannides, Georgios, et al.
Published: (2025) -
Enhancing Adverse Drug Event Detection with Multimodal Dataset: Corpus Creation and Model Development
by: Sahoo, Pranab, et al.
Published: (2024) -
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
by: Sinha, Neelabh, et al.
Published: (2024) -
Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions
by: Ghosh, Akash, et al.
Published: (2024)