VISTA: A Visual and Textual Attention Dataset for Interpreting Multimodal Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Harshit, Tasdizen, Tolga |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DuoFormer: Leveraging Hierarchical Visual Representations by Local and Global Attention
di: Tang, Xiaoya, et al.
Pubblicazione: (2024)
di: Tang, Xiaoya, et al.
Pubblicazione: (2024)
A Comparison of Object Detection and Phrase Grounding Models in Chest X-ray Abnormality Localization using Eye-tracking Data
di: Ghelichkhan, Elham, et al.
Pubblicazione: (2025)
di: Ghelichkhan, Elham, et al.
Pubblicazione: (2025)
HAVT-IVD: Heterogeneity-Aware Cross-Modal Network for Audio-Visual Surveillance: Idling Vehicles Detection With Multichannel Audio and Multiscale Visual Cues
di: Li, Xiwen, et al.
Pubblicazione: (2025)
di: Li, Xiwen, et al.
Pubblicazione: (2025)
DuoFormer: Leveraging Hierarchical Representations by Local and Global Attention Vision Transformer
di: Tang, Xiaoya, et al.
Pubblicazione: (2025)
di: Tang, Xiaoya, et al.
Pubblicazione: (2025)
SRA: A Novel Method to Improve Feature Embedding in Self-supervised Learning for Histopathological Images
di: Manoochehri, Hamid, et al.
Pubblicazione: (2024)
di: Manoochehri, Hamid, et al.
Pubblicazione: (2024)
WeakSupCon: Weakly Supervised Contrastive Learning for Encoder Pre-training
di: Zhang, Bodong, et al.
Pubblicazione: (2025)
di: Zhang, Bodong, et al.
Pubblicazione: (2025)
How to Build Robust, Scalable Models for GSV-Based Indicators in Neighborhood Research
di: Tang, Xiaoya, et al.
Pubblicazione: (2026)
di: Tang, Xiaoya, et al.
Pubblicazione: (2026)
Joint Audio-Visual Idling Vehicle Detection with Streamlined Input Dependencies
di: Li, Xiwen, et al.
Pubblicazione: (2024)
di: Li, Xiwen, et al.
Pubblicazione: (2024)
Weakly Supervised Contrastive Learning for Histopathology Patch Embeddings
di: Zhang, Bodong, et al.
Pubblicazione: (2026)
di: Zhang, Bodong, et al.
Pubblicazione: (2026)
BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment
di: Shinoda, Risa, et al.
Pubblicazione: (2026)
di: Shinoda, Risa, et al.
Pubblicazione: (2026)
VISTANet: VIsual Spoken Textual Additive Net for Interpretable Multimodal Emotion Recognition
di: Kumar, Puneet, et al.
Pubblicazione: (2022)
di: Kumar, Puneet, et al.
Pubblicazione: (2022)
F2FLDM: Latent Diffusion Models with Histopathology Pre-Trained Embeddings for Unpaired Frozen Section to FFPE Translation
di: Ho, Man M., et al.
Pubblicazione: (2024)
di: Ho, Man M., et al.
Pubblicazione: (2024)
Autonomous Imagination: Closed-Loop Decomposition of Visual-to-Textual Conversion in Visual Reasoning for Multimodal Large Language Models
di: Liu, Jingming, et al.
Pubblicazione: (2024)
di: Liu, Jingming, et al.
Pubblicazione: (2024)
VisText-Mosquito: A Unified Multimodal Dataset for Visual Detection, Segmentation, and Textual Explanation on Mosquito Breeding Sites
di: Islam, Md. Adnanul, et al.
Pubblicazione: (2025)
di: Islam, Md. Adnanul, et al.
Pubblicazione: (2025)
CLASS-M: Adaptive stain separation-based contrastive learning with pseudo-labeling for histopathological image classification
di: Zhang, Bodong, et al.
Pubblicazione: (2023)
di: Zhang, Bodong, et al.
Pubblicazione: (2023)
Cognitive-Inspired Hierarchical Attention Fusion With Visual and Textual for Cross-Domain Sequential Recommendation
di: Wu, Wangyu, et al.
Pubblicazione: (2025)
di: Wu, Wangyu, et al.
Pubblicazione: (2025)
Optimizing Multimodal Language Models through Attention-based Interpretability
di: Sergeev, Alexander, et al.
Pubblicazione: (2025)
di: Sergeev, Alexander, et al.
Pubblicazione: (2025)
VQA-MHUG: A Gaze Dataset to Study Multimodal Neural Attention in Visual Question Answering
di: Sood, Ekta, et al.
Pubblicazione: (2021)
di: Sood, Ekta, et al.
Pubblicazione: (2021)
VISTA: A Visual Analytics Framework to Enhance Foundation Model-Generated Data Labels
di: Xuan, Xiwei, et al.
Pubblicazione: (2025)
di: Xuan, Xiwei, et al.
Pubblicazione: (2025)
Explaining How Visual, Textual and Multimodal Encoders Share Concepts
di: Cornet, Clément, et al.
Pubblicazione: (2025)
di: Cornet, Clément, et al.
Pubblicazione: (2025)
Disentangled and Interpretable Multimodal Attention Fusion for Cancer Survival Prediction
di: Eijpe, Aniek, et al.
Pubblicazione: (2025)
di: Eijpe, Aniek, et al.
Pubblicazione: (2025)
A Flag Decomposition for Hierarchical Datasets
di: Mankovich, Nathan, et al.
Pubblicazione: (2025)
di: Mankovich, Nathan, et al.
Pubblicazione: (2025)
Textualized and Feature-based Models for Compound Multimodal Emotion Recognition in the Wild
di: Richet, Nicolas, et al.
Pubblicazione: (2024)
di: Richet, Nicolas, et al.
Pubblicazione: (2024)
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
di: Ataallah, Kirolos, et al.
Pubblicazione: (2024)
di: Ataallah, Kirolos, et al.
Pubblicazione: (2024)
Holistic Visual-Textual Sentiment Analysis with Prior Models
di: Chen, Junyu, et al.
Pubblicazione: (2022)
di: Chen, Junyu, et al.
Pubblicazione: (2022)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
di: Gan, Woody Haosheng, et al.
Pubblicazione: (2025)
di: Gan, Woody Haosheng, et al.
Pubblicazione: (2025)
EDITS: Enhancing Dataset Distillation with Implicit Textual Semantics
di: Xia, Qianxin, et al.
Pubblicazione: (2025)
di: Xia, Qianxin, et al.
Pubblicazione: (2025)
Visual Textualization for Image Prompted Object Detection
di: Wu, Yongjian, et al.
Pubblicazione: (2025)
di: Wu, Yongjian, et al.
Pubblicazione: (2025)
PathMR: Multimodal Visual Reasoning for Interpretable Pathology Diagnosis
di: Zhang, Ye, et al.
Pubblicazione: (2025)
di: Zhang, Ye, et al.
Pubblicazione: (2025)
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis
di: Wang, Pengfei, et al.
Pubblicazione: (2025)
di: Wang, Pengfei, et al.
Pubblicazione: (2025)
DISC: Latent Diffusion Models with Self-Distillation from Separated Conditions for Prostate Cancer Grading
di: Ho, Man M., et al.
Pubblicazione: (2024)
di: Ho, Man M., et al.
Pubblicazione: (2024)
TCP:Textual-based Class-aware Prompt tuning for Visual-Language Model
di: Yao, Hantao, et al.
Pubblicazione: (2023)
di: Yao, Hantao, et al.
Pubblicazione: (2023)
VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?
di: Liu, Qing'an, et al.
Pubblicazione: (2026)
di: Liu, Qing'an, et al.
Pubblicazione: (2026)
SignAvatars: A Large-scale 3D Sign Language Holistic Motion Dataset and Benchmark
di: Yu, Zhengdi, et al.
Pubblicazione: (2023)
di: Yu, Zhengdi, et al.
Pubblicazione: (2023)
Visual and Textual Prompts in VLLMs for Enhancing Emotion Recognition
di: Wang, Zhifeng, et al.
Pubblicazione: (2025)
di: Wang, Zhifeng, et al.
Pubblicazione: (2025)
GIFT: A Framework Towards Global Interpretable Faithful Textual Explanations of Vision Classifiers
di: Zablocki, Éloi, et al.
Pubblicazione: (2024)
di: Zablocki, Éloi, et al.
Pubblicazione: (2024)
MMIS: Multimodal Dataset for Interior Scene Visual Generation and Recognition
di: Kassab, Hozaifa, et al.
Pubblicazione: (2024)
di: Kassab, Hozaifa, et al.
Pubblicazione: (2024)
Visual Attention Drifts,but Anchors Hold:Mitigating Hallucination in Multimodal Large Language Models via Cross-Layer Visual Anchors
di: Yang, Chengxu, et al.
Pubblicazione: (2026)
di: Yang, Chengxu, et al.
Pubblicazione: (2026)
CLEAR: Character Unlearning in Textual and Visual Modalities
di: Dontsov, Alexey, et al.
Pubblicazione: (2024)
di: Dontsov, Alexey, et al.
Pubblicazione: (2024)
Probing CLIP's Comprehension of 360-Degree Textual and Visual Semantics
di: Wang, Hai, et al.
Pubblicazione: (2026)
di: Wang, Hai, et al.
Pubblicazione: (2026)
Documenti analoghi
-
DuoFormer: Leveraging Hierarchical Visual Representations by Local and Global Attention
di: Tang, Xiaoya, et al.
Pubblicazione: (2024) -
A Comparison of Object Detection and Phrase Grounding Models in Chest X-ray Abnormality Localization using Eye-tracking Data
di: Ghelichkhan, Elham, et al.
Pubblicazione: (2025) -
HAVT-IVD: Heterogeneity-Aware Cross-Modal Network for Audio-Visual Surveillance: Idling Vehicles Detection With Multichannel Audio and Multiscale Visual Cues
di: Li, Xiwen, et al.
Pubblicazione: (2025) -
DuoFormer: Leveraging Hierarchical Representations by Local and Global Attention Vision Transformer
di: Tang, Xiaoya, et al.
Pubblicazione: (2025) -
SRA: A Novel Method to Improve Feature Embedding in Self-supervised Learning for Histopathological Images
di: Manoochehri, Hamid, et al.
Pubblicazione: (2024)