Towards Explainable AI: Multi-Modal Transformer for Video-based Image Description Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Agarwal, Lakshita, Verma, Bindu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism
von: Agarwal, Lakshita, et al.
Veröffentlicht: (2025)
von: Agarwal, Lakshita, et al.
Veröffentlicht: (2025)
Advanced Chest X-Ray Analysis via Transformer-Based Image Descriptors and Cross-Model Attention Mechanism
von: Agarwal, Lakshita, et al.
Veröffentlicht: (2025)
von: Agarwal, Lakshita, et al.
Veröffentlicht: (2025)
DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning
von: Gao, Yifeng, et al.
Veröffentlicht: (2025)
von: Gao, Yifeng, et al.
Veröffentlicht: (2025)
Towards Zero-Shot & Explainable Video Description by Reasoning over Graphs of Events in Space and Time
von: Masala, Mihai, et al.
Veröffentlicht: (2025)
von: Masala, Mihai, et al.
Veröffentlicht: (2025)
MMLNB: Multi-Modal Learning for Neuroblastoma Subtyping Classification Assisted with Textual Description Generation
von: Chen, Huangwei, et al.
Veröffentlicht: (2025)
von: Chen, Huangwei, et al.
Veröffentlicht: (2025)
Brain Stroke Detection and Classification Using CT Imaging with Transformer Models and Explainable AI
von: Qari, Shomukh, et al.
Veröffentlicht: (2025)
von: Qari, Shomukh, et al.
Veröffentlicht: (2025)
XMACNet: An Explainable Lightweight Attention based CNN with Multi Modal Fusion for Chili Disease Classification
von: Ray, Tapon Kumer, et al.
Veröffentlicht: (2026)
von: Ray, Tapon Kumer, et al.
Veröffentlicht: (2026)
Multi-Head Explainer: A General Framework to Improve Explainability in CNNs and Transformers
von: Sun, Bohang, et al.
Veröffentlicht: (2025)
von: Sun, Bohang, et al.
Veröffentlicht: (2025)
Federated Transformer-GNN for Privacy-Preserving Brain Tumor Localization with Modality-Level Explainability
von: Protani, Andrea, et al.
Veröffentlicht: (2026)
von: Protani, Andrea, et al.
Veröffentlicht: (2026)
Towards Multi-Task Multi-Modal Models: A Video Generative Perspective
von: Yu, Lijun
Veröffentlicht: (2024)
von: Yu, Lijun
Veröffentlicht: (2024)
MICA: Towards Explainable Skin Lesion Diagnosis via Multi-Level Image-Concept Alignment
von: Bie, Yequan, et al.
Veröffentlicht: (2024)
von: Bie, Yequan, et al.
Veröffentlicht: (2024)
REVEAL: Reasoning-Enhanced Forensic Evidence Analysis for Explainable AI-Generated Image Detection
von: Cao, Huangsen, et al.
Veröffentlicht: (2025)
von: Cao, Huangsen, et al.
Veröffentlicht: (2025)
DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation
von: Cai, Minghong, et al.
Veröffentlicht: (2024)
von: Cai, Minghong, et al.
Veröffentlicht: (2024)
Cross-Modal Transferable Image-to-Video Attack on Video Quality Metrics
von: Gotin, Georgii, et al.
Veröffentlicht: (2025)
von: Gotin, Georgii, et al.
Veröffentlicht: (2025)
Towards Quantitative Evaluation of Explainable AI Methods for Deepfake Detection
von: Tsigos, Konstantinos, et al.
Veröffentlicht: (2024)
von: Tsigos, Konstantinos, et al.
Veröffentlicht: (2024)
PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement
von: Hu, Teng, et al.
Veröffentlicht: (2025)
von: Hu, Teng, et al.
Veröffentlicht: (2025)
A Framework for Evaluating Zero-Shot Image Generation in Concept-based Explainability
von: Astolfi, Giacomo, et al.
Veröffentlicht: (2026)
von: Astolfi, Giacomo, et al.
Veröffentlicht: (2026)
MIRAGE: Towards AI-Generated Image Detection in the Wild
von: Xia, Cheng, et al.
Veröffentlicht: (2025)
von: Xia, Cheng, et al.
Veröffentlicht: (2025)
Explainable, Multi-modal Wound Infection Classification from Images Augmented with Generated Captions
von: Busaranuvong, Palawat, et al.
Veröffentlicht: (2025)
von: Busaranuvong, Palawat, et al.
Veröffentlicht: (2025)
AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI
von: Fan, Fanda, et al.
Veröffentlicht: (2024)
von: Fan, Fanda, et al.
Veröffentlicht: (2024)
Towards Counterfactual and Contrastive Explainability and Transparency of DCNN Image Classifiers
von: Tariq, Syed Ali, et al.
Veröffentlicht: (2025)
von: Tariq, Syed Ali, et al.
Veröffentlicht: (2025)
PASSION: Towards Effective Incomplete Multi-Modal Medical Image Segmentation with Imbalanced Missing Rates
von: Shi, Junjie, et al.
Veröffentlicht: (2024)
von: Shi, Junjie, et al.
Veröffentlicht: (2024)
Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning
von: Zhang, Fanrui, et al.
Veröffentlicht: (2025)
von: Zhang, Fanrui, et al.
Veröffentlicht: (2025)
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
von: Rawal, Ishaan Singh, et al.
Veröffentlicht: (2023)
von: Rawal, Ishaan Singh, et al.
Veröffentlicht: (2023)
ProtoMedX: Towards Explainable Multi-Modal Prototype Learning for Bone Health Classification
von: Pellicer, Alvaro Lopez, et al.
Veröffentlicht: (2025)
von: Pellicer, Alvaro Lopez, et al.
Veröffentlicht: (2025)
Towards Multimodal Video Paragraph Captioning Models Robust to Missing Modality
von: Chen, Sishuo, et al.
Veröffentlicht: (2024)
von: Chen, Sishuo, et al.
Veröffentlicht: (2024)
Enhancing Osteoporosis Detection: An Explainable Multi-Modal Learning Framework with Feature Fusion and Variable Clustering
von: Chagahi, Mehdi Hosseini, et al.
Veröffentlicht: (2024)
von: Chagahi, Mehdi Hosseini, et al.
Veröffentlicht: (2024)
DIVE: Towards Descriptive and Diverse Visual Commonsense Generation
von: Park, Jun-Hyung, et al.
Veröffentlicht: (2024)
von: Park, Jun-Hyung, et al.
Veröffentlicht: (2024)
CamViG: Camera Aware Image-to-Video Generation with Multimodal Transformers
von: Marmon, Andrew, et al.
Veröffentlicht: (2024)
von: Marmon, Andrew, et al.
Veröffentlicht: (2024)
MUSES: 3D-Controllable Image Generation via Multi-Modal Agent Collaboration
von: Ding, Yanbo, et al.
Veröffentlicht: (2024)
von: Ding, Yanbo, et al.
Veröffentlicht: (2024)
Early Exit and Multi Stage Knowledge Distillation in VLMs for Video Summarization
von: Khan, Anas Anwarul Haq, et al.
Veröffentlicht: (2025)
von: Khan, Anas Anwarul Haq, et al.
Veröffentlicht: (2025)
Self-Corrected Image Generation with Explainable Latent Rewards
von: Luo, Yinyi, et al.
Veröffentlicht: (2026)
von: Luo, Yinyi, et al.
Veröffentlicht: (2026)
Red Teaming Models for Hyperspectral Image Analysis Using Explainable AI
von: Zaigrajew, Vladimir, et al.
Veröffentlicht: (2024)
von: Zaigrajew, Vladimir, et al.
Veröffentlicht: (2024)
Robust Multiple Description Neural Video Codec with Masked Transformer for Dynamic and Noisy Networks
von: Hu, Xinyue, et al.
Veröffentlicht: (2024)
von: Hu, Xinyue, et al.
Veröffentlicht: (2024)
Sensor-Adaptive Flood Mapping with Pre-trained Multi-Modal Transformers across SAR and Multispectral Modalities
von: Tanaka, Tomohiro, et al.
Veröffentlicht: (2025)
von: Tanaka, Tomohiro, et al.
Veröffentlicht: (2025)
Ivy-Fake: A Unified Explainable Framework and Benchmark for Image and Video AIGC Detection
von: Jiang, Changjiang, et al.
Veröffentlicht: (2025)
von: Jiang, Changjiang, et al.
Veröffentlicht: (2025)
Multi-language Video Subtitle Dataset for Image-based Text Recognition
von: Singkhornart, Thanadol, et al.
Veröffentlicht: (2024)
von: Singkhornart, Thanadol, et al.
Veröffentlicht: (2024)
Contrastive Learning-based Multi Modal Architecture for Emoticon Prediction by Employing Image-Text Pairs
von: Pandey, Ananya, et al.
Veröffentlicht: (2024)
von: Pandey, Ananya, et al.
Veröffentlicht: (2024)
Edge-Enhanced Vision Transformer Framework for Accurate AI-Generated Image Detection
von: Das, Dabbrata, et al.
Veröffentlicht: (2025)
von: Das, Dabbrata, et al.
Veröffentlicht: (2025)
On the Effectiveness of Methods and Metrics for Explainable AI in Remote Sensing Image Scene Classification
von: Klotz, Jonas, et al.
Veröffentlicht: (2025)
von: Klotz, Jonas, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism
von: Agarwal, Lakshita, et al.
Veröffentlicht: (2025) -
Advanced Chest X-Ray Analysis via Transformer-Based Image Descriptors and Cross-Model Attention Mechanism
von: Agarwal, Lakshita, et al.
Veröffentlicht: (2025) -
DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning
von: Gao, Yifeng, et al.
Veröffentlicht: (2025) -
Towards Zero-Shot & Explainable Video Description by Reasoning over Graphs of Events in Space and Time
von: Masala, Mihai, et al.
Veröffentlicht: (2025) -
MMLNB: Multi-Modal Learning for Neuroblastoma Subtyping Classification Assisted with Textual Description Generation
von: Chen, Huangwei, et al.
Veröffentlicht: (2025)