MAMI: Multi-Attentional Mutual-Information for Long Sequence Neuron Captioning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fauzulhaq, Alfirsa Damasyifa, Parwitayasa, Wahyu, Sugihdharma, Joseph Ananda, Ridhani, M. Fadli, Yudistira, Novanto |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Adaptive Fusion of Multimodal Deep Networks for Human Action Recognition
von: Yudistira, Novanto
Veröffentlicht: (2025)
von: Yudistira, Novanto
Veröffentlicht: (2025)
Efficient Object Detection of Marine Debris using Pruned YOLO Model
von: Aryaza, Abi, et al.
Veröffentlicht: (2025)
von: Aryaza, Abi, et al.
Veröffentlicht: (2025)
Adaptive Conformal Prediction for Reliable and Explainable Medical Image Classification
von: Octadion, One, et al.
Veröffentlicht: (2026)
von: Octadion, One, et al.
Veröffentlicht: (2026)
Input-Adaptive Visual Preprocessing for Efficient Fast Vision-Language Model Inference
von: Cahyani, Putu Indah Githa, et al.
Veröffentlicht: (2025)
von: Cahyani, Putu Indah Githa, et al.
Veröffentlicht: (2025)
IndoHerb: Indonesia Medicinal Plants Recognition using Transfer Learning and Deep Learning
von: Musyaffa, Muhammad Salman Ikrar, et al.
Veröffentlicht: (2023)
von: Musyaffa, Muhammad Salman Ikrar, et al.
Veröffentlicht: (2023)
Hybrid of DiffStride and Spectral Pooling in Convolutional Neural Networks
von: Rafif, Sulthan, et al.
Veröffentlicht: (2024)
von: Rafif, Sulthan, et al.
Veröffentlicht: (2024)
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
von: Mohamed, Abdelrahman, et al.
Veröffentlicht: (2025)
von: Mohamed, Abdelrahman, et al.
Veröffentlicht: (2025)
Training-Free Disentangled Text-Guided Image Editing via Sparse Latent Constraints
von: Shabrina, Mutiara, et al.
Veröffentlicht: (2025)
von: Shabrina, Mutiara, et al.
Veröffentlicht: (2025)
Regional Attention-Enhanced Swin Transformer for Clinically Relevant Medical Image Captioning
von: Naz, Zubia, et al.
Veröffentlicht: (2025)
von: Naz, Zubia, et al.
Veröffentlicht: (2025)
Transformer with Controlled Attention for Synchronous Motion Captioning
von: Radouane, Karim, et al.
Veröffentlicht: (2024)
von: Radouane, Karim, et al.
Veröffentlicht: (2024)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
von: Chen, Xiaofu, et al.
Veröffentlicht: (2025)
von: Chen, Xiaofu, et al.
Veröffentlicht: (2025)
Multi-LLM Collaborative Caption Generation in Scientific Documents
von: Kim, Jaeyoung, et al.
Veröffentlicht: (2025)
von: Kim, Jaeyoung, et al.
Veröffentlicht: (2025)
OmniCaptioner: One Captioner to Rule Them All
von: Lu, Yiting, et al.
Veröffentlicht: (2025)
von: Lu, Yiting, et al.
Veröffentlicht: (2025)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
von: Nguyen, Nhi Ngoc-Yen, et al.
Veröffentlicht: (2026)
von: Nguyen, Nhi Ngoc-Yen, et al.
Veröffentlicht: (2026)
ASVRI-Legal: Fine-Tuning LLMs with Retrieval Augmented Generation for Enhanced Legal Regulation
von: Octadion, One, et al.
Veröffentlicht: (2025)
von: Octadion, One, et al.
Veröffentlicht: (2025)
MIRL: Mutual Information-Guided Reinforcement Learning for Vision-Language Models
von: Zhang, Yin, et al.
Veröffentlicht: (2026)
von: Zhang, Yin, et al.
Veröffentlicht: (2026)
Inserting Faces inside Captions: Image Captioning with Attention Guided Merging
von: Tevissen, Yannis, et al.
Veröffentlicht: (2024)
von: Tevissen, Yannis, et al.
Veröffentlicht: (2024)
APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention
von: Huang, Yuxiang, et al.
Veröffentlicht: (2026)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2026)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
A Video is Worth 10,000 Words: Training and Benchmarking with Diverse Captions for Better Long Video Retrieval
von: Gwilliam, Matthew, et al.
Veröffentlicht: (2023)
von: Gwilliam, Matthew, et al.
Veröffentlicht: (2023)
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
von: Wang, Xinran, et al.
Veröffentlicht: (2026)
von: Wang, Xinran, et al.
Veröffentlicht: (2026)
From Image Captioning to Visual Storytelling
von: Passadakis, Admitos, et al.
Veröffentlicht: (2025)
von: Passadakis, Admitos, et al.
Veröffentlicht: (2025)
VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions
von: Matsuda, Kazuki, et al.
Veröffentlicht: (2025)
von: Matsuda, Kazuki, et al.
Veröffentlicht: (2025)
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
von: Gan, Chengguang, et al.
Veröffentlicht: (2025)
von: Gan, Chengguang, et al.
Veröffentlicht: (2025)
Image Captioning via Compact Bidirectional Architecture
von: Song, Zijie, et al.
Veröffentlicht: (2022)
von: Song, Zijie, et al.
Veröffentlicht: (2022)
MUNIChus: Multilingual News Image Captioning Benchmark
von: Chen, Yuji, et al.
Veröffentlicht: (2026)
von: Chen, Yuji, et al.
Veröffentlicht: (2026)
Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs
von: Fang, Hao, et al.
Veröffentlicht: (2025)
von: Fang, Hao, et al.
Veröffentlicht: (2025)
From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models
von: Wang, Qidong, et al.
Veröffentlicht: (2026)
von: Wang, Qidong, et al.
Veröffentlicht: (2026)
TriAttention: Efficient Long Reasoning with Trigonometric KV Compression
von: Mao, Weian, et al.
Veröffentlicht: (2026)
von: Mao, Weian, et al.
Veröffentlicht: (2026)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
von: Li, Wenyan, et al.
Veröffentlicht: (2024)
von: Li, Wenyan, et al.
Veröffentlicht: (2024)
Temporal Image Caption Retrieval Competition -- Description and Results
von: Pokrywka, Jakub, et al.
Veröffentlicht: (2024)
von: Pokrywka, Jakub, et al.
Veröffentlicht: (2024)
TADACap: Time-series Adaptive Domain-Aware Captioning
von: Fons, Elizabeth, et al.
Veröffentlicht: (2025)
von: Fons, Elizabeth, et al.
Veröffentlicht: (2025)
Redemption Score: A Multi-Modal Evaluation Framework for Image Captioning via Distributional, Perceptual, and Linguistic Signal Triangulation
von: Dahal, Ashim, et al.
Veröffentlicht: (2025)
von: Dahal, Ashim, et al.
Veröffentlicht: (2025)
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
von: Bai, Longju, et al.
Veröffentlicht: (2024)
von: Bai, Longju, et al.
Veröffentlicht: (2024)
VIXEN: Visual Text Comparison Network for Image Difference Captioning
von: Black, Alexander, et al.
Veröffentlicht: (2024)
von: Black, Alexander, et al.
Veröffentlicht: (2024)
Transformer based Multitask Learning for Image Captioning and Object Detection
von: Basak, Debolena, et al.
Veröffentlicht: (2024)
von: Basak, Debolena, et al.
Veröffentlicht: (2024)
Altogether: Image Captioning via Re-aligning Alt-text
von: Xu, Hu, et al.
Veröffentlicht: (2024)
von: Xu, Hu, et al.
Veröffentlicht: (2024)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025)
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025)
Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time
von: Berger, Uri, et al.
Veröffentlicht: (2025)
von: Berger, Uri, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards Adaptive Fusion of Multimodal Deep Networks for Human Action Recognition
von: Yudistira, Novanto
Veröffentlicht: (2025) -
Efficient Object Detection of Marine Debris using Pruned YOLO Model
von: Aryaza, Abi, et al.
Veröffentlicht: (2025) -
Adaptive Conformal Prediction for Reliable and Explainable Medical Image Classification
von: Octadion, One, et al.
Veröffentlicht: (2026) -
Input-Adaptive Visual Preprocessing for Efficient Fast Vision-Language Model Inference
von: Cahyani, Putu Indah Githa, et al.
Veröffentlicht: (2025) -
IndoHerb: Indonesia Medicinal Plants Recognition using Transfer Learning and Deep Learning
von: Musyaffa, Muhammad Salman Ikrar, et al.
Veröffentlicht: (2023)