Attention Lattice Adapter: Visual Explanation Generation for Visual Foundation Model
Fuente:
arXiv
Guardado en:
| Autores principales: | Hirano, Shinnosuke, Wada, Yuiga, Iida, Tsumugi, Sugiura, Komei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LLM-Free Image Captioning Evaluation in Reference-Flexible Settings
por: Hirano, Shinnosuke, et al.
Publicado: (2025)
por: Hirano, Shinnosuke, et al.
Publicado: (2025)
VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions
por: Matsuda, Kazuki, et al.
Publicado: (2025)
por: Matsuda, Kazuki, et al.
Publicado: (2025)
DENEB: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning
por: Matsuda, Kazuki, et al.
Publicado: (2024)
por: Matsuda, Kazuki, et al.
Publicado: (2024)
MLLM-as-a-Judge Exhibits Model Preference Bias
por: Koyama, Shuitsu, et al.
Publicado: (2026)
por: Koyama, Shuitsu, et al.
Publicado: (2026)
ZINA: Multimodal Fine-grained Hallucination Detection and Editing
por: Wada, Yuiga, et al.
Publicado: (2025)
por: Wada, Yuiga, et al.
Publicado: (2025)
Polos: Multimodal Metric Learning from Human Feedback for Image Captioning
por: Wada, Yuiga, et al.
Publicado: (2024)
por: Wada, Yuiga, et al.
Publicado: (2024)
Layer-Wise Relevance Propagation with Conservation Property for ResNet
por: Otsuki, Seitaro, et al.
Publicado: (2024)
por: Otsuki, Seitaro, et al.
Publicado: (2024)
Deep Space Weather Model: Long-Range Solar Flare Prediction from Multi-Wavelength Images
por: Nagashima, Shunya, et al.
Publicado: (2025)
por: Nagashima, Shunya, et al.
Publicado: (2025)
Future Success Prediction in Open-Vocabulary Object Manipulation Tasks Based on End-Effector Trajectories
por: Kambara, Motonari, et al.
Publicado: (2024)
por: Kambara, Motonari, et al.
Publicado: (2024)
Object Segmentation from Open-Vocabulary Manipulation Instructions Based on Optimal Transport Polygon Matching with Multimodal Foundation Models
por: Nishimura, Takayuki, et al.
Publicado: (2024)
por: Nishimura, Takayuki, et al.
Publicado: (2024)
Capturing Fine-Grained Alignments Improves 3D Affordance Detection
por: Tokumitsu, Junsei, et al.
Publicado: (2025)
por: Tokumitsu, Junsei, et al.
Publicado: (2025)
ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video Understanding
por: Yashima, Daichi, et al.
Publicado: (2026)
por: Yashima, Daichi, et al.
Publicado: (2026)
FLARE-SSM: Deep State Space Models with Influence-Balanced Loss for 72-Hour Solar Flare Prediction
por: Takagi, Yusuke, et al.
Publicado: (2025)
por: Takagi, Yusuke, et al.
Publicado: (2025)
Open-Vocabulary Mobile Manipulation Based on Double Relaxed Contrastive Learning with Dense Labeling
por: Yashima, Daichi, et al.
Publicado: (2024)
por: Yashima, Daichi, et al.
Publicado: (2024)
Stitch4D: Sparse Multi-Location 4D Urban Reconstruction via Spatio-Temporal Interpolation
por: Kogure, Hina, et al.
Publicado: (2026)
por: Kogure, Hina, et al.
Publicado: (2026)
EVCtrl: Efficient Control Adapter for Visual Generation
por: Yang, Zixiang, et al.
Publicado: (2025)
por: Yang, Zixiang, et al.
Publicado: (2025)
Cortical-SSM: A Deep State Space Model for EEG and ECoG Motor Imagery Decoding
por: Suzuki, Shuntaro, et al.
Publicado: (2025)
por: Suzuki, Shuntaro, et al.
Publicado: (2025)
GENNAV: Polygon Mask Generation for Generalized Referring Navigable Regions
por: Katsumata, Kei, et al.
Publicado: (2025)
por: Katsumata, Kei, et al.
Publicado: (2025)
FontAdapter: Instant Font Adaptation in Visual Text Generation
por: Koo, Myungkyu, et al.
Publicado: (2025)
por: Koo, Myungkyu, et al.
Publicado: (2025)
Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
por: Korekata, Ryosuke, et al.
Publicado: (2025)
por: Korekata, Ryosuke, et al.
Publicado: (2025)
ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning
por: Yashima, Daichi, et al.
Publicado: (2026)
por: Yashima, Daichi, et al.
Publicado: (2026)
Task Success Prediction for Open-Vocabulary Manipulation Based on Multi-Level Aligned Representations
por: Goko, Miyu, et al.
Publicado: (2024)
por: Goko, Miyu, et al.
Publicado: (2024)
Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video Captioning
por: Chen, Junan, et al.
Publicado: (2025)
por: Chen, Junan, et al.
Publicado: (2025)
From Visual Explanations to Counterfactual Explanations with Latent Diffusion
por: Luu, Tung, et al.
Publicado: (2025)
por: Luu, Tung, et al.
Publicado: (2025)
Attention Guided CAM: Visual Explanations of Vision Transformer Guided by Self-Attention
por: Leem, Saebom, et al.
Publicado: (2024)
por: Leem, Saebom, et al.
Publicado: (2024)
Robust Scene Change Detection Using Visual Foundation Models and Cross-Attention Mechanisms
por: Lin, Chun-Jung, et al.
Publicado: (2024)
por: Lin, Chun-Jung, et al.
Publicado: (2024)
ChartQA-X: Generating Explanations for Visual Chart Reasoning
por: Hegde, Shamanthak, et al.
Publicado: (2025)
por: Hegde, Shamanthak, et al.
Publicado: (2025)
Integrating Reinforcement Learning with Visual Generative Models: Foundations and Advances
por: Liang, Yuanzhi, et al.
Publicado: (2025)
por: Liang, Yuanzhi, et al.
Publicado: (2025)
RAW-Adapter: Adapting Pre-trained Visual Model to Camera RAW Images
por: Cui, Ziteng, et al.
Publicado: (2024)
por: Cui, Ziteng, et al.
Publicado: (2024)
DM2RM: Dual-Mode Multimodal Ranking for Target Objects and Receptacles Based on Open-Vocabulary Instructions
por: Korekata, Ryosuke, et al.
Publicado: (2024)
por: Korekata, Ryosuke, et al.
Publicado: (2024)
Faithful Counterfactual Visual Explanations (FCVE)
por: Khan, Bismillah, et al.
Publicado: (2025)
por: Khan, Bismillah, et al.
Publicado: (2025)
RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers
por: Gong, Yan, et al.
Publicado: (2025)
por: Gong, Yan, et al.
Publicado: (2025)
MoSA: Mixture of Sparse Adapters for Visual Efficient Tuning
por: Zhang, Qizhe, et al.
Publicado: (2023)
por: Zhang, Qizhe, et al.
Publicado: (2023)
Dyn-Adapter: Towards Disentangled Representation for Efficient Visual Recognition
por: Zhang, Yurong, et al.
Publicado: (2024)
por: Zhang, Yurong, et al.
Publicado: (2024)
Evaluating Visual Explanations of Attention Maps for Transformer-based Medical Imaging
por: Chung, Minjae, et al.
Publicado: (2025)
por: Chung, Minjae, et al.
Publicado: (2025)
Audio-Visual Intelligence in Large Foundation Models
por: Qin, You, et al.
Publicado: (2026)
por: Qin, You, et al.
Publicado: (2026)
Test-time Distribution Learning Adapter for Cross-modal Visual Reasoning
por: Zhang, Yi, et al.
Publicado: (2024)
por: Zhang, Yi, et al.
Publicado: (2024)
Pear: Pruning and Sharing Adapters in Visual Parameter-Efficient Fine-Tuning
por: Zhong, Yibo, et al.
Publicado: (2024)
por: Zhong, Yibo, et al.
Publicado: (2024)
NaiLIA: Multimodal Nail Design Retrieval Based on Dense Intent Descriptions and Palette Queries
por: Amemiya, Kanon, et al.
Publicado: (2026)
por: Amemiya, Kanon, et al.
Publicado: (2026)
RAW-Adapter: Adapting Pre-trained Visual Model to Camera RAW Images and A Benchmark
por: Cui, Ziteng, et al.
Publicado: (2025)
por: Cui, Ziteng, et al.
Publicado: (2025)
Ejemplares similares
-
LLM-Free Image Captioning Evaluation in Reference-Flexible Settings
por: Hirano, Shinnosuke, et al.
Publicado: (2025) -
VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions
por: Matsuda, Kazuki, et al.
Publicado: (2025) -
DENEB: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning
por: Matsuda, Kazuki, et al.
Publicado: (2024) -
MLLM-as-a-Judge Exhibits Model Preference Bias
por: Koyama, Shuitsu, et al.
Publicado: (2026) -
ZINA: Multimodal Fine-grained Hallucination Detection and Editing
por: Wada, Yuiga, et al.
Publicado: (2025)