Inserting Faces inside Captions: Image Captioning with Attention Guided Merging
Fuente:
arXiv
Saved in:
| Main Authors: | Tevissen, Yannis, Guetari, Khalil, Tassel, Marine, Kerleroux, Erwan, Petitpont, Frédéric |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multimodal Chaptering for Long-Form TV Newscast Video
by: Guetari, Khalil, et al.
Published: (2024)
by: Guetari, Khalil, et al.
Published: (2024)
Towards Retrieval Augmented Generation over Large Video Libraries
by: Tevissen, Yannis, et al.
Published: (2024)
by: Tevissen, Yannis, et al.
Published: (2024)
Stacked Cross-modal Feature Consolidation Attention Networks for Image Captioning
by: Pourkeshavarz, Mozhgan, et al.
Published: (2023)
by: Pourkeshavarz, Mozhgan, et al.
Published: (2023)
Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization
by: Tang, Changli, et al.
Published: (2024)
by: Tang, Changli, et al.
Published: (2024)
Whitened CLIP as a Likelihood Surrogate of Images and Captions
by: Betser, Roy, et al.
Published: (2025)
by: Betser, Roy, et al.
Published: (2025)
The Solution for the CVPR2023 NICE Image Captioning Challenge
by: Wu, Xiangyu, et al.
Published: (2023)
by: Wu, Xiangyu, et al.
Published: (2023)
RetinaLogos: Fine-Grained Synthesis of High-Resolution Retinal Images Through Captions
by: Ning, Junzhi, et al.
Published: (2025)
by: Ning, Junzhi, et al.
Published: (2025)
Omnidirectional Image Quality Captioning: A Large-scale Database and A New Model
by: Yan, Jiebin, et al.
Published: (2025)
by: Yan, Jiebin, et al.
Published: (2025)
GCS-M3VLT: Guided Context Self-Attention based Multi-modal Medical Vision Language Transformer for Retinal Image Captioning
by: Cherukuri, Teja Krishna, et al.
Published: (2024)
by: Cherukuri, Teja Krishna, et al.
Published: (2024)
MedBLIP: Fine-tuning BLIP for Medical Image Captioning
by: Limbu, Manshi, et al.
Published: (2025)
by: Limbu, Manshi, et al.
Published: (2025)
SPECTRUM: Semantic Processing and Emotion-informed video-Captioning Through Retrieval and Understanding Modalities
by: Faghihi, Ehsan, et al.
Published: (2024)
by: Faghihi, Ehsan, et al.
Published: (2024)
Transformers in Medicine: Improving Vision-Language Alignment for Medical Image Captioning
by: Suresh, Yogesh Thakku, et al.
Published: (2025)
by: Suresh, Yogesh Thakku, et al.
Published: (2025)
Listening without Looking: Modality Bias in Audio-Visual Captioning
by: Ishikawa, Yuchi, et al.
Published: (2025)
by: Ishikawa, Yuchi, et al.
Published: (2025)
CAMP-VQA: Caption-Embedded Multimodal Perception for No-Reference Quality Assessment of Compressed Video
by: Wang, Xinyi, et al.
Published: (2025)
by: Wang, Xinyi, et al.
Published: (2025)
ATTN-FIQA: Interpretable Attention-based Face Image Quality Assessment with Vision Transformers
by: Ozgur, Guray, et al.
Published: (2026)
by: Ozgur, Guray, et al.
Published: (2026)
The Importance of Facial Features in Vision-based Sign Language Recognition: Eyes, Mouth or Full Face?
by: Pham, Dinh Nam, et al.
Published: (2025)
by: Pham, Dinh Nam, et al.
Published: (2025)
Attention-Aware Laparoscopic Image Desmoking Network with Lightness Embedding and Hybrid Guided Embedding
by: Liu, Ziteng, et al.
Published: (2024)
by: Liu, Ziteng, et al.
Published: (2024)
Explainable Face Verification via Feature-Guided Gradient Backpropagation
by: Lu, Yuhang, et al.
Published: (2024)
by: Lu, Yuhang, et al.
Published: (2024)
MGA-Net: A Novel Mask-Guided Attention Neural Network for Precision Neonatal Brain Imaging
by: Jafrasteh, Bahram, et al.
Published: (2024)
by: Jafrasteh, Bahram, et al.
Published: (2024)
Q-space Guided Collaborative Attention Translation Network for Flexible Diffusion-Weighted Images Synthesis
by: Zhu, Pengli, et al.
Published: (2025)
by: Zhu, Pengli, et al.
Published: (2025)
ShieldGemma 2: Robust and Tractable Image Content Moderation
by: Zeng, Wenjun, et al.
Published: (2025)
by: Zeng, Wenjun, et al.
Published: (2025)
Mind the Context: Attention-Guided Weak-to-Strong Consistency for Enhanced Semi-Supervised Medical Image Segmentation
by: Cheng, Yuxuan, et al.
Published: (2024)
by: Cheng, Yuxuan, et al.
Published: (2024)
A Generative Framework for Bidirectional Image-Report Understanding in Chest Radiography
by: Evans, Nicholas, et al.
Published: (2025)
by: Evans, Nicholas, et al.
Published: (2025)
CSANet: Channel Spatial Attention Network for Robust 3D Face Alignment and Reconstruction
by: Liu, Yilin, et al.
Published: (2024)
by: Liu, Yilin, et al.
Published: (2024)
Towards the Detection of AI-Synthesized Human Face Images
by: Lu, Yuhang, et al.
Published: (2024)
by: Lu, Yuhang, et al.
Published: (2024)
Disability Representations: Finding Biases in Automatic Image Generation
by: Tevissen, Yannis
Published: (2024)
by: Tevissen, Yannis
Published: (2024)
FundaQ-8: A Clinically-Inspired Scoring Framework for Automated Fundus Image Quality Assessment
by: Zun, Lee Qi, et al.
Published: (2025)
by: Zun, Lee Qi, et al.
Published: (2025)
MedVisionLlama: Leveraging Pre-Trained Large Language Model Layers to Enhance Medical Image Segmentation
by: Kumar, Gurucharan Marthi Krishna, et al.
Published: (2024)
by: Kumar, Gurucharan Marthi Krishna, et al.
Published: (2024)
Dilated Strip Attention Network for Image Restoration
by: Hao, Fangwei, et al.
Published: (2024)
by: Hao, Fangwei, et al.
Published: (2024)
Region Attention Transformer for Medical Image Restoration
by: Yang, Zhiwen, et al.
Published: (2024)
by: Yang, Zhiwen, et al.
Published: (2024)
Unsupervised Deformable Image Registration with Local-Global Attention and Image Decomposition
by: Huang, Zhengyong, et al.
Published: (2026)
by: Huang, Zhengyong, et al.
Published: (2026)
Frame Sampling Strategies Matter: A Benchmark for small vision language models
by: Brkic, Marija, et al.
Published: (2025)
by: Brkic, Marija, et al.
Published: (2025)
FANCL: Feature-Guided Attention Network with Curriculum Learning for Brain Metastases Segmentation
by: Liu, Zijiang, et al.
Published: (2024)
by: Liu, Zijiang, et al.
Published: (2024)
Medical Image Segmentation Using Directional Window Attention
by: Kareem, Daniya Najiha Abdul, et al.
Published: (2024)
by: Kareem, Daniya Najiha Abdul, et al.
Published: (2024)
Hybrid Convolutional and Attention Network for Hyperspectral Image Denoising
by: Hu, Shuai, et al.
Published: (2024)
by: Hu, Shuai, et al.
Published: (2024)
EvaNet: Elevation-Guided Flood Extent Mapping on Earth Imagery (Extended Version)
by: Sami, Mirza Tanzim, et al.
Published: (2024)
by: Sami, Mirza Tanzim, et al.
Published: (2024)
ComFace: Facial Representation Learning with Synthetic Data for Comparing Faces
by: Akamatsu, Yusuke, et al.
Published: (2024)
by: Akamatsu, Yusuke, et al.
Published: (2024)
CANAMRF: An Attention-Based Model for Multimodal Depression Detection
by: Wei, Yuntao, et al.
Published: (2024)
by: Wei, Yuntao, et al.
Published: (2024)
Leveraging Depth Maps and Attention Mechanisms for Enhanced Image Inpainting
by: Park, Jin Hyun, et al.
Published: (2025)
by: Park, Jin Hyun, et al.
Published: (2025)
Multi-scale Attention Network for Single Image Super-Resolution
by: Wang, Yan, et al.
Published: (2022)
by: Wang, Yan, et al.
Published: (2022)
Similar Items
-
Multimodal Chaptering for Long-Form TV Newscast Video
by: Guetari, Khalil, et al.
Published: (2024) -
Towards Retrieval Augmented Generation over Large Video Libraries
by: Tevissen, Yannis, et al.
Published: (2024) -
Stacked Cross-modal Feature Consolidation Attention Networks for Image Captioning
by: Pourkeshavarz, Mozhgan, et al.
Published: (2023) -
Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization
by: Tang, Changli, et al.
Published: (2024) -
Whitened CLIP as a Likelihood Surrogate of Images and Captions
by: Betser, Roy, et al.
Published: (2025)