Emergent Natural Language with Communication Games for Improving Image Captioning Capabilities without Additional Data
Fuente:
arXiv
Guardado en:
| Autores principales: | Dutta, Parag, Dukkipati, Ambedkar |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Semi-supervised Deep Transfer for Regression without Domain Alignment
por: Biswas, Mainak, et al.
Publicado: (2025)
por: Biswas, Mainak, et al.
Publicado: (2025)
Saddle-Free Guidance: Improved On-Manifold Sampling without Labels or Additional Training
por: Yeats, Eric, et al.
Publicado: (2025)
por: Yeats, Eric, et al.
Publicado: (2025)
Image Captions are Natural Prompts for Text-to-Image Models
por: Lei, Shiye, et al.
Publicado: (2023)
por: Lei, Shiye, et al.
Publicado: (2023)
Image-Caption Encoding for Improving Zero-Shot Generalization
por: Yu, Eric Yang, et al.
Publicado: (2024)
por: Yu, Eric Yang, et al.
Publicado: (2024)
Pixels to Prose: Understanding the art of Image Captioning
por: Singh, Hrishikesh, et al.
Publicado: (2024)
por: Singh, Hrishikesh, et al.
Publicado: (2024)
Pretrained Image-Text Models are Secretly Video Captioners
por: Zhang, Chunhui, et al.
Publicado: (2025)
por: Zhang, Chunhui, et al.
Publicado: (2025)
Differentially Private Representation Learning via Image Captioning
por: Sander, Tom, et al.
Publicado: (2024)
por: Sander, Tom, et al.
Publicado: (2024)
Modeling Image-Caption Rating from Comparative Judgments
por: Minni, Kezia, et al.
Publicado: (2026)
por: Minni, Kezia, et al.
Publicado: (2026)
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
por: Merchant, Nicholas, et al.
Publicado: (2025)
por: Merchant, Nicholas, et al.
Publicado: (2025)
CompCap: Improving Multimodal Large Language Models with Composite Captions
por: Chen, Xiaohui, et al.
Publicado: (2024)
por: Chen, Xiaohui, et al.
Publicado: (2024)
ICC: Quantifying Image Caption Concreteness for Multimodal Dataset Curation
por: Yanuka, Moran, et al.
Publicado: (2024)
por: Yanuka, Moran, et al.
Publicado: (2024)
Using Images from a Video Game to Improve the Detection of Truck Axles
por: Marcomini, Leandro Arab, et al.
Publicado: (2025)
por: Marcomini, Leandro Arab, et al.
Publicado: (2025)
Infusing Environmental Captions for Long-Form Video Language Grounding
por: Lee, Hyogun, et al.
Publicado: (2024)
por: Lee, Hyogun, et al.
Publicado: (2024)
Hyperdimensional Cross-Modal Alignment of Frozen Language and Image Models for Efficient Image Captioning
por: Dalvi, Abhishek, et al.
Publicado: (2026)
por: Dalvi, Abhishek, et al.
Publicado: (2026)
AIM: Additional Image Guided Generation of Transferable Adversarial Attacks
por: Li, Teng, et al.
Publicado: (2025)
por: Li, Teng, et al.
Publicado: (2025)
Vision4PPG: Emergent PPG Analysis Capability of Vision Foundation Models for Vital Signs like Blood Pressure
por: Kataria, Saurabh, et al.
Publicado: (2025)
por: Kataria, Saurabh, et al.
Publicado: (2025)
Linear Alignment of Vision-language Models for Image Captioning
por: Paischer, Fabian, et al.
Publicado: (2023)
por: Paischer, Fabian, et al.
Publicado: (2023)
Generalizable Geometric Image Caption Synthesis
por: Xin, Yue, et al.
Publicado: (2025)
por: Xin, Yue, et al.
Publicado: (2025)
PICS: Pipeline for Image Captioning and Search
por: Rosario, Grant, et al.
Publicado: (2024)
por: Rosario, Grant, et al.
Publicado: (2024)
It's a (Blind) Match! Towards Vision-Language Correspondence without Parallel Data
por: Schnaus, Dominik, et al.
Publicado: (2025)
por: Schnaus, Dominik, et al.
Publicado: (2025)
Improving Image Captioning Descriptiveness by Ranking and LLM-based Fusion
por: Celona, Luigi, et al.
Publicado: (2023)
por: Celona, Luigi, et al.
Publicado: (2023)
Image Captioning as an Assistive Technology: Lessons Learned from VizWiz 2020 Challenge
por: Dognin, Pierre, et al.
Publicado: (2020)
por: Dognin, Pierre, et al.
Publicado: (2020)
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
por: Lai, Zhengfeng, et al.
Publicado: (2024)
por: Lai, Zhengfeng, et al.
Publicado: (2024)
STAND: Semantic Anchoring Constraint with Dual-Granularity Disambiguation for Remote Sensing Image Change Captioning
por: Gong, Yanpei, et al.
Publicado: (2026)
por: Gong, Yanpei, et al.
Publicado: (2026)
Concept-Aware Batch Sampling Improves Language-Image Pretraining
por: Ghosh, Adhiraj, et al.
Publicado: (2025)
por: Ghosh, Adhiraj, et al.
Publicado: (2025)
Learning an Image Editing Model without Image Editing Pairs
por: Kumari, Nupur, et al.
Publicado: (2025)
por: Kumari, Nupur, et al.
Publicado: (2025)
Reasoning-Guided Grounding: Elevating Video Anomaly Detection through Multimodal Large Language Models
por: Agarwal, Sakshi, et al.
Publicado: (2026)
por: Agarwal, Sakshi, et al.
Publicado: (2026)
From Pixels to Prose: A Large Dataset of Dense Image Captions
por: Singla, Vasu, et al.
Publicado: (2024)
por: Singla, Vasu, et al.
Publicado: (2024)
OCT Data is All You Need: How Vision Transformers with and without Pre-training Benefit Imaging
por: Han, Zihao, et al.
Publicado: (2025)
por: Han, Zihao, et al.
Publicado: (2025)
Improved Generation of Synthetic Imaging Data Using Feature-Aligned Diffusion
por: Nair, Lakshmi
Publicado: (2024)
por: Nair, Lakshmi
Publicado: (2024)
No Hard Negatives Required: Concept Centric Learning Leads to Compositionality without Degrading Zero-shot Capabilities of Contrastive Models
por: Pham, Hai X., et al.
Publicado: (2026)
por: Pham, Hai X., et al.
Publicado: (2026)
Good Representation, Better Explanation: Role of Convolutional Neural Networks in Transformer-Based Remote Sensing Image Captioning
por: Das, Swadhin, et al.
Publicado: (2025)
por: Das, Swadhin, et al.
Publicado: (2025)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
por: Luo, Jianjie, et al.
Publicado: (2024)
por: Luo, Jianjie, et al.
Publicado: (2024)
Large VLM-based Stylized Sports Captioning
por: Dhar, Sauptik, et al.
Publicado: (2025)
por: Dhar, Sauptik, et al.
Publicado: (2025)
Bridge the Modality and Capability Gaps in Vision-Language Model Selection
por: Yi, Chao, et al.
Publicado: (2024)
por: Yi, Chao, et al.
Publicado: (2024)
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
por: Kim, Si-Woo, et al.
Publicado: (2025)
por: Kim, Si-Woo, et al.
Publicado: (2025)
Learning without Forgetting for Vision-Language Models
por: Zhou, Da-Wei, et al.
Publicado: (2023)
por: Zhou, Da-Wei, et al.
Publicado: (2023)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
por: Lai, Zhengfeng, et al.
Publicado: (2023)
por: Lai, Zhengfeng, et al.
Publicado: (2023)
GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning
por: Raju, S M Taslim Uddin, et al.
Publicado: (2025)
por: Raju, S M Taslim Uddin, et al.
Publicado: (2025)
Learning to Rank Caption Chains for Video-Text Alignment
por: Blume, Ansel, et al.
Publicado: (2026)
por: Blume, Ansel, et al.
Publicado: (2026)
Ejemplares similares
-
Semi-supervised Deep Transfer for Regression without Domain Alignment
por: Biswas, Mainak, et al.
Publicado: (2025) -
Saddle-Free Guidance: Improved On-Manifold Sampling without Labels or Additional Training
por: Yeats, Eric, et al.
Publicado: (2025) -
Image Captions are Natural Prompts for Text-to-Image Models
por: Lei, Shiye, et al.
Publicado: (2023) -
Image-Caption Encoding for Improving Zero-Shot Generalization
por: Yu, Eric Yang, et al.
Publicado: (2024) -
Pixels to Prose: Understanding the art of Image Captioning
por: Singh, Hrishikesh, et al.
Publicado: (2024)