ANNA: Abstractive Text-to-Image Synthesis with Filtered News Captions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ramakrishnan, Aashish Anantha, Huang, Sharon X., Lee, Dongwon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ANCHOR: LLM-driven Subject Conditioning for Text-to-Image Synthesis
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2024)
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2024)
RONA: Pragmatically Diverse Image Captioning with Coherence Relations
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025)
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025)
CORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships?
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025)
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025)
IRONIC: Coherence-Aware Reasoning Chains for Multi-Modal Sarcasm Detection
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025)
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025)
IMM-MOT: A Novel 3D Multi-object Tracking Framework with Interacting Multiple Model Filter
von: Liu, Xiaohong, et al.
Veröffentlicht: (2025)
von: Liu, Xiaohong, et al.
Veröffentlicht: (2025)
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles
von: Ng, Ho Yin 'Sam', et al.
Veröffentlicht: (2025)
von: Ng, Ho Yin 'Sam', et al.
Veröffentlicht: (2025)
A review of Recent Techniques for Person Re-Identification
von: Asperti, Andrea, et al.
Veröffentlicht: (2025)
von: Asperti, Andrea, et al.
Veröffentlicht: (2025)
A Lightweight Measure of Classification Difficulty from Application Dataset Characteristics
von: Cao, Bryan Bo, et al.
Veröffentlicht: (2024)
von: Cao, Bryan Bo, et al.
Veröffentlicht: (2024)
Enhancement of 3D Camera Synthetic Training Data with Noise Models
von: Osvaldová, Katarína, et al.
Veröffentlicht: (2024)
von: Osvaldová, Katarína, et al.
Veröffentlicht: (2024)
Spectral Fidelity and Spatial Enhancement: An Assessment and Cascading of Pan-Sharpening Techniques for Satellite Imagery
von: B, Abdul Aziz A., et al.
Veröffentlicht: (2024)
von: B, Abdul Aziz A., et al.
Veröffentlicht: (2024)
Text-only Synthesis for Image Captioning
von: Zhou, Qing, et al.
Veröffentlicht: (2024)
von: Zhou, Qing, et al.
Veröffentlicht: (2024)
Mathematical Analysis of Image Matching Techniques
von: Samoilenko, Oleh
Veröffentlicht: (2026)
von: Samoilenko, Oleh
Veröffentlicht: (2026)
MUNIChus: Multilingual News Image Captioning Benchmark
von: Chen, Yuji, et al.
Veröffentlicht: (2026)
von: Chen, Yuji, et al.
Veröffentlicht: (2026)
Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
VIXEN: Visual Text Comparison Network for Image Difference Captioning
von: Black, Alexander, et al.
Veröffentlicht: (2024)
von: Black, Alexander, et al.
Veröffentlicht: (2024)
Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model
von: Cheng, Sheng, et al.
Veröffentlicht: (2024)
von: Cheng, Sheng, et al.
Veröffentlicht: (2024)
On Representation of 3D Rotation in the Context of Deep Learning
von: Pravdová, Viktória, et al.
Veröffentlicht: (2024)
von: Pravdová, Viktória, et al.
Veröffentlicht: (2024)
Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
von: Fonseca, Rui, et al.
Veröffentlicht: (2025)
von: Fonseca, Rui, et al.
Veröffentlicht: (2025)
Beyond Filtering: Adaptive Image-Text Quality Enhancement for MLLM Pretraining
von: Huang, Han, et al.
Veröffentlicht: (2024)
von: Huang, Han, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Long-Context Translation for Cultural Image Captioning: Gators submission for AmericasNLP 2026 shared task
von: Dhawan, Aashish, et al.
Veröffentlicht: (2026)
von: Dhawan, Aashish, et al.
Veröffentlicht: (2026)
EAMA : Entity-Aware Multimodal Alignment Based Approach for News Image Captioning
von: Zhang, Junzhe, et al.
Veröffentlicht: (2024)
von: Zhang, Junzhe, et al.
Veröffentlicht: (2024)
From Image Captioning to Visual Storytelling
von: Passadakis, Admitos, et al.
Veröffentlicht: (2025)
von: Passadakis, Admitos, et al.
Veröffentlicht: (2025)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
DiffChat: Learning to Chat with Text-to-Image Synthesis Models for Interactive Image Creation
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025)
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025)
Multimodal Abstractive Summarization of Instructional Videos with Vision-Language Models
von: Nazir, Maham, et al.
Veröffentlicht: (2026)
von: Nazir, Maham, et al.
Veröffentlicht: (2026)
From Intentions to Techniques: A Comprehensive Taxonomy and Challenges in Text Watermarking for Large Language Models
von: Lalai, Harsh Nishant, et al.
Veröffentlicht: (2024)
von: Lalai, Harsh Nishant, et al.
Veröffentlicht: (2024)
See or Guess: Counterfactually Regularized Image Captioning
von: Cao, Qian, et al.
Veröffentlicht: (2024)
von: Cao, Qian, et al.
Veröffentlicht: (2024)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
von: Nguyen, Nhi Ngoc-Yen, et al.
Veröffentlicht: (2026)
von: Nguyen, Nhi Ngoc-Yen, et al.
Veröffentlicht: (2026)
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
von: Liu, Zheng, et al.
Veröffentlicht: (2024)
von: Liu, Zheng, et al.
Veröffentlicht: (2024)
IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning
von: Lee, Soeun, et al.
Veröffentlicht: (2024)
von: Lee, Soeun, et al.
Veröffentlicht: (2024)
Image Captioning via Compact Bidirectional Architecture
von: Song, Zijie, et al.
Veröffentlicht: (2022)
von: Song, Zijie, et al.
Veröffentlicht: (2022)
Finetuned Multimodal Language Models Are High-Quality Image-Text Data Filters
von: Wang, Weizhi, et al.
Veröffentlicht: (2024)
von: Wang, Weizhi, et al.
Veröffentlicht: (2024)
Altogether: Image Captioning via Re-aligning Alt-text
von: Xu, Hu, et al.
Veröffentlicht: (2024)
von: Xu, Hu, et al.
Veröffentlicht: (2024)
Multi-LLM Collaborative Caption Generation in Scientific Documents
von: Kim, Jaeyoung, et al.
Veröffentlicht: (2025)
von: Kim, Jaeyoung, et al.
Veröffentlicht: (2025)
DocSum: Domain-Adaptive Pre-training for Document Abstractive Summarization
von: Chau, Phan Phuong Mai, et al.
Veröffentlicht: (2024)
von: Chau, Phan Phuong Mai, et al.
Veröffentlicht: (2024)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
von: Luo, Jianjie, et al.
Veröffentlicht: (2024)
von: Luo, Jianjie, et al.
Veröffentlicht: (2024)
Enhancing Journalism with AI: A Study of Contextualized Image Captioning for News Articles using LLMs and LMMs
von: Anagnostopoulou, Aliki, et al.
Veröffentlicht: (2024)
von: Anagnostopoulou, Aliki, et al.
Veröffentlicht: (2024)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
von: Li, Wenyan, et al.
Veröffentlicht: (2024)
von: Li, Wenyan, et al.
Veröffentlicht: (2024)
Temporal Image Caption Retrieval Competition -- Description and Results
von: Pokrywka, Jakub, et al.
Veröffentlicht: (2024)
von: Pokrywka, Jakub, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ANCHOR: LLM-driven Subject Conditioning for Text-to-Image Synthesis
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2024) -
RONA: Pragmatically Diverse Image Captioning with Coherence Relations
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025) -
CORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships?
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025) -
IRONIC: Coherence-Aware Reasoning Chains for Multi-Modal Sarcasm Detection
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2025) -
IMM-MOT: A Novel 3D Multi-object Tracking Framework with Interacting Multiple Model Filter
von: Liu, Xiaohong, et al.
Veröffentlicht: (2025)