Where Do Images Come From? Analyzing Captions to Geographically Profile Datasets
Fuente:
arXiv
Saved in:
| Main Authors: | Basu, Abhipsa, Bahl, Yugam, Bhagat, Kirti, Seshadri, Preethi, Babu, R. Venkatesh, Pruthi, Danish |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GeoDiv: Framework For Measuring Geographical Diversity In Text-To-Image Models
by: Basu, Abhipsa, et al.
Published: (2026)
by: Basu, Abhipsa, et al.
Published: (2026)
Harnessing Diffusion-Generated Synthetic Images for Fair Image Classification
by: Basu, Abhipsa, et al.
Published: (2025)
by: Basu, Abhipsa, et al.
Published: (2025)
Balancing Act: Distribution-Guided Debiasing in Diffusion Models
by: Parihar, Rishubh, et al.
Published: (2024)
by: Parihar, Rishubh, et al.
Published: (2024)
Richer Output for Richer Countries: Uncovering Geographical Disparities in Generated Stories and Travel Recommendations
by: Bhagat, Kirti, et al.
Published: (2024)
by: Bhagat, Kirti, et al.
Published: (2024)
Do Vision Language Models Need to Process Image Tokens?
by: Ghosh, Sambit, et al.
Published: (2026)
by: Ghosh, Sambit, et al.
Published: (2026)
PhraseStereo: The First Open-Vocabulary Stereo Image Segmentation Dataset
by: Campagnolo, Thomas, et al.
Published: (2025)
by: Campagnolo, Thomas, et al.
Published: (2025)
Analyzing Transformer Models and Knowledge Distillation Approaches for Image Captioning on Edge AI
by: Kwok, Wing Man Casca, et al.
Published: (2025)
by: Kwok, Wing Man Casca, et al.
Published: (2025)
Analyzing Image Beyond Visual Aspect: Image Emotion Classification via Multiple-Affective Captioning
by: Zhou, Zibo, et al.
Published: (2025)
by: Zhou, Zibo, et al.
Published: (2025)
Compass Control: Multi Object Orientation Control for Text-to-Image Generation
by: Parihar, Rishubh, et al.
Published: (2025)
by: Parihar, Rishubh, et al.
Published: (2025)
Sieve: Multimodal Dataset Pruning Using Image Captioning Models
by: Mahmoud, Anas, et al.
Published: (2023)
by: Mahmoud, Anas, et al.
Published: (2023)
How Many Images Does It Take? Estimating Imitation Thresholds in Text-to-Image Models
by: Verma, Sahil, et al.
Published: (2024)
by: Verma, Sahil, et al.
Published: (2024)
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
by: Cheng, Kanzhi, et al.
Published: (2025)
by: Cheng, Kanzhi, et al.
Published: (2025)
Leveraging Vision-Language Models for Improving Domain Generalization in Image Classification
by: Addepalli, Sravanti, et al.
Published: (2023)
by: Addepalli, Sravanti, et al.
Published: (2023)
From Pixels to Prose: A Large Dataset of Dense Image Captions
by: Singla, Vasu, et al.
Published: (2024)
by: Singla, Vasu, et al.
Published: (2024)
From Image Captioning to Visual Storytelling
by: Passadakis, Admitos, et al.
Published: (2025)
by: Passadakis, Admitos, et al.
Published: (2025)
PreciseControl: Enhancing Text-To-Image Diffusion Models with Fine-Grained Attribute Control
by: Parihar, Rishubh, et al.
Published: (2024)
by: Parihar, Rishubh, et al.
Published: (2024)
MirrorVerse: Pushing Diffusion Models to Realistically Reflect the World
by: Dhiman, Ankit, et al.
Published: (2025)
by: Dhiman, Ankit, et al.
Published: (2025)
CaptionQA: Is Your Caption as Useful as the Image Itself?
by: Yang, Shijia, et al.
Published: (2025)
by: Yang, Shijia, et al.
Published: (2025)
CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning
by: Saito, Kuniaki, et al.
Published: (2025)
by: Saito, Kuniaki, et al.
Published: (2025)
ICC: Quantifying Image Caption Concreteness for Multimodal Dataset Curation
by: Yanuka, Moran, et al.
Published: (2024)
by: Yanuka, Moran, et al.
Published: (2024)
KTVIC: A Vietnamese Image Captioning Dataset on the Life Domain
by: Pham, Anh-Cuong, et al.
Published: (2024)
by: Pham, Anh-Cuong, et al.
Published: (2024)
Rethinking Dataset Distillation: Hard Truths about Soft Labels
by: Dey, Priyam, et al.
Published: (2026)
by: Dey, Priyam, et al.
Published: (2026)
Generating an Image From 1,000 Words: Enhancing Text-to-Image With Structured Captions
by: Gutflaish, Eyal, et al.
Published: (2025)
by: Gutflaish, Eyal, et al.
Published: (2025)
Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization
by: Bharadwaj, Siddhant, et al.
Published: (2026)
by: Bharadwaj, Siddhant, et al.
Published: (2026)
DEJIMA: A Novel Large-scale Japanese Dataset for Image Captioning and Visual Question Answering
by: Katsube, Toshiki, et al.
Published: (2025)
by: Katsube, Toshiki, et al.
Published: (2025)
SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning
by: Zhang, Lin, et al.
Published: (2025)
by: Zhang, Lin, et al.
Published: (2025)
Small Dents, Big Impact: A Dataset and Deep Learning Approach for Vehicle Dent Detection
by: Baig, Danish Zia, et al.
Published: (2025)
by: Baig, Danish Zia, et al.
Published: (2025)
PixLore: A Dataset-driven Approach to Rich Image Captioning
by: Bonilla-Salvador, Diego, et al.
Published: (2023)
by: Bonilla-Salvador, Diego, et al.
Published: (2023)
#PraCegoVer: A Large Dataset for Image Captioning in Portuguese
by: Santos, Gabriel Oliveira dos, et al.
Published: (2021)
by: Santos, Gabriel Oliveira dos, et al.
Published: (2021)
A Manually Annotated Image-Caption Dataset for Detecting Children in the Wild
by: Kireev, Klim, et al.
Published: (2025)
by: Kireev, Klim, et al.
Published: (2025)
Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning
by: Tosato, Lucrezia, et al.
Published: (2026)
by: Tosato, Lucrezia, et al.
Published: (2026)
Text2Place: Affordance-aware Text Guided Human Placement
by: Parihar, Rishubh, et al.
Published: (2024)
by: Parihar, Rishubh, et al.
Published: (2024)
A Comparative Study of Image Denoising Algorithms
by: Danish, Muhammad Umair
Published: (2024)
by: Danish, Muhammad Umair
Published: (2024)
What Makes for Good Image Captions?
by: Chen, Delong, et al.
Published: (2024)
by: Chen, Delong, et al.
Published: (2024)
Benchmarking and Improving Detail Image Caption
by: Dong, Hongyuan, et al.
Published: (2024)
by: Dong, Hongyuan, et al.
Published: (2024)
LuoJiaHOG: A Hierarchy Oriented Geo-aware Image Caption Dataset for Remote Sensing Image-Text Retrival
by: Zhao, Yuanxin, et al.
Published: (2024)
by: Zhao, Yuanxin, et al.
Published: (2024)
From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment
by: Hirota, Yusuke, et al.
Published: (2024)
by: Hirota, Yusuke, et al.
Published: (2024)
Image Generation from Image Captioning -- Invertible Approach
by: Menon, Nandakishore S, et al.
Published: (2024)
by: Menon, Nandakishore S, et al.
Published: (2024)
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
by: Peng, Cihang, et al.
Published: (2025)
by: Peng, Cihang, et al.
Published: (2025)
ProFeAT: Projected Feature Adversarial Training for Self-Supervised Learning of Robust Representations
by: Addepalli, Sravanti, et al.
Published: (2024)
by: Addepalli, Sravanti, et al.
Published: (2024)
Similar Items
-
GeoDiv: Framework For Measuring Geographical Diversity In Text-To-Image Models
by: Basu, Abhipsa, et al.
Published: (2026) -
Harnessing Diffusion-Generated Synthetic Images for Fair Image Classification
by: Basu, Abhipsa, et al.
Published: (2025) -
Balancing Act: Distribution-Guided Debiasing in Diffusion Models
by: Parihar, Rishubh, et al.
Published: (2024) -
Richer Output for Richer Countries: Uncovering Geographical Disparities in Generated Stories and Travel Recommendations
by: Bhagat, Kirti, et al.
Published: (2024) -
Do Vision Language Models Need to Process Image Tokens?
by: Ghosh, Sambit, et al.
Published: (2026)