Where Do Images Come From? Analyzing Captions to Geographically Profile Datasets
Fuente:
arXiv
Guardado en:
| Autores principales: | Basu, Abhipsa, Bahl, Yugam, Bhagat, Kirti, Seshadri, Preethi, Babu, R. Venkatesh, Pruthi, Danish |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
GeoDiv: Framework For Measuring Geographical Diversity In Text-To-Image Models
por: Basu, Abhipsa, et al.
Publicado: (2026)
por: Basu, Abhipsa, et al.
Publicado: (2026)
Harnessing Diffusion-Generated Synthetic Images for Fair Image Classification
por: Basu, Abhipsa, et al.
Publicado: (2025)
por: Basu, Abhipsa, et al.
Publicado: (2025)
Balancing Act: Distribution-Guided Debiasing in Diffusion Models
por: Parihar, Rishubh, et al.
Publicado: (2024)
por: Parihar, Rishubh, et al.
Publicado: (2024)
Richer Output for Richer Countries: Uncovering Geographical Disparities in Generated Stories and Travel Recommendations
por: Bhagat, Kirti, et al.
Publicado: (2024)
por: Bhagat, Kirti, et al.
Publicado: (2024)
Do Vision Language Models Need to Process Image Tokens?
por: Ghosh, Sambit, et al.
Publicado: (2026)
por: Ghosh, Sambit, et al.
Publicado: (2026)
PhraseStereo: The First Open-Vocabulary Stereo Image Segmentation Dataset
por: Campagnolo, Thomas, et al.
Publicado: (2025)
por: Campagnolo, Thomas, et al.
Publicado: (2025)
Analyzing Transformer Models and Knowledge Distillation Approaches for Image Captioning on Edge AI
por: Kwok, Wing Man Casca, et al.
Publicado: (2025)
por: Kwok, Wing Man Casca, et al.
Publicado: (2025)
Analyzing Image Beyond Visual Aspect: Image Emotion Classification via Multiple-Affective Captioning
por: Zhou, Zibo, et al.
Publicado: (2025)
por: Zhou, Zibo, et al.
Publicado: (2025)
Compass Control: Multi Object Orientation Control for Text-to-Image Generation
por: Parihar, Rishubh, et al.
Publicado: (2025)
por: Parihar, Rishubh, et al.
Publicado: (2025)
Sieve: Multimodal Dataset Pruning Using Image Captioning Models
por: Mahmoud, Anas, et al.
Publicado: (2023)
por: Mahmoud, Anas, et al.
Publicado: (2023)
How Many Images Does It Take? Estimating Imitation Thresholds in Text-to-Image Models
por: Verma, Sahil, et al.
Publicado: (2024)
por: Verma, Sahil, et al.
Publicado: (2024)
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
por: Cheng, Kanzhi, et al.
Publicado: (2025)
por: Cheng, Kanzhi, et al.
Publicado: (2025)
Leveraging Vision-Language Models for Improving Domain Generalization in Image Classification
por: Addepalli, Sravanti, et al.
Publicado: (2023)
por: Addepalli, Sravanti, et al.
Publicado: (2023)
From Pixels to Prose: A Large Dataset of Dense Image Captions
por: Singla, Vasu, et al.
Publicado: (2024)
por: Singla, Vasu, et al.
Publicado: (2024)
From Image Captioning to Visual Storytelling
por: Passadakis, Admitos, et al.
Publicado: (2025)
por: Passadakis, Admitos, et al.
Publicado: (2025)
PreciseControl: Enhancing Text-To-Image Diffusion Models with Fine-Grained Attribute Control
por: Parihar, Rishubh, et al.
Publicado: (2024)
por: Parihar, Rishubh, et al.
Publicado: (2024)
MirrorVerse: Pushing Diffusion Models to Realistically Reflect the World
por: Dhiman, Ankit, et al.
Publicado: (2025)
por: Dhiman, Ankit, et al.
Publicado: (2025)
CaptionQA: Is Your Caption as Useful as the Image Itself?
por: Yang, Shijia, et al.
Publicado: (2025)
por: Yang, Shijia, et al.
Publicado: (2025)
CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning
por: Saito, Kuniaki, et al.
Publicado: (2025)
por: Saito, Kuniaki, et al.
Publicado: (2025)
ICC: Quantifying Image Caption Concreteness for Multimodal Dataset Curation
por: Yanuka, Moran, et al.
Publicado: (2024)
por: Yanuka, Moran, et al.
Publicado: (2024)
KTVIC: A Vietnamese Image Captioning Dataset on the Life Domain
por: Pham, Anh-Cuong, et al.
Publicado: (2024)
por: Pham, Anh-Cuong, et al.
Publicado: (2024)
Rethinking Dataset Distillation: Hard Truths about Soft Labels
por: Dey, Priyam, et al.
Publicado: (2026)
por: Dey, Priyam, et al.
Publicado: (2026)
Generating an Image From 1,000 Words: Enhancing Text-to-Image With Structured Captions
por: Gutflaish, Eyal, et al.
Publicado: (2025)
por: Gutflaish, Eyal, et al.
Publicado: (2025)
Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization
por: Bharadwaj, Siddhant, et al.
Publicado: (2026)
por: Bharadwaj, Siddhant, et al.
Publicado: (2026)
DEJIMA: A Novel Large-scale Japanese Dataset for Image Captioning and Visual Question Answering
por: Katsube, Toshiki, et al.
Publicado: (2025)
por: Katsube, Toshiki, et al.
Publicado: (2025)
SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning
por: Zhang, Lin, et al.
Publicado: (2025)
por: Zhang, Lin, et al.
Publicado: (2025)
Small Dents, Big Impact: A Dataset and Deep Learning Approach for Vehicle Dent Detection
por: Baig, Danish Zia, et al.
Publicado: (2025)
por: Baig, Danish Zia, et al.
Publicado: (2025)
PixLore: A Dataset-driven Approach to Rich Image Captioning
por: Bonilla-Salvador, Diego, et al.
Publicado: (2023)
por: Bonilla-Salvador, Diego, et al.
Publicado: (2023)
#PraCegoVer: A Large Dataset for Image Captioning in Portuguese
por: Santos, Gabriel Oliveira dos, et al.
Publicado: (2021)
por: Santos, Gabriel Oliveira dos, et al.
Publicado: (2021)
A Manually Annotated Image-Caption Dataset for Detecting Children in the Wild
por: Kireev, Klim, et al.
Publicado: (2025)
por: Kireev, Klim, et al.
Publicado: (2025)
Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning
por: Tosato, Lucrezia, et al.
Publicado: (2026)
por: Tosato, Lucrezia, et al.
Publicado: (2026)
Text2Place: Affordance-aware Text Guided Human Placement
por: Parihar, Rishubh, et al.
Publicado: (2024)
por: Parihar, Rishubh, et al.
Publicado: (2024)
A Comparative Study of Image Denoising Algorithms
por: Danish, Muhammad Umair
Publicado: (2024)
por: Danish, Muhammad Umair
Publicado: (2024)
What Makes for Good Image Captions?
por: Chen, Delong, et al.
Publicado: (2024)
por: Chen, Delong, et al.
Publicado: (2024)
Benchmarking and Improving Detail Image Caption
por: Dong, Hongyuan, et al.
Publicado: (2024)
por: Dong, Hongyuan, et al.
Publicado: (2024)
LuoJiaHOG: A Hierarchy Oriented Geo-aware Image Caption Dataset for Remote Sensing Image-Text Retrival
por: Zhao, Yuanxin, et al.
Publicado: (2024)
por: Zhao, Yuanxin, et al.
Publicado: (2024)
From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment
por: Hirota, Yusuke, et al.
Publicado: (2024)
por: Hirota, Yusuke, et al.
Publicado: (2024)
Image Generation from Image Captioning -- Invertible Approach
por: Menon, Nandakishore S, et al.
Publicado: (2024)
por: Menon, Nandakishore S, et al.
Publicado: (2024)
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
por: Peng, Cihang, et al.
Publicado: (2025)
por: Peng, Cihang, et al.
Publicado: (2025)
ProFeAT: Projected Feature Adversarial Training for Self-Supervised Learning of Robust Representations
por: Addepalli, Sravanti, et al.
Publicado: (2024)
por: Addepalli, Sravanti, et al.
Publicado: (2024)
Ejemplares similares
-
GeoDiv: Framework For Measuring Geographical Diversity In Text-To-Image Models
por: Basu, Abhipsa, et al.
Publicado: (2026) -
Harnessing Diffusion-Generated Synthetic Images for Fair Image Classification
por: Basu, Abhipsa, et al.
Publicado: (2025) -
Balancing Act: Distribution-Guided Debiasing in Diffusion Models
por: Parihar, Rishubh, et al.
Publicado: (2024) -
Richer Output for Richer Countries: Uncovering Geographical Disparities in Generated Stories and Travel Recommendations
por: Bhagat, Kirti, et al.
Publicado: (2024) -
Do Vision Language Models Need to Process Image Tokens?
por: Ghosh, Sambit, et al.
Publicado: (2026)