Pixels to Prose: Understanding the art of Image Captioning
Fuente:
arXiv
Salvato in:
| Autori principali: | Singh, Hrishikesh, Sharma, Aarti, Pant, Millie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Pixels to Prose: A Large Dataset of Dense Image Captions
di: Singla, Vasu, et al.
Pubblicazione: (2024)
di: Singla, Vasu, et al.
Pubblicazione: (2024)
From Pixels to Prose: Advancing Multi-Modal Language Models for Remote Sensing
di: Sun, Xintian, et al.
Pubblicazione: (2024)
di: Sun, Xintian, et al.
Pubblicazione: (2024)
Good Representation, Better Explanation: Role of Convolutional Neural Networks in Transformer-Based Remote Sensing Image Captioning
di: Das, Swadhin, et al.
Pubblicazione: (2025)
di: Das, Swadhin, et al.
Pubblicazione: (2025)
Low-light Pedestrian Detection in Visible and Infrared Image Feeds: Issues and Challenges
di: Akilan, Thangarajah, et al.
Pubblicazione: (2023)
di: Akilan, Thangarajah, et al.
Pubblicazione: (2023)
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
di: Woo, Byeongju, et al.
Pubblicazione: (2026)
di: Woo, Byeongju, et al.
Pubblicazione: (2026)
Differentially Private Representation Learning via Image Captioning
di: Sander, Tom, et al.
Pubblicazione: (2024)
di: Sander, Tom, et al.
Pubblicazione: (2024)
Modeling Image-Caption Rating from Comparative Judgments
di: Minni, Kezia, et al.
Pubblicazione: (2026)
di: Minni, Kezia, et al.
Pubblicazione: (2026)
Pretrained Image-Text Models are Secretly Video Captioners
di: Zhang, Chunhui, et al.
Pubblicazione: (2025)
di: Zhang, Chunhui, et al.
Pubblicazione: (2025)
MsEdF: A Multi-stream Encoder-decoder Framework for Remote Sensing Image Captioning
di: Das, Swadhin, et al.
Pubblicazione: (2025)
di: Das, Swadhin, et al.
Pubblicazione: (2025)
Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning
di: You, Zuyao, et al.
Pubblicazione: (2025)
di: You, Zuyao, et al.
Pubblicazione: (2025)
Surveying Facial Recognition Models for Diverse Indian Demographics: A Comparative Analysis on LFW and Custom Dataset
di: Pant, Pranav, et al.
Pubblicazione: (2024)
di: Pant, Pranav, et al.
Pubblicazione: (2024)
ICC: Quantifying Image Caption Concreteness for Multimodal Dataset Curation
di: Yanuka, Moran, et al.
Pubblicazione: (2024)
di: Yanuka, Moran, et al.
Pubblicazione: (2024)
Introducing Resizable Region Packing Problem in Image Generation, with a Heuristic Solution
di: Sharma, Hrishikesh
Pubblicazione: (2025)
di: Sharma, Hrishikesh
Pubblicazione: (2025)
Role of Locality and Weight Sharing in Image-Based Tasks: A Sample Complexity Separation between CNNs, LCNs, and FCNs
di: Lahoti, Aakash, et al.
Pubblicazione: (2024)
di: Lahoti, Aakash, et al.
Pubblicazione: (2024)
From Semantics to Pixels: Coarse-to-Fine Masked Autoencoders for Hierarchical Visual Understanding
di: Xiang, Wenzhao, et al.
Pubblicazione: (2026)
di: Xiang, Wenzhao, et al.
Pubblicazione: (2026)
Pixel-level Counterfactual Contrastive Learning for Medical Image Segmentation
di: Lafargue-Hauret, Marceau, et al.
Pubblicazione: (2026)
di: Lafargue-Hauret, Marceau, et al.
Pubblicazione: (2026)
Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models
di: NVIDIA, et al.
Pubblicazione: (2024)
di: NVIDIA, et al.
Pubblicazione: (2024)
Latent Forcing: Reordering the Diffusion Trajectory for Pixel-Space Image Generation
di: Baade, Alan, et al.
Pubblicazione: (2026)
di: Baade, Alan, et al.
Pubblicazione: (2026)
FREPix: Frequency-Heterogeneous Flow Matching for Pixel-Space Image Generation
di: Lin, Mingfeng, et al.
Pubblicazione: (2026)
di: Lin, Mingfeng, et al.
Pubblicazione: (2026)
Image-Caption Encoding for Improving Zero-Shot Generalization
di: Yu, Eric Yang, et al.
Pubblicazione: (2024)
di: Yu, Eric Yang, et al.
Pubblicazione: (2024)
Linear Alignment of Vision-language Models for Image Captioning
di: Paischer, Fabian, et al.
Pubblicazione: (2023)
di: Paischer, Fabian, et al.
Pubblicazione: (2023)
PICS: Pipeline for Image Captioning and Search
di: Rosario, Grant, et al.
Pubblicazione: (2024)
di: Rosario, Grant, et al.
Pubblicazione: (2024)
Generalizable Geometric Image Caption Synthesis
di: Xin, Yue, et al.
Pubblicazione: (2025)
di: Xin, Yue, et al.
Pubblicazione: (2025)
Image Captioning as an Assistive Technology: Lessons Learned from VizWiz 2020 Challenge
di: Dognin, Pierre, et al.
Pubblicazione: (2020)
di: Dognin, Pierre, et al.
Pubblicazione: (2020)
Image Captions are Natural Prompts for Text-to-Image Models
di: Lei, Shiye, et al.
Pubblicazione: (2023)
di: Lei, Shiye, et al.
Pubblicazione: (2023)
From Pixels to Graphs: Deep Graph-Level Anomaly Detection on Dermoscopic Images
di: Xu, Dehn, et al.
Pubblicazione: (2025)
di: Xu, Dehn, et al.
Pubblicazione: (2025)
Pixel-Wise Recognition for Holistic Surgical Scene Understanding
di: Ayobi, Nicolás, et al.
Pubblicazione: (2024)
di: Ayobi, Nicolás, et al.
Pubblicazione: (2024)
STAND: Semantic Anchoring Constraint with Dual-Granularity Disambiguation for Remote Sensing Image Change Captioning
di: Gong, Yanpei, et al.
Pubblicazione: (2026)
di: Gong, Yanpei, et al.
Pubblicazione: (2026)
Emergent Natural Language with Communication Games for Improving Image Captioning Capabilities without Additional Data
di: Dutta, Parag, et al.
Pubblicazione: (2025)
di: Dutta, Parag, et al.
Pubblicazione: (2025)
Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual Fidelity
di: Lee, Hagyeong, et al.
Pubblicazione: (2024)
di: Lee, Hagyeong, et al.
Pubblicazione: (2024)
Pixel-Wise Symbol Spotting via Progressive Points Location for Parsing CAD Images
di: Pang, Junbiao, et al.
Pubblicazione: (2024)
di: Pang, Junbiao, et al.
Pubblicazione: (2024)
Pixel Distillation: A New Knowledge Distillation Scheme for Low-Resolution Image Recognition
di: Guo, Guangyu, et al.
Pubblicazione: (2021)
di: Guo, Guangyu, et al.
Pubblicazione: (2021)
Pixels Versus Priors: Controlling Knowledge Priors in Vision-Language Models through Visual Counterfacts
di: Golovanevsky, Michal, et al.
Pubblicazione: (2025)
di: Golovanevsky, Michal, et al.
Pubblicazione: (2025)
An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels
di: Nguyen, Duy-Kien, et al.
Pubblicazione: (2024)
di: Nguyen, Duy-Kien, et al.
Pubblicazione: (2024)
FigCaps-HF: A Figure-to-Caption Generative Framework and Benchmark with Human Feedback
di: Singh, Ashish, et al.
Pubblicazione: (2023)
di: Singh, Ashish, et al.
Pubblicazione: (2023)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
di: Luo, Jianjie, et al.
Pubblicazione: (2024)
di: Luo, Jianjie, et al.
Pubblicazione: (2024)
Large VLM-based Stylized Sports Captioning
di: Dhar, Sauptik, et al.
Pubblicazione: (2025)
di: Dhar, Sauptik, et al.
Pubblicazione: (2025)
Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?
di: Yan, Xinchen, et al.
Pubblicazione: (2025)
di: Yan, Xinchen, et al.
Pubblicazione: (2025)
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
di: Kim, Si-Woo, et al.
Pubblicazione: (2025)
di: Kim, Si-Woo, et al.
Pubblicazione: (2025)
From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
di: Chen, Yiming, et al.
Pubblicazione: (2025)
di: Chen, Yiming, et al.
Pubblicazione: (2025)
Documenti analoghi
-
From Pixels to Prose: A Large Dataset of Dense Image Captions
di: Singla, Vasu, et al.
Pubblicazione: (2024) -
From Pixels to Prose: Advancing Multi-Modal Language Models for Remote Sensing
di: Sun, Xintian, et al.
Pubblicazione: (2024) -
Good Representation, Better Explanation: Role of Convolutional Neural Networks in Transformer-Based Remote Sensing Image Captioning
di: Das, Swadhin, et al.
Pubblicazione: (2025) -
Low-light Pedestrian Detection in Visible and Infrared Image Feeds: Issues and Challenges
di: Akilan, Thangarajah, et al.
Pubblicazione: (2023) -
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
di: Woo, Byeongju, et al.
Pubblicazione: (2026)