From Pixels to Prose: A Large Dataset of Dense Image Captions
Fuente:
arXiv
Saved in:
| Main Authors: | Singla, Vasu, Yue, Kaiyu, Paul, Sukriti, Shirkavand, Reza, Jayawardhana, Mayuka, Ganjdanesh, Alireza, Huang, Heng, Bhatele, Abhinav, Somepalli, Gowthami, Goldstein, Tom |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ARGUS: Hallucination and Omission Evaluation in Video-LLMs
by: Rawal, Ruchit, et al.
Published: (2025)
by: Rawal, Ruchit, et al.
Published: (2025)
Efficient Fine-Tuning and Concept Suppression for Pruned Diffusion Models
by: Shirkavand, Reza, et al.
Published: (2024)
by: Shirkavand, Reza, et al.
Published: (2024)
Not All Prompts Are Made Equal: Prompt-based Pruning of Text-to-Image Diffusion Models
by: Ganjdanesh, Alireza, et al.
Published: (2024)
by: Ganjdanesh, Alireza, et al.
Published: (2024)
PUP 3D-GS: Principled Uncertainty Pruning for 3D Gaussian Splatting
by: Hanson, Alex, et al.
Published: (2024)
by: Hanson, Alex, et al.
Published: (2024)
Zero-Shot Vision Encoder Grafting via LLM Surrogates
by: Yue, Kaiyu, et al.
Published: (2025)
by: Yue, Kaiyu, et al.
Published: (2025)
Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs
by: Hans, Abhimanyu, et al.
Published: (2024)
by: Hans, Abhimanyu, et al.
Published: (2024)
CinePile: A Long Video Question Answering Dataset and Benchmark
by: Rawal, Ruchit, et al.
Published: (2024)
by: Rawal, Ruchit, et al.
Published: (2024)
Pixels to Prose: Understanding the art of Image Captioning
by: Singh, Hrishikesh, et al.
Published: (2024)
by: Singh, Hrishikesh, et al.
Published: (2024)
Measuring Style Similarity in Diffusion Models
by: Somepalli, Gowthami, et al.
Published: (2024)
by: Somepalli, Gowthami, et al.
Published: (2024)
FineGRAIN: Evaluating Failure Modes of Text-to-Image Models with Vision Language Model Judges
by: Hayes, Kevin David, et al.
Published: (2025)
by: Hayes, Kevin David, et al.
Published: (2025)
Speedy-Splat: Fast 3D Gaussian Splatting with Sparse Pixels and Sparse Primitives
by: Hanson, Alex, et al.
Published: (2024)
by: Hanson, Alex, et al.
Published: (2024)
Jointly Training and Pruning CNNs via Learnable Agent Guidance and Alignment
by: Ganjdanesh, Alireza, et al.
Published: (2024)
by: Ganjdanesh, Alireza, et al.
Published: (2024)
LMD3: Language Model Data Density Dependence
by: Kirchenbauer, John, et al.
Published: (2024)
by: Kirchenbauer, John, et al.
Published: (2024)
ToMoE: Converting Dense Large Language Models to Mixture-of-Experts through Dynamic Structural Pruning
by: Gao, Shangqian, et al.
Published: (2025)
by: Gao, Shangqian, et al.
Published: (2025)
Image Generation with a Sphere Encoder
by: Yue, Kaiyu, et al.
Published: (2026)
by: Yue, Kaiyu, et al.
Published: (2026)
Mixture of Efficient Diffusion Experts Through Automatic Interval and Sub-Network Selection
by: Ganjdanesh, Alireza, et al.
Published: (2024)
by: Ganjdanesh, Alireza, et al.
Published: (2024)
From Pixels to Prose: Advancing Multi-Modal Language Models for Remote Sensing
by: Sun, Xintian, et al.
Published: (2024)
by: Sun, Xintian, et al.
Published: (2024)
Analytics of Longitudinal System Monitoring Data for Performance Prediction
by: Costello, Ian J., et al.
Published: (2020)
by: Costello, Ian J., et al.
Published: (2020)
Transformers Boost the Performance of Decision Trees on Tabular Data across Sample Sizes
by: Jayawardhana, Mayuka, et al.
Published: (2025)
by: Jayawardhana, Mayuka, et al.
Published: (2025)
UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders
by: Walmer, Matthew, et al.
Published: (2026)
by: Walmer, Matthew, et al.
Published: (2026)
BFMD: A Full-Match Badminton Dense Dataset for Dense Shot Captioning
by: Ding, Ning, et al.
Published: (2026)
by: Ding, Ning, et al.
Published: (2026)
Zero-shot Multivariate Time Series Forecasting Using Tabular Prior Fitted Networks
by: Jayawardhana, Mayuka, et al.
Published: (2026)
by: Jayawardhana, Mayuka, et al.
Published: (2026)
A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions
by: Urbanek, Jack, et al.
Published: (2023)
by: Urbanek, Jack, et al.
Published: (2023)
From Pixels to Posts: Retrieval-Augmented Fashion Captioning and Hashtag Generation
by: Gondal, Moazzam Umer, et al.
Published: (2025)
by: Gondal, Moazzam Umer, et al.
Published: (2025)
Catalog-Native LLM: Speaking Item-ID Dialect with Less Entanglement for Recommendation
by: Shirkavand, Reza, et al.
Published: (2025)
by: Shirkavand, Reza, et al.
Published: (2025)
Characterizing Production GPU Workloads using System-wide Telemetry Data
by: Cankur, Onur, et al.
Published: (2025)
by: Cankur, Onur, et al.
Published: (2025)
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
by: Geiping, Jonas, et al.
Published: (2025)
by: Geiping, Jonas, et al.
Published: (2025)
ParEval-Repo: A Benchmark Suite for Evaluating LLMs with Repository-level HPC Translation Tasks
by: Davis, Joshua H., et al.
Published: (2025)
by: Davis, Joshua H., et al.
Published: (2025)
Streaming Dense Video Captioning
by: Zhou, Xingyi, et al.
Published: (2024)
by: Zhou, Xingyi, et al.
Published: (2024)
HPC-Coder-V2: Studying Code LLMs Across Low-Resource Parallel Languages
by: Chaturvedi, Aman, et al.
Published: (2024)
by: Chaturvedi, Aman, et al.
Published: (2024)
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
by: Li, Ang, et al.
Published: (2025)
by: Li, Ang, et al.
Published: (2025)
From Panels to Prose: Generating Literary Narratives from Comics
by: Sachdeva, Ragav, et al.
Published: (2025)
by: Sachdeva, Ragav, et al.
Published: (2025)
Object Recognition as Next Token Prediction
by: Yue, Kaiyu, et al.
Published: (2023)
by: Yue, Kaiyu, et al.
Published: (2023)
ML-based Modeling to Predict I/O Performance on Different Storage Sub-systems
by: Xu, Yiheng, et al.
Published: (2023)
by: Xu, Yiheng, et al.
Published: (2023)
SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries
by: Mkhallati, Hassan, et al.
Published: (2023)
by: Mkhallati, Hassan, et al.
Published: (2023)
RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning
by: Huang, Tzu-Heng, et al.
Published: (2026)
by: Huang, Tzu-Heng, et al.
Published: (2026)
Speech-Gesture Mapping and Engagement Evaluation in Human Robot Interaction
by: Ghosh, Bishal, et al.
Published: (2018)
by: Ghosh, Bishal, et al.
Published: (2018)
StegaVision: Enhancing Steganography with Attention Mechanism
by: Kumar, Abhinav, et al.
Published: (2024)
by: Kumar, Abhinav, et al.
Published: (2024)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
by: Lim, Junyoung, et al.
Published: (2025)
by: Lim, Junyoung, et al.
Published: (2025)
Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN Training
by: Ranjan, Aditya K., et al.
Published: (2025)
by: Ranjan, Aditya K., et al.
Published: (2025)
Similar Items
-
ARGUS: Hallucination and Omission Evaluation in Video-LLMs
by: Rawal, Ruchit, et al.
Published: (2025) -
Efficient Fine-Tuning and Concept Suppression for Pruned Diffusion Models
by: Shirkavand, Reza, et al.
Published: (2024) -
Not All Prompts Are Made Equal: Prompt-based Pruning of Text-to-Image Diffusion Models
by: Ganjdanesh, Alireza, et al.
Published: (2024) -
PUP 3D-GS: Principled Uncertainty Pruning for 3D Gaussian Splatting
by: Hanson, Alex, et al.
Published: (2024) -
Zero-Shot Vision Encoder Grafting via LLM Surrogates
by: Yue, Kaiyu, et al.
Published: (2025)