ImageInWords: Unlocking Hyper-Detailed Image Descriptions
Fuente:
arXiv
Guardado en:
| Autores principales: | Garg, Roopal, Burns, Andrea, Ayan, Burcu Karagol, Bitton, Yonatan, Montgomery, Ceslee, Onoe, Yasumasa, Bunner, Andrew, Krishna, Ranjay, Baldridge, Jason, Soricut, Radu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DOCCI: Descriptions of Connected and Contrasting Images
por: Onoe, Yasumasa, et al.
Publicado: (2024)
por: Onoe, Yasumasa, et al.
Publicado: (2024)
TECCI: Tricky Edits of Collected and Curated Images
por: Agrawal, Aishwarya, et al.
Publicado: (2026)
por: Agrawal, Aishwarya, et al.
Publicado: (2026)
Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline
por: Gordon, Brian, et al.
Publicado: (2025)
por: Gordon, Brian, et al.
Publicado: (2025)
Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
por: Gordon, Brian, et al.
Publicado: (2023)
por: Gordon, Brian, et al.
Publicado: (2023)
Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
por: Cho, Jaemin, et al.
Publicado: (2023)
por: Cho, Jaemin, et al.
Publicado: (2023)
Wavelet-Based Image Tokenizer for Vision Transformers
por: Zhu, Zhenhai, et al.
Publicado: (2024)
por: Zhu, Zhenhai, et al.
Publicado: (2024)
Harm Amplification in Text-to-Image Models
por: Hao, Susan, et al.
Publicado: (2024)
por: Hao, Susan, et al.
Publicado: (2024)
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits
por: Yosef, Ron, et al.
Publicado: (2025)
por: Yosef, Ron, et al.
Publicado: (2025)
Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis
por: Ventura, Mor, et al.
Publicado: (2026)
por: Ventura, Mor, et al.
Publicado: (2026)
NL-Eye: Abductive NLI for Images
por: Ventura, Mor, et al.
Publicado: (2024)
por: Ventura, Mor, et al.
Publicado: (2024)
Greedy Growing Enables High-Resolution Pixel-Based Diffusion Models
por: Vasconcelos, Cristina N., et al.
Publicado: (2024)
por: Vasconcelos, Cristina N., et al.
Publicado: (2024)
Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions
por: Pi, Renjie, et al.
Publicado: (2024)
por: Pi, Renjie, et al.
Publicado: (2024)
Agonistic Image Generation: Unsettling the Hegemony of Intention
por: Shaw, Andrew, et al.
Publicado: (2025)
por: Shaw, Andrew, et al.
Publicado: (2025)
Painting with Words: Elevating Detailed Image Captioning with Benchmark and Alignment Learning
por: Ye, Qinghao, et al.
Publicado: (2025)
por: Ye, Qinghao, et al.
Publicado: (2025)
Method of Weighted Words on Cylindric Partitions
por: Barsakçı, Burcu
Publicado: (2025)
por: Barsakçı, Burcu
Publicado: (2025)
Stanford oceanographic expedition 17, Galapagos Islands and vicinity, 22 February-23 March 1968: observations on birds, the Galapagos fur seal, and cetaceans
por: Baldridge, Alan
Publicado: (1968)
por: Baldridge, Alan
Publicado: (1968)
See More Details: Efficient Image Super-Resolution by Experts Mining
por: Zamfir, Eduard, et al.
Publicado: (2024)
por: Zamfir, Eduard, et al.
Publicado: (2024)
CausalLM is not optimal for in-context learning
por: Ding, Nan, et al.
Publicado: (2023)
por: Ding, Nan, et al.
Publicado: (2023)
Omni-SMoLA: Boosting Generalist Multimodal Models with Soft Mixture of Low-rank Experts
por: Wu, Jialin, et al.
Publicado: (2023)
por: Wu, Jialin, et al.
Publicado: (2023)
ParallelPARC: A Scalable Pipeline for Generating Natural-Language Analogies
por: Sultan, Oren, et al.
Publicado: (2024)
por: Sultan, Oren, et al.
Publicado: (2024)
Semantic and Expressive Variation in Image Captions Across Languages
por: Ye, Andre, et al.
Publicado: (2023)
por: Ye, Andre, et al.
Publicado: (2023)
Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation
por: Garcia, Fernando Gabriela, et al.
Publicado: (2025)
por: Garcia, Fernando Gabriela, et al.
Publicado: (2025)
Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage
por: Lee, Saehyung, et al.
Publicado: (2024)
por: Lee, Saehyung, et al.
Publicado: (2024)
Into the Unknown: Generating Geospatial Descriptions for New Environments
por: Paz-Argaman, Tzuf, et al.
Publicado: (2024)
por: Paz-Argaman, Tzuf, et al.
Publicado: (2024)
Heterogeneous graph neural networks for species distribution modeling
por: Harrell, Lauren, et al.
Publicado: (2025)
por: Harrell, Lauren, et al.
Publicado: (2025)
GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
por: Kamath, Amita, et al.
Publicado: (2025)
por: Kamath, Amita, et al.
Publicado: (2025)
From Images to Words: Efficient Cross-Modal Knowledge Distillation to Language Models from Black-box Teachers
por: Sengupta, Ayan, et al.
Publicado: (2026)
por: Sengupta, Ayan, et al.
Publicado: (2026)
The Unmet Promise of Synthetic Training Images: Using Retrieved Real Images Performs Better
por: Geng, Scott, et al.
Publicado: (2024)
por: Geng, Scott, et al.
Publicado: (2024)
Revisit Anything: Visual Place Recognition via Image Segment Retrieval
por: Garg, Kartik, et al.
Publicado: (2024)
por: Garg, Kartik, et al.
Publicado: (2024)
Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis
por: Ramos, Vasco, et al.
Publicado: (2024)
por: Ramos, Vasco, et al.
Publicado: (2024)
MIMIC: Masked Image Modeling with Image Correspondences
por: Marathe, Kalyani, et al.
Publicado: (2023)
por: Marathe, Kalyani, et al.
Publicado: (2023)
PoSh: Using Scene Graphs To Guide LLMs-as-a-Judge For Detailed Image Descriptions
por: Ananthram, Amith, et al.
Publicado: (2025)
por: Ananthram, Amith, et al.
Publicado: (2025)
Gravitational collapse of matter fields in de Sitter spacetimes
por: Garg, Akriti, et al.
Publicado: (2025)
por: Garg, Akriti, et al.
Publicado: (2025)
Hawking radiation from black holes in 2+1 dimensions
por: Garg, Akriti, et al.
Publicado: (2026)
por: Garg, Akriti, et al.
Publicado: (2026)
Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings
por: Wiles, Olivia, et al.
Publicado: (2024)
por: Wiles, Olivia, et al.
Publicado: (2024)
Devil is in the Details: Density Guidance for Detail-Aware Generation with Flow Models
por: Karczewski, Rafał, et al.
Publicado: (2025)
por: Karczewski, Rafał, et al.
Publicado: (2025)
Detecting Stylistic Fingerprints of Large Language Models
por: Bitton, Yehonatan, et al.
Publicado: (2025)
por: Bitton, Yehonatan, et al.
Publicado: (2025)
Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion
por: Fan, Xiang, et al.
Publicado: (2024)
por: Fan, Xiang, et al.
Publicado: (2024)
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
por: Zhang, Tianyi, et al.
Publicado: (2026)
por: Zhang, Tianyi, et al.
Publicado: (2026)
DetailCLIP: Injecting Image Details into CLIP's Feature Space
por: Zhang, Zilun, et al.
Publicado: (2022)
por: Zhang, Zilun, et al.
Publicado: (2022)
Ejemplares similares
-
DOCCI: Descriptions of Connected and Contrasting Images
por: Onoe, Yasumasa, et al.
Publicado: (2024) -
TECCI: Tricky Edits of Collected and Curated Images
por: Agrawal, Aishwarya, et al.
Publicado: (2026) -
Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline
por: Gordon, Brian, et al.
Publicado: (2025) -
Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
por: Gordon, Brian, et al.
Publicado: (2023) -
Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
por: Cho, Jaemin, et al.
Publicado: (2023)