Generating Accurate and Detailed Captions for High-Resolution Images
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Hankyeol, Seo, Gawon, Lee, Kyounggyu, Kim, Dogun, Song, Kyungwoo, Jung, Jiyoung |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Visual Classification using Comparative Descriptors
by: Lee, Hankyeol, et al.
Published: (2024)
by: Lee, Hankyeol, et al.
Published: (2024)
ReflectCAP: Detailed Image Captioning with Reflective Memory
by: Min, Kyungmin, et al.
Published: (2026)
by: Min, Kyungmin, et al.
Published: (2026)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
by: Jung, Mingi, et al.
Published: (2025)
by: Jung, Mingi, et al.
Published: (2025)
Leveraging Textual Compositional Reasoning for Robust Change Captioning
by: Park, Kyu Ri, et al.
Published: (2025)
by: Park, Kyu Ri, et al.
Published: (2025)
P3T: Prototypical Point-level Prompt Tuning with Enhanced Generalization for 3D Vision-Language Models
by: Jung, Geunyoung, et al.
Published: (2026)
by: Jung, Geunyoung, et al.
Published: (2026)
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
by: Kim, Dongwon, et al.
Published: (2026)
by: Kim, Dongwon, et al.
Published: (2026)
Describe Anything: Detailed Localized Image and Video Captioning
by: Lian, Long, et al.
Published: (2025)
by: Lian, Long, et al.
Published: (2025)
good4cir: Generating Detailed Synthetic Captions for Composed Image Retrieval
by: Kolouju, Pranavi, et al.
Published: (2025)
by: Kolouju, Pranavi, et al.
Published: (2025)
Image Tiling for High-Resolution Reasoning: Balancing Local Detail with Global Context
by: de Margerie, Anatole Jacquin, et al.
Published: (2025)
by: de Margerie, Anatole Jacquin, et al.
Published: (2025)
Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?
by: Feng, Mingqian, et al.
Published: (2024)
by: Feng, Mingqian, et al.
Published: (2024)
A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images
by: Lee, Jaeseong, et al.
Published: (2025)
by: Lee, Jaeseong, et al.
Published: (2025)
Accurate and Fast Compressed Video Captioning
by: Shen, Yaojie, et al.
Published: (2023)
by: Shen, Yaojie, et al.
Published: (2023)
Fourier-Guided Attention Upsampling for Image Super-Resolution
by: Choi, Daejune, et al.
Published: (2025)
by: Choi, Daejune, et al.
Published: (2025)
URECA: Unique Region Caption Anything
by: Lim, Sangbeom, et al.
Published: (2025)
by: Lim, Sangbeom, et al.
Published: (2025)
VizECGNet: Visual ECG Image Network for Cardiovascular Diseases Classification with Multi-Modal Training and Knowledge Distillation
by: Nam, Ju-Hyeon, et al.
Published: (2024)
by: Nam, Ju-Hyeon, et al.
Published: (2024)
Language-guided Recursive Spatiotemporal Graph Modeling for Video Summarization
by: Park, Jungin, et al.
Published: (2025)
by: Park, Jungin, et al.
Published: (2025)
Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations
by: Park, Jungin, et al.
Published: (2025)
by: Park, Jungin, et al.
Published: (2025)
V-LynX: Token Interface Alignment for Video+X LLMs
by: Park, Jungin, et al.
Published: (2026)
by: Park, Jungin, et al.
Published: (2026)
SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
by: Dang, Jisheng, et al.
Published: (2025)
by: Dang, Jisheng, et al.
Published: (2025)
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
by: Kim, Si-Woo, et al.
Published: (2025)
by: Kim, Si-Woo, et al.
Published: (2025)
APT: Improving Diffusion Models for High Resolution Image Generation with Adaptive Path Tracing
by: Han, Sangmin, et al.
Published: (2025)
by: Han, Sangmin, et al.
Published: (2025)
Geometrical Properties of Text Token Embeddings for Strong Semantic Binding in Text-to-Image Generation
by: Seo, Hoigi, et al.
Published: (2025)
by: Seo, Hoigi, et al.
Published: (2025)
RaDL: Relation-aware Disentangled Learning for Multi-Instance Text-to-Image Generation
by: Park, Geon, et al.
Published: (2025)
by: Park, Geon, et al.
Published: (2025)
EditSplat: Multi-View Fusion and Attention-Guided Optimization for View-Consistent 3D Scene Editing with 3D Gaussian Splatting
by: Lee, Dong In, et al.
Published: (2024)
by: Lee, Dong In, et al.
Published: (2024)
Self-Cascaded Diffusion Models for Arbitrary-Scale Image Super-Resolution
by: Bang, Junseo, et al.
Published: (2025)
by: Bang, Junseo, et al.
Published: (2025)
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
by: Kim, Taewhan, et al.
Published: (2024)
by: Kim, Taewhan, et al.
Published: (2024)
BeyondScene: Higher-Resolution Human-Centric Scene Generation With Pretrained Diffusion
by: Kim, Gwanghyun, et al.
Published: (2024)
by: Kim, Gwanghyun, et al.
Published: (2024)
TCAN: Animating Human Images with Temporally Consistent Pose Guidance using Diffusion Models
by: Kim, Jeongho, et al.
Published: (2024)
by: Kim, Jeongho, et al.
Published: (2024)
CaptionFool: Universal Image Captioning Model Attacks
by: Parekh, Swapnil
Published: (2026)
by: Parekh, Swapnil
Published: (2026)
Local Representative Token Guided Merging for Text-to-Image Generation
by: Lee, Min-Jeong, et al.
Published: (2025)
by: Lee, Min-Jeong, et al.
Published: (2025)
When Eyes Betray AI: Social Gaze Consistency as a Semantic Cue for AI-Generated Image Detection
by: Kim, Jihyeon, et al.
Published: (2026)
by: Kim, Jihyeon, et al.
Published: (2026)
IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning
by: Lee, Soeun, et al.
Published: (2024)
by: Lee, Soeun, et al.
Published: (2024)
AGIC: Attention-Guided Image Captioning to Improve Caption Relevance
by: Teja, L. D. M. S. Sai, et al.
Published: (2025)
by: Teja, L. D. M. S. Sai, et al.
Published: (2025)
Bridging the Domain Gap: A Simple Domain Matching Method for Reference-based Image Super-Resolution in Remote Sensing
by: Min, Jeongho, et al.
Published: (2024)
by: Min, Jeongho, et al.
Published: (2024)
DyRA: Portable Dynamic Resolution Adjustment Network for Existing Detectors
by: Seo, Daeun, et al.
Published: (2023)
by: Seo, Daeun, et al.
Published: (2023)
XMeCap: Meme Caption Generation with Sub-Image Adaptability
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
Make VLM Recognize Visual Hallucination on Cartoon Character Image with Pose Information
by: Kim, Bumsoo, et al.
Published: (2024)
by: Kim, Bumsoo, et al.
Published: (2024)
Siamese-Driven Optimization for Low-Resolution Image Latent Embedding in Image Captioning
by: Tan, Jing Jie, et al.
Published: (2025)
by: Tan, Jing Jie, et al.
Published: (2025)
Beta Sampling is All You Need: Efficient Image Generation Strategy for Diffusion Models using Stepwise Spectral Analysis
by: Lee, Haeil, et al.
Published: (2024)
by: Lee, Haeil, et al.
Published: (2024)
MAMS: Model-Agnostic Module Selection Framework for Video Captioning
by: Lee, Sangho, et al.
Published: (2025)
by: Lee, Sangho, et al.
Published: (2025)
Similar Items
-
Enhancing Visual Classification using Comparative Descriptors
by: Lee, Hankyeol, et al.
Published: (2024) -
ReflectCAP: Detailed Image Captioning with Reflective Memory
by: Min, Kyungmin, et al.
Published: (2026) -
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
by: Jung, Mingi, et al.
Published: (2025) -
Leveraging Textual Compositional Reasoning for Robust Change Captioning
by: Park, Kyu Ri, et al.
Published: (2025) -
P3T: Prototypical Point-level Prompt Tuning with Enhanced Generalization for 3D Vision-Language Models
by: Jung, Geunyoung, et al.
Published: (2026)