FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity
Fuente:
arXiv
Saved in:
| Main Authors: | Hua, Hang, Liu, Qing, Zhang, Lingzhi, Shi, Jing, Zhang, Zhifei, Wang, Yilin, Zhang, Jianming, Luo, Jiebo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Detecting Human Artifacts from Text-to-Image Models
by: Wang, Kaihong, et al.
Published: (2024)
by: Wang, Kaihong, et al.
Published: (2024)
MIRA: Multimodal Iterative Reasoning Agent for Image Editing
by: Zeng, Ziyun, et al.
Published: (2025)
by: Zeng, Ziyun, et al.
Published: (2025)
PromptFix: You Prompt and We Fix the Photo
by: Yu, Yongsheng, et al.
Published: (2024)
by: Yu, Yongsheng, et al.
Published: (2024)
ZipIR: Latent Pyramid Diffusion Transformer for High-Resolution Image Restoration
by: Yu, Yongsheng, et al.
Published: (2025)
by: Yu, Yongsheng, et al.
Published: (2025)
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
by: Cui, Tianyu, et al.
Published: (2025)
by: Cui, Tianyu, et al.
Published: (2025)
FINEMATCH: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction
by: Hua, Hang, et al.
Published: (2024)
by: Hua, Hang, et al.
Published: (2024)
SAM 2++: Tracking Anything at Any Granularity
by: Zhang, Jiaming, et al.
Published: (2025)
by: Zhang, Jiaming, et al.
Published: (2025)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
by: Tang, Yunlong, et al.
Published: (2025)
by: Tang, Yunlong, et al.
Published: (2025)
Detecting Origin Attribution for Text-to-Image Diffusion Models
by: Xu, Katherine, et al.
Published: (2024)
by: Xu, Katherine, et al.
Published: (2024)
Good Seed Makes a Good Crop: Discovering Secret Seeds in Text-to-Image Diffusion Models
by: Xu, Katherine, et al.
Published: (2024)
by: Xu, Katherine, et al.
Published: (2024)
Count Anything at Any Granularity
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
SwapAnything: Enabling Arbitrary Object Swapping in Personalized Visual Editing
by: Gu, Jing, et al.
Published: (2024)
by: Gu, Jing, et al.
Published: (2024)
MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models
by: Hua, Hang, et al.
Published: (2024)
by: Hua, Hang, et al.
Published: (2024)
STAND: Semantic Anchoring Constraint with Dual-Granularity Disambiguation for Remote Sensing Image Change Captioning
by: Gong, Yanpei, et al.
Published: (2026)
by: Gong, Yanpei, et al.
Published: (2026)
Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity
by: Zhang, Huaxin, et al.
Published: (2024)
by: Zhang, Huaxin, et al.
Published: (2024)
Get In Video: Add Anything You Want to the Video
by: Zhuang, Shaobin, et al.
Published: (2025)
by: Zhuang, Shaobin, et al.
Published: (2025)
Aesthetic Image Captioning with Saliency Enhanced MLLMs
by: Tao, Yilin, et al.
Published: (2025)
by: Tao, Yilin, et al.
Published: (2025)
Structure-Guided Image Completion with Image-level and Object-level Semantic Discriminators
by: Zheng, Haitian, et al.
Published: (2022)
by: Zheng, Haitian, et al.
Published: (2022)
Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
by: Lin, Weifeng, et al.
Published: (2024)
by: Lin, Weifeng, et al.
Published: (2024)
UniRestorer: Universal Image Restoration via Adaptively Estimating Image Degradation at Proper Granularity
by: Lin, Jingbo, et al.
Published: (2024)
by: Lin, Jingbo, et al.
Published: (2024)
AIComposer: Any Style and Content Image Composition via Feature Integration
by: Li, Haowen, et al.
Published: (2025)
by: Li, Haowen, et al.
Published: (2025)
Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models
by: Li, Senmao, et al.
Published: (2024)
by: Li, Senmao, et al.
Published: (2024)
V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction Tuning
by: Hua, Hang, et al.
Published: (2024)
by: Hua, Hang, et al.
Published: (2024)
IMPRINT: Generative Object Compositing by Learning Identity-Preserving Representation
by: Song, Yizhi, et al.
Published: (2024)
by: Song, Yizhi, et al.
Published: (2024)
UniHuman: A Unified Model for Editing Human Images in the Wild
by: Li, Nannan, et al.
Published: (2023)
by: Li, Nannan, et al.
Published: (2023)
MaskFocus: Focusing Policy Optimization on Critical Steps for Masked Image Generation
by: Zhang, Guohui, et al.
Published: (2025)
by: Zhang, Guohui, et al.
Published: (2025)
MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models
by: Hua, Hang, et al.
Published: (2025)
by: Hua, Hang, et al.
Published: (2025)
NeIn: Telling What You Don't Want
by: Bui, Nhat-Tan, et al.
Published: (2024)
by: Bui, Nhat-Tan, et al.
Published: (2024)
GaussianStyle: Gaussian Head Avatar via StyleGAN
by: Liu, Pinxin, et al.
Published: (2024)
by: Liu, Pinxin, et al.
Published: (2024)
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
by: Zeng, Ziyun, et al.
Published: (2026)
by: Zeng, Ziyun, et al.
Published: (2026)
Multi-Granularity Hand Action Detection
by: Zhe, Ting, et al.
Published: (2023)
by: Zhe, Ting, et al.
Published: (2023)
Track Any Anomalous Object: A Granular Video Anomaly Detection Pipeline
by: Huang, Yuzhi, et al.
Published: (2025)
by: Huang, Yuzhi, et al.
Published: (2025)
SeFi-CD: A Semantic First Change Detection Paradigm That Can Detect Any Change You Want
by: Zhao, Ling, et al.
Published: (2024)
by: Zhao, Ling, et al.
Published: (2024)
SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning
by: Zhang, Lin, et al.
Published: (2025)
by: Zhang, Lin, et al.
Published: (2025)
Do You Guys Want to Dance: Zero-Shot Compositional Human Dance Generation with Multiple Persons
by: Xu, Zhe, et al.
Published: (2024)
by: Xu, Zhe, et al.
Published: (2024)
A Lightweight Sparse Focus Transformer for Remote Sensing Image Change Captioning
by: Sun, Dongwei, et al.
Published: (2024)
by: Sun, Dongwei, et al.
Published: (2024)
Top-Down Semantic Refinement for Image Captioning
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
Thinking Outside the BBox: Unconstrained Generative Object Compositing
by: Tarrés, Gemma Canet, et al.
Published: (2024)
by: Tarrés, Gemma Canet, et al.
Published: (2024)
Similar Items
-
Detecting Human Artifacts from Text-to-Image Models
by: Wang, Kaihong, et al.
Published: (2024) -
MIRA: Multimodal Iterative Reasoning Agent for Image Editing
by: Zeng, Ziyun, et al.
Published: (2025) -
PromptFix: You Prompt and We Fix the Photo
by: Yu, Yongsheng, et al.
Published: (2024) -
ZipIR: Latent Pyramid Diffusion Transformer for High-Resolution Image Restoration
by: Yu, Yongsheng, et al.
Published: (2025) -
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
by: Cui, Tianyu, et al.
Published: (2025)