ControlCap: Controllable Region-level Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Yuzhong, Liu, Yue, Guo, Zonghao, Wu, Weijia, Gong, Chen, Wan, Fang, Ye, Qixiang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolution
by: Zhao, Yuzhong, et al.
Published: (2024)
by: Zhao, Yuzhong, et al.
Published: (2024)
Thinking with Images via Self-Calling Agent
by: Yang, Wenxi, et al.
Published: (2025)
by: Yang, Wenxi, et al.
Published: (2025)
CC-Diff: Enhancing Contextual Coherence in Remote Sensing Image Synthesis
by: Zhang, Mu, et al.
Published: (2024)
by: Zhang, Mu, et al.
Published: (2024)
Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
by: Liu, Feng, et al.
Published: (2024)
by: Liu, Feng, et al.
Published: (2024)
DiffuMask: Synthesizing Images with Pixel-level Annotations for Semantic Segmentation Using Diffusion Models
by: Wu, Weijia, et al.
Published: (2023)
by: Wu, Weijia, et al.
Published: (2023)
Global2Local: A Joint-Hierarchical Attention for Video Captioning
by: Dai, Chengpeng, et al.
Published: (2022)
by: Dai, Chengpeng, et al.
Published: (2022)
GLaVE-Cap: Global-Local Aligned Video Captioning with Vision Expert Integration
by: Xu, Wan, et al.
Published: (2025)
by: Xu, Wan, et al.
Published: (2025)
Evaluation of Text-to-Video Generation Models: A Dynamics Perspective
by: Liao, Mingxiang, et al.
Published: (2024)
by: Liao, Mingxiang, et al.
Published: (2024)
FingerCap: Fine-grained Finger-level Hand Motion Captioning
by: Shen, Xin, et al.
Published: (2025)
by: Shen, Xin, et al.
Published: (2025)
AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning
by: Ren, Yiming, et al.
Published: (2025)
by: Ren, Yiming, et al.
Published: (2025)
VMamba: Visual State Space Model
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
SnapCap: Efficient Snapshot Compressive Video Captioning
by: Sun, Jianqiao, et al.
Published: (2024)
by: Sun, Jianqiao, et al.
Published: (2024)
MeaCap: Memory-Augmented Zero-shot Image Captioning
by: Zeng, Zequn, et al.
Published: (2024)
by: Zeng, Zequn, et al.
Published: (2024)
IF-VidCap: Can Video Caption Models Follow Instructions?
by: Li, Shihao, et al.
Published: (2025)
by: Li, Shihao, et al.
Published: (2025)
TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
by: Jin, Bu, et al.
Published: (2024)
by: Jin, Bu, et al.
Published: (2024)
CapS-Adapter: Caption-based MultiModal Adapter in Zero-Shot Classification
by: Wang, Qijie, et al.
Published: (2024)
by: Wang, Qijie, et al.
Published: (2024)
Video ReCap: Recursive Captioning of Hour-Long Videos
by: Islam, Md Mohaiminul, et al.
Published: (2024)
by: Islam, Md Mohaiminul, et al.
Published: (2024)
SuperCap: Multi-resolution Superpixel-based Image Captioning
by: Senior, Henry, et al.
Published: (2025)
by: Senior, Henry, et al.
Published: (2025)
CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning
by: Saito, Kuniaki, et al.
Published: (2025)
by: Saito, Kuniaki, et al.
Published: (2025)
CycleCap: Improving VLMs Captioning Performance via Self-Supervised Cycle Consistency Fine-Tuning
by: Krestenitis, Marios, et al.
Published: (2026)
by: Krestenitis, Marios, et al.
Published: (2026)
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
by: Cheng, Kanzhi, et al.
Published: (2025)
by: Cheng, Kanzhi, et al.
Published: (2025)
VoCap: Video Object Captioning and Segmentation from Any Prompt
by: Uijlings, Jasper, et al.
Published: (2025)
by: Uijlings, Jasper, et al.
Published: (2025)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
by: Li, Yuying, et al.
Published: (2025)
by: Li, Yuying, et al.
Published: (2025)
XMeCap: Meme Caption Generation with Sub-Image Adaptability
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
CodecCap: High-Fidelity Codec-Inspired Residual Modeling for Dense Video Captioning
by: Lin, Zihan, et al.
Published: (2026)
by: Lin, Zihan, et al.
Published: (2026)
CapHDR2IR: Caption-Driven Transfer from Visible Light to Infrared Domain
by: Peng, Jingchao, et al.
Published: (2024)
by: Peng, Jingchao, et al.
Published: (2024)
OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward
by: Zhong, Chunlin, et al.
Published: (2025)
by: Zhong, Chunlin, et al.
Published: (2025)
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
by: Meng, Desen, et al.
Published: (2025)
by: Meng, Desen, et al.
Published: (2025)
ProCap: Projection-Aware Captioning for Spatial Augmented Reality
by: Cao, Zimo, et al.
Published: (2026)
by: Cao, Zimo, et al.
Published: (2026)
AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
by: Chai, Wenhao, et al.
Published: (2024)
by: Chai, Wenhao, et al.
Published: (2024)
HyperCap: Hyperspectral Land Cover Captioning Dataset for Vision Language Models
by: Das, Aryan, et al.
Published: (2025)
by: Das, Aryan, et al.
Published: (2025)
Delving Deep into Semantic Relation Distillation
by: Yan, Zhaoyi, et al.
Published: (2025)
by: Yan, Zhaoyi, et al.
Published: (2025)
COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation
by: Deng, Xueqing, et al.
Published: (2025)
by: Deng, Xueqing, et al.
Published: (2025)
Scene Graph-guided SegCaptioning Transformer with Fine-grained Alignment for Controllable Video Segmentation and Captioning
by: Zhang, Xu, et al.
Published: (2026)
by: Zhang, Xu, et al.
Published: (2026)
BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning
by: Ye, Shaokai, et al.
Published: (2026)
by: Ye, Shaokai, et al.
Published: (2026)
Expandable Residual Approximation for Knowledge Distillation
by: Yan, Zhaoyi, et al.
Published: (2025)
by: Yan, Zhaoyi, et al.
Published: (2025)
OrthCaps: An Orthogonal CapsNet with Sparse Attention Routing and Pruning
by: Geng, Xinyu, et al.
Published: (2024)
by: Geng, Xinyu, et al.
Published: (2024)
ReCap: Event-Aware Image Captioning with Article Retrieval and Semantic Gaussian Normalization
by: Nguyen, Thinh-Phuc, et al.
Published: (2025)
by: Nguyen, Thinh-Phuc, et al.
Published: (2025)
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
by: Xing, Long, et al.
Published: (2025)
by: Xing, Long, et al.
Published: (2025)
Similar Items
-
DynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolution
by: Zhao, Yuzhong, et al.
Published: (2024) -
Thinking with Images via Self-Calling Agent
by: Yang, Wenxi, et al.
Published: (2025) -
CC-Diff: Enhancing Contextual Coherence in Remote Sensing Image Synthesis
by: Zhang, Mu, et al.
Published: (2024) -
Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
by: Liu, Feng, et al.
Published: (2024) -
DiffuMask: Synthesizing Images with Pixel-level Annotations for Semantic Segmentation Using Diffusion Models
by: Wu, Weijia, et al.
Published: (2023)