ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Peng, Cihang, Hou, Qiming, Ren, Zhong, Zhou, Kun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Panoptic Captioning: An Equivalence Bridge for Image and Text
by: Lin, Kun-Yu, et al.
Published: (2025)
by: Lin, Kun-Yu, et al.
Published: (2025)
GHOST: Grounded Human Motion Generation with Open Vocabulary Scene-and-Text Contexts
by: Milacski, Zoltán Á., et al.
Published: (2024)
by: Milacski, Zoltán Á., et al.
Published: (2024)
Unified Embedding Alignment for Open-Vocabulary Video Instance Segmentation
by: Fang, Hao, et al.
Published: (2024)
by: Fang, Hao, et al.
Published: (2024)
IFAdapter: Instance Feature Control for Grounded Text-to-Image Generation
by: Wu, Yinwei, et al.
Published: (2024)
by: Wu, Yinwei, et al.
Published: (2024)
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
by: Fan, Tiehan, et al.
Published: (2024)
by: Fan, Tiehan, et al.
Published: (2024)
Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation
by: Lebailly, Tim, et al.
Published: (2025)
by: Lebailly, Tim, et al.
Published: (2025)
Mitigating Open-Vocabulary Caption Hallucinations
by: Ben-Kish, Assaf, et al.
Published: (2023)
by: Ben-Kish, Assaf, et al.
Published: (2023)
Open-Vocabulary Scene Text Recognition via Pseudo-Image Labeling and Margin Loss
by: Ren, Xuhua, et al.
Published: (2024)
by: Ren, Xuhua, et al.
Published: (2024)
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection
by: Bao, Wentao, et al.
Published: (2024)
by: Bao, Wentao, et al.
Published: (2024)
OVI-MAP:Open-Vocabulary Instance-Semantic Mapping
by: Deng, Zilong, et al.
Published: (2026)
by: Deng, Zilong, et al.
Published: (2026)
Progressive Radiance Distillation for Inverse Rendering with Gaussian Splatting
by: Ye, Keyang, et al.
Published: (2024)
by: Ye, Keyang, et al.
Published: (2024)
Open-Vocabulary Object Detection with Meta Prompt Representation and Instance Contrastive Optimization
by: Wang, Zhao, et al.
Published: (2024)
by: Wang, Zhao, et al.
Published: (2024)
Instance Brownian Bridge as Texts for Open-vocabulary Video Instance Segmentation
by: Cheng, Zesen, et al.
Published: (2024)
by: Cheng, Zesen, et al.
Published: (2024)
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
by: Wysoczańska, Monika, et al.
Published: (2025)
by: Wysoczańska, Monika, et al.
Published: (2025)
Improving Text Generation on Images with Synthetic Captions
by: Koh, Jun Young, et al.
Published: (2024)
by: Koh, Jun Young, et al.
Published: (2024)
Towards Real-Time Open-Vocabulary Video Instance Segmentation
by: Yan, Bin, et al.
Published: (2024)
by: Yan, Bin, et al.
Published: (2024)
M$^{2}$Chat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation
by: Chi, Xiaowei, et al.
Published: (2023)
by: Chi, Xiaowei, et al.
Published: (2023)
OpenTrack3D: Towards Accurate and Generalizable Open-Vocabulary 3D Instance Segmentation
by: Zhou, Zhishan, et al.
Published: (2025)
by: Zhou, Zhishan, et al.
Published: (2025)
Open-Vocabulary 3D Semantic Segmentation with Text-to-Image Diffusion Models
by: Zhu, Xiaoyu, et al.
Published: (2024)
by: Zhu, Xiaoyu, et al.
Published: (2024)
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
by: Liu, Zheng, et al.
Published: (2024)
by: Liu, Zheng, et al.
Published: (2024)
Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2024)
by: Kazakos, Evangelos, et al.
Published: (2024)
MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
by: Zhou, Dewei, et al.
Published: (2024)
by: Zhou, Dewei, et al.
Published: (2024)
PhraseStereo: The First Open-Vocabulary Stereo Image Segmentation Dataset
by: Campagnolo, Thomas, et al.
Published: (2025)
by: Campagnolo, Thomas, et al.
Published: (2025)
3D Gaussian Splatting with Deferred Reflection
by: Ye, Keyang, et al.
Published: (2024)
by: Ye, Keyang, et al.
Published: (2024)
CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation
by: Zhu, Wenqi, et al.
Published: (2024)
by: Zhu, Wenqi, et al.
Published: (2024)
Visual Programming for Zero-shot Open-Vocabulary 3D Visual Grounding
by: Yuan, Zhihao, et al.
Published: (2023)
by: Yuan, Zhihao, et al.
Published: (2023)
Filter & Align: Leveraging Human Knowledge to Curate Image-Text Data
by: Zhang, Lei, et al.
Published: (2023)
by: Zhang, Lei, et al.
Published: (2023)
Open-Vocabulary Video Anomaly Detection
by: Wu, Peng, et al.
Published: (2023)
by: Wu, Peng, et al.
Published: (2023)
OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding
by: Huang, Sheng-Yu, et al.
Published: (2026)
by: Huang, Sheng-Yu, et al.
Published: (2026)
Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning
by: Hu, Zhe, et al.
Published: (2025)
by: Hu, Zhe, et al.
Published: (2025)
Affogato: Learning Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale
by: Lee, Junha, et al.
Published: (2025)
by: Lee, Junha, et al.
Published: (2025)
ExpAlign: Expectation-Guided Vision-Language Alignment for Open-Vocabulary Grounding
by: Hu, Junyi, et al.
Published: (2026)
by: Hu, Junyi, et al.
Published: (2026)
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
by: Cui, Tianyu, et al.
Published: (2025)
by: Cui, Tianyu, et al.
Published: (2025)
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
by: Liu, Yanqing, et al.
Published: (2024)
by: Liu, Yanqing, et al.
Published: (2024)
RTGen: Generating Region-Text Pairs for Open-Vocabulary Object Detection
by: Chen, Fangyi, et al.
Published: (2024)
by: Chen, Fangyi, et al.
Published: (2024)
SCORE: Scene Context Matters in Open-Vocabulary Remote Sensing Instance Segmentation
by: Huang, Shiqi, et al.
Published: (2025)
by: Huang, Shiqi, et al.
Published: (2025)
Open-Vocabulary Domain Generalization in Urban-Scene Segmentation
by: Zhao, Dong, et al.
Published: (2026)
by: Zhao, Dong, et al.
Published: (2026)
Zero-Shot Open-Vocabulary Human Motion Grounding with Test-Time Training
by: Zhou, Yunjiao, et al.
Published: (2025)
by: Zhou, Yunjiao, et al.
Published: (2025)
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
by: Wang, Xinran, et al.
Published: (2025)
by: Wang, Xinran, et al.
Published: (2025)
Re$^2$MoGen: Open-Vocabulary Motion Generation via LLM Reasoning and Physics-Aware Refinement
by: Zheng, Jiakun, et al.
Published: (2026)
by: Zheng, Jiakun, et al.
Published: (2026)
Similar Items
-
Panoptic Captioning: An Equivalence Bridge for Image and Text
by: Lin, Kun-Yu, et al.
Published: (2025) -
GHOST: Grounded Human Motion Generation with Open Vocabulary Scene-and-Text Contexts
by: Milacski, Zoltán Á., et al.
Published: (2024) -
Unified Embedding Alignment for Open-Vocabulary Video Instance Segmentation
by: Fang, Hao, et al.
Published: (2024) -
IFAdapter: Instance Feature Control for Grounded Text-to-Image Generation
by: Wu, Yinwei, et al.
Published: (2024) -
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
by: Fan, Tiehan, et al.
Published: (2024)