Harlequin: Color-driven Generation of Synthetic Data for Referring Expression Comprehension
Fuente:
arXiv
Saved in:
| Main Authors: | Parolari, Luca, Izzo, Elena, Ballan, Lamberto |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
7Bench: a Comprehensive Benchmark for Layout-guided Text-to-image Models
by: Izzo, Elena, et al.
Published: (2025)
by: Izzo, Elena, et al.
Published: (2025)
Benchmarking Layout-Guided Diffusion Models through Unified Semantic-Spatial Evaluation in Closed and Open Settings
by: Parolari, Luca, et al.
Published: (2026)
by: Parolari, Luca, et al.
Published: (2026)
Towards Polyp Counting In Full-Procedure Colonoscopy Videos
by: Parolari, Luca, et al.
Published: (2025)
by: Parolari, Luca, et al.
Published: (2025)
Temporally-Aware Supervised Contrastive Learning for Polyp Counting in Colonoscopy
by: Parolari, Luca, et al.
Published: (2025)
by: Parolari, Luca, et al.
Published: (2025)
Contrastive Learning under Noisy Temporal Self-Supervision for Colonoscopy Videos
by: Parolari, Luca, et al.
Published: (2026)
by: Parolari, Luca, et al.
Published: (2026)
The Solution for the 5th GCAIAC Zero-shot Referring Expression Comprehension Challenge
by: Huang, Longfei, et al.
Published: (2024)
by: Huang, Longfei, et al.
Published: (2024)
Assessing the Visual Enumeration Abilities of Specialized Counting Architectures and Vision-Language Models
by: Hou, Kuinan, et al.
Published: (2025)
by: Hou, Kuinan, et al.
Published: (2025)
Distilling Knowledge for Short-to-Long Term Trajectory Prediction
by: Das, Sourav, et al.
Published: (2023)
by: Das, Sourav, et al.
Published: (2023)
Learning from Synthetic Data for Visual Grounding
by: He, Ruozhen, et al.
Published: (2024)
by: He, Ruozhen, et al.
Published: (2024)
Phrase-Instance Alignment for Generalized Referring Segmentation
by: Nguyen, E-Ro, et al.
Published: (2024)
by: Nguyen, E-Ro, et al.
Published: (2024)
Including Facial Expressions in Contextual Embeddings for Sign Language Generation
by: Viegas, Carla, et al.
Published: (2022)
by: Viegas, Carla, et al.
Published: (2022)
Multiview Progress Prediction of Robot Activities
by: Zoppellari, Elena, et al.
Published: (2026)
by: Zoppellari, Elena, et al.
Published: (2026)
R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model?
by: Zhang, Jingyi, et al.
Published: (2026)
by: Zhang, Jingyi, et al.
Published: (2026)
ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness
by: Liang, Yijun, et al.
Published: (2025)
by: Liang, Yijun, et al.
Published: (2025)
SynRES: Towards Referring Expression Segmentation in the Wild via Synthetic Data
by: Kim, Dong-Hee, et al.
Published: (2025)
by: Kim, Dong-Hee, et al.
Published: (2025)
PromptTA: Prompt-driven Text Adapter for Source-free Domain Generalization
by: Zhang, Haoran, et al.
Published: (2024)
by: Zhang, Haoran, et al.
Published: (2024)
SafaRi:Adaptive Sequence Transformer for Weakly Supervised Referring Expression Segmentation
by: Nag, Sayan, et al.
Published: (2024)
by: Nag, Sayan, et al.
Published: (2024)
Controlled Training Data Generation with Diffusion Models
by: Yeo, Teresa, et al.
Published: (2024)
by: Yeo, Teresa, et al.
Published: (2024)
Data Redaction from Conditional Generative Models
by: Kong, Zhifeng, et al.
Published: (2023)
by: Kong, Zhifeng, et al.
Published: (2023)
Detect Fake with Fake: Leveraging Synthetic Data-driven Representation for Synthetic Image Detection
by: Otake, Hina, et al.
Published: (2024)
by: Otake, Hina, et al.
Published: (2024)
DreamLLM: Synergistic Multimodal Comprehension and Creation
by: Dong, Runpei, et al.
Published: (2023)
by: Dong, Runpei, et al.
Published: (2023)
Data Alignment for Zero-Shot Concept Generation in Dermatology AI
by: Gadgil, Soham, et al.
Published: (2024)
by: Gadgil, Soham, et al.
Published: (2024)
BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation Models
by: Zhang, Ziheng, et al.
Published: (2025)
by: Zhang, Ziheng, et al.
Published: (2025)
BloomVQA: Assessing Hierarchical Multi-modal Comprehension
by: Gong, Yunye, et al.
Published: (2023)
by: Gong, Yunye, et al.
Published: (2023)
Comprehensive Exploration of Synthetic Data Generation: A Survey
by: Bauer, André, et al.
Published: (2024)
by: Bauer, André, et al.
Published: (2024)
A Comprehensive Information-Decomposition Analysis of Large Vision-Language Models
by: Xiu, Lixin, et al.
Published: (2026)
by: Xiu, Lixin, et al.
Published: (2026)
MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
by: Xia, Peng, et al.
Published: (2024)
by: Xia, Peng, et al.
Published: (2024)
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
by: Wang, Xiyao, et al.
Published: (2024)
by: Wang, Xiyao, et al.
Published: (2024)
Rethinking Comprehensive Benchmark for Chart Understanding: A Perspective from Scientific Literature
by: Shen, Lingdong, et al.
Published: (2024)
by: Shen, Lingdong, et al.
Published: (2024)
SciGA: A Comprehensive Dataset for Designing Graphical Abstracts in Academic Papers
by: Kawada, Takuro, et al.
Published: (2025)
by: Kawada, Takuro, et al.
Published: (2025)
CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness
by: Liu, Zhihang, et al.
Published: (2025)
by: Liu, Zhihang, et al.
Published: (2025)
Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies
by: Pathak, Surendra, et al.
Published: (2026)
by: Pathak, Surendra, et al.
Published: (2026)
Referring Expression Generation in Visually Grounded Dialogue with Discourse-aware Comprehension Guiding
by: Willemsen, Bram, et al.
Published: (2024)
by: Willemsen, Bram, et al.
Published: (2024)
CoGen: Learning from Feedback with Coupled Comprehension and Generation
by: Gul, Mustafa Omer, et al.
Published: (2024)
by: Gul, Mustafa Omer, et al.
Published: (2024)
Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models
by: Zheng, Kening, et al.
Published: (2024)
by: Zheng, Kening, et al.
Published: (2024)
Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme
by: Ma, Yan, et al.
Published: (2025)
by: Ma, Yan, et al.
Published: (2025)
EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models
by: Wang, Zekun, et al.
Published: (2025)
by: Wang, Zekun, et al.
Published: (2025)
A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation
by: Wang, Andrew Z., et al.
Published: (2025)
by: Wang, Andrew Z., et al.
Published: (2025)
PersONAL: Towards a Comprehensive Benchmark for Personalized Embodied Agents
by: Ziliotto, Filippo, et al.
Published: (2025)
by: Ziliotto, Filippo, et al.
Published: (2025)
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
by: Deitke, Matt, et al.
Published: (2024)
by: Deitke, Matt, et al.
Published: (2024)
Similar Items
-
7Bench: a Comprehensive Benchmark for Layout-guided Text-to-image Models
by: Izzo, Elena, et al.
Published: (2025) -
Benchmarking Layout-Guided Diffusion Models through Unified Semantic-Spatial Evaluation in Closed and Open Settings
by: Parolari, Luca, et al.
Published: (2026) -
Towards Polyp Counting In Full-Procedure Colonoscopy Videos
by: Parolari, Luca, et al.
Published: (2025) -
Temporally-Aware Supervised Contrastive Learning for Polyp Counting in Colonoscopy
by: Parolari, Luca, et al.
Published: (2025) -
Contrastive Learning under Noisy Temporal Self-Supervision for Colonoscopy Videos
by: Parolari, Luca, et al.
Published: (2026)