Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Du, Zilin, Li, Haoxin, Yu, Jianfei, Li, Boyang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Difficulty of Learning a Meta-network for Training Data Selection
by: Du, Zilin, et al.
Published: (2026)
by: Du, Zilin, et al.
Published: (2026)
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
by: Woo, Byeongju, et al.
Published: (2026)
by: Woo, Byeongju, et al.
Published: (2026)
Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision
by: Cao, Shengcao, et al.
Published: (2024)
by: Cao, Shengcao, et al.
Published: (2024)
Pre-Training Meta-Rule Selection Policy for Visual Generative Abductive Learning
by: Jin, Yu, et al.
Published: (2025)
by: Jin, Yu, et al.
Published: (2025)
Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
by: Li, Jialuo, et al.
Published: (2025)
by: Li, Jialuo, et al.
Published: (2025)
Oscillation-Reduced MXFP4 Training for Vision Transformers
by: Chen, Yuxiang, et al.
Published: (2025)
by: Chen, Yuxiang, et al.
Published: (2025)
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
by: Wang, Shihao, et al.
Published: (2026)
by: Wang, Shihao, et al.
Published: (2026)
HoloLLM: Multisensory Foundation Model for Language-Grounded Human Sensing and Reasoning
by: Zhou, Chuhao, et al.
Published: (2025)
by: Zhou, Chuhao, et al.
Published: (2025)
Learning to Learn from APIs: Black-Box Data-Free Meta-Learning
by: Hu, Zixuan, et al.
Published: (2023)
by: Hu, Zixuan, et al.
Published: (2023)
Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding
by: Wang, Wenkai, et al.
Published: (2026)
by: Wang, Wenkai, et al.
Published: (2026)
Visual Generation Without Guidance
by: Chen, Huayu, et al.
Published: (2025)
by: Chen, Huayu, et al.
Published: (2025)
Toward an Artificial General Teacher: Procedural Geometry Data Generation and Visual Grounding with Vision-Language Models
by: Nguyen-Truong, Hai, et al.
Published: (2026)
by: Nguyen-Truong, Hai, et al.
Published: (2026)
Shape-Guided Diffusion with Inside-Outside Attention
by: Park, Dong Huk, et al.
Published: (2022)
by: Park, Dong Huk, et al.
Published: (2022)
VG3T: Visual Geometry Grounded Gaussian Transformer
by: Kim, Junho, et al.
Published: (2025)
by: Kim, Junho, et al.
Published: (2025)
Visual Test-time Scaling for GUI Agent Grounding
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning
by: Li, Chenghao, et al.
Published: (2026)
by: Li, Chenghao, et al.
Published: (2026)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
by: Gao, Ziqi, et al.
Published: (2024)
by: Gao, Ziqi, et al.
Published: (2024)
Closing the Gap in Human Behavior Analysis: A Pipeline for Synthesizing Trimodal Data
by: Stippel, Christian, et al.
Published: (2024)
by: Stippel, Christian, et al.
Published: (2024)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
by: Lai, Zhengfeng, et al.
Published: (2023)
by: Lai, Zhengfeng, et al.
Published: (2023)
Aligning Logits Generatively for Principled Black-Box Knowledge Distillation
by: Ma, Jing, et al.
Published: (2022)
by: Ma, Jing, et al.
Published: (2022)
EfficientTrain++: Generalized Curriculum Learning for Efficient Visual Backbone Training
by: Wang, Yulin, et al.
Published: (2024)
by: Wang, Yulin, et al.
Published: (2024)
SparVAR: Exploring Sparsity in Visual AutoRegressive Modeling for Training-Free Acceleration
by: Li, Zekun, et al.
Published: (2026)
by: Li, Zekun, et al.
Published: (2026)
Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction Tuning
by: Safaei, Bardia, et al.
Published: (2025)
by: Safaei, Bardia, et al.
Published: (2025)
Turn That Frown Upside Down: FaceID Customization via Cross-Training Data
by: Wang, Shuhe, et al.
Published: (2025)
by: Wang, Shuhe, et al.
Published: (2025)
Towards Visual Text Grounding of Multimodal Large Language Model
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
CellCLIP -- Learning Perturbation Effects in Cell Painting via Text-Guided Contrastive Learning
by: Lu, Mingyu, et al.
Published: (2025)
by: Lu, Mingyu, et al.
Published: (2025)
Large-image Object Detection for Fine-grained Recognition of Punches Patterns in Medieval Panel Painting
by: Bruegger, Josh, et al.
Published: (2025)
by: Bruegger, Josh, et al.
Published: (2025)
The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding
by: Luo, Jiayun, et al.
Published: (2024)
by: Luo, Jiayun, et al.
Published: (2024)
Prompting the Unseen: Detecting Hidden Backdoors in Black-Box Models
by: Huang, Zi-Xuan, et al.
Published: (2024)
by: Huang, Zi-Xuan, et al.
Published: (2024)
Diversify, Don't Fine-Tune: Scaling Up Visual Recognition Training with Synthetic Images
by: Yu, Zhuoran, et al.
Published: (2023)
by: Yu, Zhuoran, et al.
Published: (2023)
ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
by: Kang, Weitai, et al.
Published: (2025)
by: Kang, Weitai, et al.
Published: (2025)
By My Eyes: Grounding Multimodal Large Language Models with Sensor Data via Visual Prompting
by: Yoon, Hyungjun, et al.
Published: (2024)
by: Yoon, Hyungjun, et al.
Published: (2024)
PointSSC: A Cooperative Vehicle-Infrastructure Point Cloud Benchmark for Semantic Scene Completion
by: Yan, Yuxiang, et al.
Published: (2023)
by: Yan, Yuxiang, et al.
Published: (2023)
Adversarial Exploitation of Data Diversity Improves Visual Localization
by: Li, Sihang, et al.
Published: (2024)
by: Li, Sihang, et al.
Published: (2024)
Box6D : Zero-shot Category-level 6D Pose Estimation of Warehouse Boxes
by: Ma, Yintao, et al.
Published: (2025)
by: Ma, Yintao, et al.
Published: (2025)
Cortex-Grounded Diffusion Models for Brain Image Generation
by: Bongratz, Fabian, et al.
Published: (2026)
by: Bongratz, Fabian, et al.
Published: (2026)
Weblica: Scalable and Reproducible Training Environments for Visual Web Agents
by: Kar, Oğuzhan Fatih, et al.
Published: (2026)
by: Kar, Oğuzhan Fatih, et al.
Published: (2026)
When Dynamic Data Selection Meets Data Augmentation
by: Yang, Suorong, et al.
Published: (2025)
by: Yang, Suorong, et al.
Published: (2025)
Grounding Video Models to Actions through Goal Conditioned Exploration
by: Luo, Yunhao, et al.
Published: (2024)
by: Luo, Yunhao, et al.
Published: (2024)
VideoGEM: Training-free Action Grounding in Videos
by: Vogel, Felix, et al.
Published: (2025)
by: Vogel, Felix, et al.
Published: (2025)
Similar Items
-
On the Difficulty of Learning a Meta-network for Training Data Selection
by: Du, Zilin, et al.
Published: (2026) -
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
by: Woo, Byeongju, et al.
Published: (2026) -
Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision
by: Cao, Shengcao, et al.
Published: (2024) -
Pre-Training Meta-Rule Selection Policy for Visual Generative Abductive Learning
by: Jin, Yu, et al.
Published: (2025) -
Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
by: Li, Jialuo, et al.
Published: (2025)