Concept-as-Tree: A Controllable Synthetic Data Framework Makes Stronger Personalized VLMs
Fuente:
arXiv
Saved in:
| Main Authors: | An, Ruichuan, Zeng, Kai, Lu, Ming, Yang, Sihan, Zhang, Renrui, Ji, Huitong, Liang, Hao, Zhang, Wentao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
by: An, Ruichuan, et al.
Published: (2025)
by: An, Ruichuan, et al.
Published: (2025)
Uni-Synergy: Bridging Understanding and Generation for Personalized Reasoning via Co-operative Reinforcement Learning
by: Shen, Zijun, et al.
Published: (2026)
by: Shen, Zijun, et al.
Published: (2026)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
by: An, Ruichuan, et al.
Published: (2024)
by: An, Ruichuan, et al.
Published: (2024)
UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens
by: An, Ruichuan, et al.
Published: (2025)
by: An, Ruichuan, et al.
Published: (2025)
Small-Large Collaboration: Training-efficient Concept Personalization for Large VLM using a Meta Personalized Small VLM
by: Yang, Sihan, et al.
Published: (2025)
by: Yang, Sihan, et al.
Published: (2025)
PEARL: Personalized Streaming Video Understanding Model
by: Zheng, Yuanhong, et al.
Published: (2026)
by: Zheng, Yuanhong, et al.
Published: (2026)
MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders
by: Cao, Jiajun, et al.
Published: (2025)
by: Cao, Jiajun, et al.
Published: (2025)
GENIUS: Generative Fluid Intelligence Evaluation Suite
by: An, Ruichuan, et al.
Published: (2026)
by: An, Ruichuan, et al.
Published: (2026)
Exploring Stronger Transformer Representation Learning for Occluded Person Re-Identification
by: Ji, Zhangjian, et al.
Published: (2024)
by: Ji, Zhangjian, et al.
Published: (2024)
VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
by: Zhong, Ming, et al.
Published: (2025)
by: Zhong, Ming, et al.
Published: (2025)
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
by: Zhang, Qizhe, et al.
Published: (2024)
by: Zhang, Qizhe, et al.
Published: (2024)
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
by: Lin, Weifeng, et al.
Published: (2025)
by: Lin, Weifeng, et al.
Published: (2025)
GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
by: Li, Bozhou, et al.
Published: (2025)
by: Li, Bozhou, et al.
Published: (2025)
VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Field Prediction
by: Zhu, Kaixin, et al.
Published: (2026)
by: Zhu, Kaixin, et al.
Published: (2026)
Make Continual Learning Stronger via C-Flat
by: Bian, Ang, et al.
Published: (2024)
by: Bian, Ang, et al.
Published: (2024)
Part-Attention Based Model Make Occluded Person Re-Identification Stronger
by: Chen, Zhihao, et al.
Published: (2024)
by: Chen, Zhihao, et al.
Published: (2024)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
by: Li, Yuying, et al.
Published: (2025)
by: Li, Yuying, et al.
Published: (2025)
Can World Models Benefit VLMs for World Dynamics?
by: Zhang, Kevin, et al.
Published: (2025)
by: Zhang, Kevin, et al.
Published: (2025)
Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks
by: Zeng, Kai, et al.
Published: (2025)
by: Zeng, Kai, et al.
Published: (2025)
Re-ranking Reasoning Context with Tree Search Makes Large Vision-Language Models Stronger
by: Yang, Qi, et al.
Published: (2025)
by: Yang, Qi, et al.
Published: (2025)
An Empirical Study of Validating Synthetic Data for Text-Based Person Retrieval
by: Cao, Min, et al.
Published: (2025)
by: Cao, Min, et al.
Published: (2025)
SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
by: Song, Zijian, et al.
Published: (2025)
by: Song, Zijian, et al.
Published: (2025)
Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning
by: Tan, Zhangyun, et al.
Published: (2026)
by: Tan, Zhangyun, et al.
Published: (2026)
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
by: Qiao, Yuxuan, et al.
Published: (2024)
by: Qiao, Yuxuan, et al.
Published: (2024)
Temporal Regularization Makes Your Video Generator Stronger
by: Chen, Harold Haodong, et al.
Published: (2025)
by: Chen, Harold Haodong, et al.
Published: (2025)
MedConcept: Unsupervised Concept Discovery for Interpretability in Medical VLMs
by: Haque, Md Rakibul, et al.
Published: (2026)
by: Haque, Md Rakibul, et al.
Published: (2026)
Contrastive Masked Autoencoders are Stronger Vision Learners
by: Huang, Zhicheng, et al.
Published: (2022)
by: Huang, Zhicheng, et al.
Published: (2022)
LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
by: Cai, Qifeng, et al.
Published: (2025)
by: Cai, Qifeng, et al.
Published: (2025)
What Makes Synthetic Data Effective in Image Segmentation
by: Zhang, Jinjin, et al.
Published: (2026)
by: Zhang, Jinjin, et al.
Published: (2026)
SpikeGen: Decoupled "Rods and Cones" Visual Representation Processing with Latent Generative Framework
by: Dai, Gaole, et al.
Published: (2025)
by: Dai, Gaole, et al.
Published: (2025)
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
by: Li, Shuo, et al.
Published: (2024)
by: Li, Shuo, et al.
Published: (2024)
vMFCoOp: Towards Equilibrium on a Unified Hyperspherical Manifold for Prompting Biomedical VLMs
by: Shao, Minye, et al.
Published: (2025)
by: Shao, Minye, et al.
Published: (2025)
EVQAScore: A Fine-grained Metric for Video Question Answering Data Quality Evaluation
by: Liang, Hao, et al.
Published: (2024)
by: Liang, Hao, et al.
Published: (2024)
SGDM: Static-Guided Dynamic Module Make Stronger Visual Models
by: Xing, Wenjie, et al.
Published: (2024)
by: Xing, Wenjie, et al.
Published: (2024)
VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
by: Zhang, Nonghai, et al.
Published: (2025)
by: Zhang, Nonghai, et al.
Published: (2025)
ViDA: Homeostatic Visual Domain Adapter for Continual Test Time Adaptation
by: Liu, Jiaming, et al.
Published: (2023)
by: Liu, Jiaming, et al.
Published: (2023)
VLMs Guided Interpretable Decision Making for Autonomous Driving
by: Hu, Xin, et al.
Published: (2025)
by: Hu, Xin, et al.
Published: (2025)
NTO3D: Neural Target Object 3D Reconstruction with Segment Anything
by: Wei, Xiaobao, et al.
Published: (2023)
by: Wei, Xiaobao, et al.
Published: (2023)
Adversarial Concept Distillation for One-Step Diffusion Personalization
by: Yang, Yixiong, et al.
Published: (2025)
by: Yang, Yixiong, et al.
Published: (2025)
Diffusion-based Synthetic Data Generation for Visible-Infrared Person Re-Identification
by: Dai, Wenbo, et al.
Published: (2025)
by: Dai, Wenbo, et al.
Published: (2025)
Similar Items
-
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
by: An, Ruichuan, et al.
Published: (2025) -
Uni-Synergy: Bridging Understanding and Generation for Personalized Reasoning via Co-operative Reinforcement Learning
by: Shen, Zijun, et al.
Published: (2026) -
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
by: An, Ruichuan, et al.
Published: (2024) -
UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens
by: An, Ruichuan, et al.
Published: (2025) -
Small-Large Collaboration: Training-efficient Concept Personalization for Large VLM using a Meta Personalized Small VLM
by: Yang, Sihan, et al.
Published: (2025)