SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Zheng, Liang, Hao, Li, Bozhou, Xiong, Wentao, Chen, Chong, He, Conghui, Zhang, Wentao, Cui, Bin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Synth-Empathy: Towards High-Quality Synthetic Empathy Data
di: Liang, Hao, et al.
Pubblicazione: (2024)
di: Liang, Hao, et al.
Pubblicazione: (2024)
ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models
di: Li, Bozhou, et al.
Pubblicazione: (2025)
di: Li, Bozhou, et al.
Pubblicazione: (2025)
Are Bigger Encoders Always Better in Vision Large Models?
di: Li, Bozhou, et al.
Pubblicazione: (2024)
di: Li, Bozhou, et al.
Pubblicazione: (2024)
Gradual Learning: Optimizing Fine-Tuning with Partially Mastered Knowledge in Large Language Models
di: Li, Bozhou, et al.
Pubblicazione: (2024)
di: Li, Bozhou, et al.
Pubblicazione: (2024)
FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding
di: Liu, Zheng, et al.
Pubblicazione: (2025)
di: Liu, Zheng, et al.
Pubblicazione: (2025)
BRACE: A Benchmark for Robust Audio Caption Quality Evaluation
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
KeyVideoLLM: Towards Large-scale Video Keyframe Selection
di: Liang, Hao, et al.
Pubblicazione: (2024)
di: Liang, Hao, et al.
Pubblicazione: (2024)
Efficient-Empathy: Towards Efficient and Effective Selection of Empathy Data
di: Sun, Linzhuang, et al.
Pubblicazione: (2024)
di: Sun, Linzhuang, et al.
Pubblicazione: (2024)
Heterogeneous Adaptive Policy Optimization: Tailoring Optimization to Every Token's Nature
di: Liu, Zheng, et al.
Pubblicazione: (2025)
di: Liu, Zheng, et al.
Pubblicazione: (2025)
A Survey of Multimodal Large Language Model from A Data-centric Perspective
di: Bai, Tianyi, et al.
Pubblicazione: (2024)
di: Bai, Tianyi, et al.
Pubblicazione: (2024)
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models
di: Niu, Junbo, et al.
Pubblicazione: (2025)
di: Niu, Junbo, et al.
Pubblicazione: (2025)
Text2VectorSQL: Towards a Unified Interface for Vector Search and SQL Queries
di: Wang, Zhengren, et al.
Pubblicazione: (2025)
di: Wang, Zhengren, et al.
Pubblicazione: (2025)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
di: Li, Yuying, et al.
Pubblicazione: (2025)
di: Li, Yuying, et al.
Pubblicazione: (2025)
TPC-ViT: Token Propagation Controller for Efficient Vision Transformer
di: Zhu, Wentao
Pubblicazione: (2024)
di: Zhu, Wentao
Pubblicazione: (2024)
AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
di: Wang, Zhengren, et al.
Pubblicazione: (2026)
di: Wang, Zhengren, et al.
Pubblicazione: (2026)
BEATS: Optimizing LLM Mathematical Capabilities with BackVerify and Adaptive Disambiguate based Efficient Tree Search
di: Sun, Linzhuang, et al.
Pubblicazione: (2024)
di: Sun, Linzhuang, et al.
Pubblicazione: (2024)
DOCR-Inspector: Fine-Grained and Automated Evaluation of Document Parsing with VLM
di: Zhang, Qintong, et al.
Pubblicazione: (2025)
di: Zhang, Qintong, et al.
Pubblicazione: (2025)
EVQAScore: A Fine-grained Metric for Video Question Answering Data Quality Evaluation
di: Liang, Hao, et al.
Pubblicazione: (2024)
di: Liang, Hao, et al.
Pubblicazione: (2024)
PilotRL: Training Language Model Agents via Global Planning-Guided Progressive Reinforcement Learning
di: Lu, Keer, et al.
Pubblicazione: (2025)
di: Lu, Keer, et al.
Pubblicazione: (2025)
Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings
di: Sharifzadeh, Sahand, et al.
Pubblicazione: (2024)
di: Sharifzadeh, Sahand, et al.
Pubblicazione: (2024)
MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding
di: Dong, Hejun, et al.
Pubblicazione: (2026)
di: Dong, Hejun, et al.
Pubblicazione: (2026)
Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction
di: Zhang, Qintong, et al.
Pubblicazione: (2024)
di: Zhang, Qintong, et al.
Pubblicazione: (2024)
Data Proportion Detection for Optimized Data Management for Large Language Models
di: Liang, Hao, et al.
Pubblicazione: (2024)
di: Liang, Hao, et al.
Pubblicazione: (2024)
Parrot Captions Teach CLIP to Spot Text
di: Lin, Yiqi, et al.
Pubblicazione: (2023)
di: Lin, Yiqi, et al.
Pubblicazione: (2023)
UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
di: Xue, Zhucun, et al.
Pubblicazione: (2025)
di: Xue, Zhucun, et al.
Pubblicazione: (2025)
Harnessing Vision-Language Pretrained Models with Temporal-Aware Adaptation for Referring Video Object Segmentation
di: Zhou, Zikun, et al.
Pubblicazione: (2024)
di: Zhou, Zikun, et al.
Pubblicazione: (2024)
DARO: Difficulty-Aware Reweighting Policy Optimization
di: Zhou, Jingyu, et al.
Pubblicazione: (2025)
di: Zhou, Jingyu, et al.
Pubblicazione: (2025)
Characterization of Political Polarized Users Attacked by Language Toxicity on Twitter
di: Xu, Wentao
Pubblicazione: (2024)
di: Xu, Wentao
Pubblicazione: (2024)
pQuant: Towards Effective Low-Bit Language Models via Decoupled Linear Quantization-Aware Training
di: Zhang, Wenzheng, et al.
Pubblicazione: (2026)
di: Zhang, Wenzheng, et al.
Pubblicazione: (2026)
Text2SQL-Flow: A Robust SQL-Aware Data Augmentation Framework for Text-to-SQL
di: Cai, Qifeng, et al.
Pubblicazione: (2025)
di: Cai, Qifeng, et al.
Pubblicazione: (2025)
Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search
di: Shi, Wentao, et al.
Pubblicazione: (2025)
di: Shi, Wentao, et al.
Pubblicazione: (2025)
BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
di: Lozano, Alejandro, et al.
Pubblicazione: (2025)
di: Lozano, Alejandro, et al.
Pubblicazione: (2025)
FastVLM: Efficient Vision Encoding for Vision Language Models
di: Vasu, Pavan Kumar Anasosalu, et al.
Pubblicazione: (2024)
di: Vasu, Pavan Kumar Anasosalu, et al.
Pubblicazione: (2024)
GeoSynth: Contextually-Aware High-Resolution Satellite Image Synthesis
di: Sastry, Srikumar, et al.
Pubblicazione: (2024)
di: Sastry, Srikumar, et al.
Pubblicazione: (2024)
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models
di: Huang, Mingzhe, et al.
Pubblicazione: (2026)
di: Huang, Mingzhe, et al.
Pubblicazione: (2026)
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration
di: Bai, Tianyi, et al.
Pubblicazione: (2024)
di: Bai, Tianyi, et al.
Pubblicazione: (2024)
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
di: Peng, Cihang, et al.
Pubblicazione: (2025)
di: Peng, Cihang, et al.
Pubblicazione: (2025)
Aesthetic Assessment of Chinese Handwritings Based on Vision Language Models
di: Zheng, Chen, et al.
Pubblicazione: (2026)
di: Zheng, Chen, et al.
Pubblicazione: (2026)
"It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language Models
di: Garg, Kapil, et al.
Pubblicazione: (2025)
di: Garg, Kapil, et al.
Pubblicazione: (2025)
SynthForge: Synthesizing High-Quality Face Dataset with Controllable 3D Generative Models
di: Rawat, Abhay, et al.
Pubblicazione: (2024)
di: Rawat, Abhay, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Synth-Empathy: Towards High-Quality Synthetic Empathy Data
di: Liang, Hao, et al.
Pubblicazione: (2024) -
ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models
di: Li, Bozhou, et al.
Pubblicazione: (2025) -
Are Bigger Encoders Always Better in Vision Large Models?
di: Li, Bozhou, et al.
Pubblicazione: (2024) -
Gradual Learning: Optimizing Fine-Tuning with Partially Mastered Knowledge in Large Language Models
di: Li, Bozhou, et al.
Pubblicazione: (2024) -
FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding
di: Liu, Zheng, et al.
Pubblicazione: (2025)