DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Hengyu, Gu, Tiancheng, Qin, Bin, Wu, Lan, Wu, Yuling, Tan, Shuo, Sun, Zelong, Wang, Jun, Wu, Nan, An, Xiang, Cai, Weidong, Feng, Ziyong, Yang, Kaicheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards
by: Wang, Jun, et al.
Published: (2026)
by: Wang, Jun, et al.
Published: (2026)
RWKV-CLIP: A Robust Vision-Language Representation Learner
by: Gu, Tiancheng, et al.
Published: (2024)
by: Gu, Tiancheng, et al.
Published: (2024)
ORID: Organ-Regional Information Driven Framework for Radiology Report Generation
by: Gu, Tiancheng, et al.
Published: (2024)
by: Gu, Tiancheng, et al.
Published: (2024)
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
by: Yang, Kaicheng, et al.
Published: (2024)
by: Yang, Kaicheng, et al.
Published: (2024)
LaPA: Latent Prompt Assist Model For Medical Visual Question Answering
by: Gu, Tiancheng, et al.
Published: (2024)
by: Gu, Tiancheng, et al.
Published: (2024)
RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm
by: Gu, Tiancheng, et al.
Published: (2025)
by: Gu, Tiancheng, et al.
Published: (2025)
UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
by: Gu, Tiancheng, et al.
Published: (2025)
by: Gu, Tiancheng, et al.
Published: (2025)
Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs
by: Gu, Tiancheng, et al.
Published: (2025)
by: Gu, Tiancheng, et al.
Published: (2025)
1st Place Solution to the 1st SkatingVerse Challenge
by: Sun, Tao, et al.
Published: (2024)
by: Sun, Tao, et al.
Published: (2024)
UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models
by: Chen, Lan, et al.
Published: (2025)
by: Chen, Lan, et al.
Published: (2025)
Pre-training of Lightweight Vision Transformers on Small Datasets with Minimally Scaled Images
by: Tan, Jen Hong
Published: (2024)
by: Tan, Jen Hong
Published: (2024)
SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge
by: Li, Chuanhao, et al.
Published: (2024)
by: Li, Chuanhao, et al.
Published: (2024)
PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training
by: Xie, Yin, et al.
Published: (2025)
by: Xie, Yin, et al.
Published: (2025)
Efficient, Validation-Free Intrinsic Quality Estimation for Large-Scale Face Recognition Datasets
by: Chen, Zhichao, et al.
Published: (2026)
by: Chen, Zhichao, et al.
Published: (2026)
BUS:Efficient and Effective Vision-language Pre-training with Bottom-Up Patch Summarization
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training
by: Bawazir, Ameera, et al.
Published: (2024)
by: Bawazir, Ameera, et al.
Published: (2024)
Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-trained BERT
by: Dai, Dongyang, et al.
Published: (2025)
by: Dai, Dongyang, et al.
Published: (2025)
Multi-label Cluster Discrimination for Visual Representation Learning
by: An, Xiang, et al.
Published: (2024)
by: An, Xiang, et al.
Published: (2024)
VIP: Vision Instructed Pre-training for Robotic Manipulation
by: Li, Zhuoling, et al.
Published: (2024)
by: Li, Zhuoling, et al.
Published: (2024)
Continual Retinal Vision-Language Pre-training upon Incremental Imaging Modalities
by: Yao, Yuang, et al.
Published: (2025)
by: Yao, Yuang, et al.
Published: (2025)
CoMP: Continual Multimodal Pre-training for Vision Foundation Models
by: Chen, Yitong, et al.
Published: (2025)
by: Chen, Yitong, et al.
Published: (2025)
Pre-Big-Bang Cosmology Cannot Explain NANOGrav 15-year Signal
by: Tan, Qin, et al.
Published: (2024)
by: Tan, Qin, et al.
Published: (2024)
Vehicle-centric Perception via Multimodal Structured Pre-training
by: Wu, Wentao, et al.
Published: (2025)
by: Wu, Wentao, et al.
Published: (2025)
How Fast Could Supermassive Black Holes Grow At the Epoch of Reionization?
by: Wu, Ziyong, et al.
Published: (2025)
by: Wu, Ziyong, et al.
Published: (2025)
Redshift Evolution of the Ratio of Supermassive Black Hole Mass to Stellar Mass
by: Wu, Ziyong, et al.
Published: (2026)
by: Wu, Ziyong, et al.
Published: (2026)
Fresh2comm: Information Freshness Optimized Collaborative Perception
by: Wu, Ziyong, et al.
Published: (2025)
by: Wu, Ziyong, et al.
Published: (2025)
Membership Inference Attacks on Large-Scale Models: A Survey
by: Wu, Hengyu, et al.
Published: (2025)
by: Wu, Hengyu, et al.
Published: (2025)
SLIM: Stealthy Low-Coverage Black-Box Watermarking via Latent-Space Confusion Zones
by: Wu, Hengyu, et al.
Published: (2026)
by: Wu, Hengyu, et al.
Published: (2026)
Non‐invasive analysis of Chinese traditional dyes in royal textiles of the Qing dynasty by excitation–emission matrix fluorescence
by: Li Zhao, et al.
Published: (2024)
by: Li Zhao, et al.
Published: (2024)
One Perturbation is Enough: On Generating Universal Adversarial Perturbations against Vision-Language Pre-training Models
by: Fang, Hao, et al.
Published: (2024)
by: Fang, Hao, et al.
Published: (2024)
AutoTVG: A New Vision-language Pre-training Paradigm for Temporal Video Grounding
by: Zhang, Xing, et al.
Published: (2024)
by: Zhang, Xing, et al.
Published: (2024)
Qing Tang
by: Qing Tang
Published: (2026)
by: Qing Tang
Published: (2026)
Qing Tang
by: Qing Tang
Published: (2026)
by: Qing Tang
Published: (2026)
Distributionally Robust Alignment for Medical Federated Vision-Language Pre-training Under Data Heterogeneity
by: Shuai, Zitao, et al.
Published: (2024)
by: Shuai, Zitao, et al.
Published: (2024)
Hán Dān Xué Bù (Mimicry) or Qīng Chū Yú Lán (Mastery)? A Cognitive Perspective on Reasoning Distillation in Large Language Models
by: Hu, Yueqing, et al.
Published: (2026)
by: Hu, Yueqing, et al.
Published: (2026)
Cross-layer Attention Sharing for Pre-trained Large Language Models
by: Mu, Yongyu, et al.
Published: (2024)
by: Mu, Yongyu, et al.
Published: (2024)
Downstream Transfer Attack: Adversarial Attacks on Downstream Models with Pre-trained Vision Transformers
by: Zheng, Weijie, et al.
Published: (2024)
by: Zheng, Weijie, et al.
Published: (2024)
Twisted Poincaré duality for orientable Poisson manifolds
by: Qi, Tiancheng, et al.
Published: (2025)
by: Qi, Tiancheng, et al.
Published: (2025)
Vertical Federated Learning Hybrid Local Pre-training
by: Li, Wenguo, et al.
Published: (2024)
by: Li, Wenguo, et al.
Published: (2024)
CERD: A Comprehensive Chinese Rhetoric Dataset for Rhetorical Understanding and Generation in Essays
by: Liu, Nuowei, et al.
Published: (2024)
by: Liu, Nuowei, et al.
Published: (2024)
Similar Items
-
UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards
by: Wang, Jun, et al.
Published: (2026) -
RWKV-CLIP: A Robust Vision-Language Representation Learner
by: Gu, Tiancheng, et al.
Published: (2024) -
ORID: Organ-Regional Information Driven Framework for Radiology Report Generation
by: Gu, Tiancheng, et al.
Published: (2024) -
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
by: Yang, Kaicheng, et al.
Published: (2024) -
LaPA: Latent Prompt Assist Model For Medical Visual Question Answering
by: Gu, Tiancheng, et al.
Published: (2024)