EVLF: Early Vision-Language Fusion for Generative Dataset Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Cai, Wenqi, Zou, Yawen, Li, Guang, Gu, Chunzhi, Zhang, Chao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Label-Consistent Dataset Distillation with Detector-Guided Refinement
by: Zou, Yawen, et al.
Published: (2025)
by: Zou, Yawen, et al.
Published: (2025)
Dataset Distillation via Vision-Language Category Prototype
by: Zou, Yawen, et al.
Published: (2025)
by: Zou, Yawen, et al.
Published: (2025)
VID-AD: A Dataset for Image-Level Logical Anomaly Detection under Vision-Induced Distraction
by: Nakata, Hiroto, et al.
Published: (2026)
by: Nakata, Hiroto, et al.
Published: (2026)
EVLF-FM: Explainable Vision Language Foundation Model for Medicine
by: Bai, Yang, et al.
Published: (2025)
by: Bai, Yang, et al.
Published: (2025)
Incremental Pseudo-Labeling for Black-Box Unsupervised Domain Adaptation
by: Zou, Yawen, et al.
Published: (2024)
by: Zou, Yawen, et al.
Published: (2024)
Multi-Scale Distillation for RGB-D Anomaly Detection on the PD-REAL Dataset
by: Qin, Jianjian, et al.
Published: (2023)
by: Qin, Jianjian, et al.
Published: (2023)
Orientation-Aware Leg Movement Learning for Action-Driven Human Motion Prediction
by: Gu, Chunzhi, et al.
Published: (2023)
by: Gu, Chunzhi, et al.
Published: (2023)
Vision-Language Dataset Distillation
by: Wu, Xindi, et al.
Published: (2023)
by: Wu, Xindi, et al.
Published: (2023)
VideoDistill: Language-aware Vision Distillation for Video Question Answering
by: Zou, Bo, et al.
Published: (2024)
by: Zou, Bo, et al.
Published: (2024)
Few-shot Human Action Anomaly Detection via a Unified Contrastive Learning Framework
by: Kamide, Koichiro, et al.
Published: (2025)
by: Kamide, Koichiro, et al.
Published: (2025)
Diverse Code Query Learning for Speech-Driven Facial Animation
by: Gu, Chunzhi, et al.
Published: (2024)
by: Gu, Chunzhi, et al.
Published: (2024)
3D Human-Human Interaction Anomaly Detection
by: Maeda, Shun, et al.
Published: (2025)
by: Maeda, Shun, et al.
Published: (2025)
Frequency-Guided Multi-Level Human Action Anomaly Detection with Normalizing Flows
by: Maeda, Shun, et al.
Published: (2024)
by: Maeda, Shun, et al.
Published: (2024)
Data-Efficient Generation for Dataset Distillation
by: Li, Zhe, et al.
Published: (2024)
by: Li, Zhe, et al.
Published: (2024)
VisionLLM-based Multimodal Fusion Network for Glottic Carcinoma Early Detection
by: Jin, Zhaohui, et al.
Published: (2024)
by: Jin, Zhaohui, et al.
Published: (2024)
Generative Dataset Distillation Based on Self-knowledge Distillation
by: Li, Longzhen, et al.
Published: (2025)
by: Li, Longzhen, et al.
Published: (2025)
DRKF: Distilled Rotated Kernel Fusion for Efficient Rotation Invariant Descriptors in Local Feature Matching
by: Huang, Ranran, et al.
Published: (2022)
by: Huang, Ranran, et al.
Published: (2022)
InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models
by: Deng, Nianchen, et al.
Published: (2025)
by: Deng, Nianchen, et al.
Published: (2025)
EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model
by: Zhang, Yuxuan, et al.
Published: (2024)
by: Zhang, Yuxuan, et al.
Published: (2024)
Multimodal Distribution Matching for Vision-Language Dataset Distillation
by: Jeong, Jongoh, et al.
Published: (2026)
by: Jeong, Jongoh, et al.
Published: (2026)
ControlFusion: A Controllable Image Fusion Framework with Language-Vision Degradation Prompts
by: Tang, Linfeng, et al.
Published: (2025)
by: Tang, Linfeng, et al.
Published: (2025)
BcQLM: Efficient Vision-Language Understanding with Distilled Q-Gated Cross-Modal Fusion
by: Xiang, Sike, et al.
Published: (2025)
by: Xiang, Sike, et al.
Published: (2025)
Visual-Advantage On-Policy Distillation for Vision-Language Models
by: Liu, Ruiqi, et al.
Published: (2026)
by: Liu, Ruiqi, et al.
Published: (2026)
Dataset Distillation by Automatic Training Trajectories
by: Liu, Dai, et al.
Published: (2024)
by: Liu, Dai, et al.
Published: (2024)
HIERAMP: Coarse-to-Fine Autoregressive Amplification for Generative Dataset Distillation
by: Zhao, Lin, et al.
Published: (2026)
by: Zhao, Lin, et al.
Published: (2026)
PromptKD: Unsupervised Prompt Distillation for Vision-Language Models
by: Li, Zheng, et al.
Published: (2024)
by: Li, Zheng, et al.
Published: (2024)
Hyperbolic Dataset Distillation
by: Li, Wenyuan, et al.
Published: (2025)
by: Li, Wenyuan, et al.
Published: (2025)
Closed-Form Linear-Probe Dataset Distillation for Pre-trained Vision Models
by: Peng, Bincheng, et al.
Published: (2026)
by: Peng, Bincheng, et al.
Published: (2026)
Enhancing Diffusion-based Dataset Distillation via Adversary-Guided Curriculum Sampling
by: Zou, Lexiao, et al.
Published: (2025)
by: Zou, Lexiao, et al.
Published: (2025)
Latent Video Dataset Distillation
by: Li, Ning, et al.
Published: (2025)
by: Li, Ning, et al.
Published: (2025)
Decoupled Audio-Visual Dataset Distillation
by: Li, Wenyuan, et al.
Published: (2025)
by: Li, Wenyuan, et al.
Published: (2025)
Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives
by: Weng, Xingxing, et al.
Published: (2025)
by: Weng, Xingxing, et al.
Published: (2025)
Image Fusion via Vision-Language Model
by: Zhao, Zixiang, et al.
Published: (2024)
by: Zhao, Zixiang, et al.
Published: (2024)
InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models
by: Sakai, Shunsuke, et al.
Published: (2025)
by: Sakai, Shunsuke, et al.
Published: (2025)
Generative Dataset Distillation Based on Diffusion Model
by: Su, Duo, et al.
Published: (2024)
by: Su, Duo, et al.
Published: (2024)
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
by: Zhang, Tiezheng, et al.
Published: (2025)
by: Zhang, Tiezheng, et al.
Published: (2025)
GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer
by: Jia, Ding, et al.
Published: (2024)
by: Jia, Ding, et al.
Published: (2024)
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
by: Ye, Wencheng, et al.
Published: (2025)
by: Ye, Wencheng, et al.
Published: (2025)
Cross-View Consistency Regularisation for Knowledge Distillation
by: Zhang, Weijia, et al.
Published: (2024)
by: Zhang, Weijia, et al.
Published: (2024)
LM4LV: A Frozen Large Language Model for Low-level Vision Tasks
by: Zheng, Boyang, et al.
Published: (2024)
by: Zheng, Boyang, et al.
Published: (2024)
Similar Items
-
Label-Consistent Dataset Distillation with Detector-Guided Refinement
by: Zou, Yawen, et al.
Published: (2025) -
Dataset Distillation via Vision-Language Category Prototype
by: Zou, Yawen, et al.
Published: (2025) -
VID-AD: A Dataset for Image-Level Logical Anomaly Detection under Vision-Induced Distraction
by: Nakata, Hiroto, et al.
Published: (2026) -
EVLF-FM: Explainable Vision Language Foundation Model for Medicine
by: Bai, Yang, et al.
Published: (2025) -
Incremental Pseudo-Labeling for Black-Box Unsupervised Domain Adaptation
by: Zou, Yawen, et al.
Published: (2024)