Distilling Vision-Language Models on Millions of Videos
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhao, Yue, Zhao, Long, Zhou, Xingyi, Wu, Jialin, Chu, Chun-Te, Miao, Hui, Schroff, Florian, Adam, Hartwig, Liu, Ting, Gong, Boqing, Krähenbühl, Philipp, Yuan, Liangzhe |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding
di: Xiong, Yuanhao, et al.
Pubblicazione: (2023)
di: Xiong, Yuanhao, et al.
Pubblicazione: (2023)
VideoGLUE: Video General Understanding Evaluation of Foundation Models
di: Yuan, Liangzhe, et al.
Pubblicazione: (2023)
di: Yuan, Liangzhe, et al.
Pubblicazione: (2023)
VideoPrism: A Foundational Visual Encoder for Video Understanding
di: Zhao, Long, et al.
Pubblicazione: (2024)
di: Zhao, Long, et al.
Pubblicazione: (2024)
Image and Video Tokenization with Binary Spherical Quantization
di: Zhao, Yue, et al.
Pubblicazione: (2024)
di: Zhao, Yue, et al.
Pubblicazione: (2024)
Interactive Post-Training for Vision-Language-Action Models
di: Tan, Shuhan, et al.
Pubblicazione: (2025)
di: Tan, Shuhan, et al.
Pubblicazione: (2025)
Video Creation by Demonstration
di: Sun, Yihong, et al.
Pubblicazione: (2024)
di: Sun, Yihong, et al.
Pubblicazione: (2024)
Domain Adaptation Through Task Distillation
di: Zhou, Brady, et al.
Pubblicazione: (2020)
di: Zhou, Brady, et al.
Pubblicazione: (2020)
Epsilon-VAE: Denoising as Visual Decoding
di: Zhao, Long, et al.
Pubblicazione: (2024)
di: Zhao, Long, et al.
Pubblicazione: (2024)
The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
di: Tan, Yuwen, et al.
Pubblicazione: (2025)
di: Tan, Yuwen, et al.
Pubblicazione: (2025)
Image Diffusion Preview with Consistency Solver
di: Wang, Fu-Yun, et al.
Pubblicazione: (2025)
di: Wang, Fu-Yun, et al.
Pubblicazione: (2025)
Compressed Map Priors for 3D Perception
di: Zhou, Brady, et al.
Pubblicazione: (2025)
di: Zhou, Brady, et al.
Pubblicazione: (2025)
Moiré Video Authentication: A Physical Signature Against AI Video Generation
di: Qing, Yuan, et al.
Pubblicazione: (2026)
di: Qing, Yuan, et al.
Pubblicazione: (2026)
VideoDistill: Language-aware Vision Distillation for Video Question Answering
di: Zou, Bo, et al.
Pubblicazione: (2024)
di: Zou, Bo, et al.
Pubblicazione: (2024)
Spherical Leech Quantization for Visual Tokenization and Generation
di: Zhao, Yue, et al.
Pubblicazione: (2025)
di: Zhao, Yue, et al.
Pubblicazione: (2025)
Lifting Data-Tracing Machine Unlearning to Knowledge-Tracing for Foundation Models
di: Tan, Yuwen, et al.
Pubblicazione: (2025)
di: Tan, Yuwen, et al.
Pubblicazione: (2025)
GEOMETRIA ÓSSEA E ATIVIDADE FÍSICA EM CRIANÇAS E ADOLESCENTES: REVISÃO SISTEMÁTICA
di: Tathyane Krahenbühl
Pubblicazione: (2018)
di: Tathyane Krahenbühl
Pubblicazione: (2018)
The use of the additional field player in handball: analysis of the Rio 2016 Olympic Games
di: Tathyane Krahenbühl
Pubblicazione: (2019)
di: Tathyane Krahenbühl
Pubblicazione: (2019)
Fatores que influenciam a massa óssea de crianças e adolescentes saudáveis mensurada pelo ultrassom quantitativo de falanges: revisão sistemática
di: Tathyane Krahenbühl
Pubblicazione: (2014)
di: Tathyane Krahenbühl
Pubblicazione: (2014)
Long-term Traffic Simulation with Interleaved Autoregressive Motion and Scenario Generation
di: Yang, Xiuyu, et al.
Pubblicazione: (2025)
di: Yang, Xiuyu, et al.
Pubblicazione: (2025)
HypDAE: Hyperbolic Diffusion Autoencoders for Hierarchical Few-shot Image Generation
di: Li, Lingxiao, et al.
Pubblicazione: (2024)
di: Li, Lingxiao, et al.
Pubblicazione: (2024)
The transcription factor LbYABBY1 of Limonium bicolor raises salt sensitivity by repressing the LbSAD2 pathway
di: Boqing Zhao, et al.
Pubblicazione: (2025)
di: Boqing Zhao, et al.
Pubblicazione: (2025)
VideoAds for Fast-Paced Video Understanding
di: Zhang, Zheyuan, et al.
Pubblicazione: (2025)
di: Zhang, Zheyuan, et al.
Pubblicazione: (2025)
On Discrete Prompt Optimization for Diffusion Models
di: Wang, Ruochen, et al.
Pubblicazione: (2024)
di: Wang, Ruochen, et al.
Pubblicazione: (2024)
LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs
di: Kim, Jihwan, et al.
Pubblicazione: (2026)
di: Kim, Jihwan, et al.
Pubblicazione: (2026)
Attention to Neural Plagiarism: Diffusion Models Can Plagiarize Your Copyrighted Images!
di: Zou, Zihang, et al.
Pubblicazione: (2026)
di: Zou, Zihang, et al.
Pubblicazione: (2026)
Culture in Action: Evaluating Text-to-Image Models through Social Activities
di: Malakouti, Sina, et al.
Pubblicazione: (2025)
di: Malakouti, Sina, et al.
Pubblicazione: (2025)
1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training
di: Zhao, Han, et al.
Pubblicazione: (2025)
di: Zhao, Han, et al.
Pubblicazione: (2025)
Deranged Perfect Matchings on complete graph and balanced complete r-partite graph
di: Deng, Boqing
Pubblicazione: (2025)
di: Deng, Boqing
Pubblicazione: (2025)
Distill Video Datasets into Images
di: Zhao, Zhenghao, et al.
Pubblicazione: (2025)
di: Zhao, Zhenghao, et al.
Pubblicazione: (2025)
Performance and accuracy of cross‐section tracking methods for hydromorphological habitat assessment in wadable rivers with sparse canopy conditions
di: Robin Schroff, et al.
Pubblicazione: (2024)
di: Robin Schroff, et al.
Pubblicazione: (2024)
Neptune: The Long Orbit to Benchmarking Long Video Understanding
di: Nagrani, Arsha, et al.
Pubblicazione: (2024)
di: Nagrani, Arsha, et al.
Pubblicazione: (2024)
Accelerating Attention with Basis Decomposition
di: Zhao, Jialin
Pubblicazione: (2025)
di: Zhao, Jialin
Pubblicazione: (2025)
Compositional Video Generation as Flow Equalization
di: Yang, Xingyi, et al.
Pubblicazione: (2024)
di: Yang, Xingyi, et al.
Pubblicazione: (2024)
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
di: Zhao, Yue, et al.
Pubblicazione: (2025)
di: Zhao, Yue, et al.
Pubblicazione: (2025)
Rapid stochastic spatial light modulator calibration and pixel crosstalk optimisation
di: Schroff, P., et al.
Pubblicazione: (2024)
di: Schroff, P., et al.
Pubblicazione: (2024)
Continual Adapter Tuning with Semantic Shift Compensation for Class-Incremental Learning
di: Zhou, Qinhao, et al.
Pubblicazione: (2024)
di: Zhou, Qinhao, et al.
Pubblicazione: (2024)
TASeg: Text-aware RGB-T Semantic Segmentation based on Fine-tuning Vision Foundation Models
di: Yu, Meng, et al.
Pubblicazione: (2025)
di: Yu, Meng, et al.
Pubblicazione: (2025)
Generalization of a density‐dependent ecosystem function in dominant aquatic macroinvertebrates
di: Roman Alther, et al.
Pubblicazione: (2024)
di: Roman Alther, et al.
Pubblicazione: (2024)
Think over Trajectories: Leveraging Video Generation to Reconstruct GPS Trajectories from Cellular Signaling
di: Zhang, Ruixing, et al.
Pubblicazione: (2026)
di: Zhang, Ruixing, et al.
Pubblicazione: (2026)
The Wisdom of Agent Crowds: A Human-AI Interaction Innovation Ignition Framework
di: Yang, Senhao, et al.
Pubblicazione: (2025)
di: Yang, Senhao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding
di: Xiong, Yuanhao, et al.
Pubblicazione: (2023) -
VideoGLUE: Video General Understanding Evaluation of Foundation Models
di: Yuan, Liangzhe, et al.
Pubblicazione: (2023) -
VideoPrism: A Foundational Visual Encoder for Video Understanding
di: Zhao, Long, et al.
Pubblicazione: (2024) -
Image and Video Tokenization with Binary Spherical Quantization
di: Zhao, Yue, et al.
Pubblicazione: (2024) -
Interactive Post-Training for Vision-Language-Action Models
di: Tan, Shuhan, et al.
Pubblicazione: (2025)