A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Wangbo, Han, Yizeng, Tang, Jiasheng, Li, Zhikai, Song, Yibing, Wang, Kai, Wang, Zhangyang, You, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dynamic Diffusion Transformer
von: Zhao, Wangbo, et al.
Veröffentlicht: (2024)
von: Zhao, Wangbo, et al.
Veröffentlicht: (2024)
Dynamic Tuning Towards Parameter and Inference Efficiency for ViT Adaptation
von: Zhao, Wangbo, et al.
Veröffentlicht: (2024)
von: Zhao, Wangbo, et al.
Veröffentlicht: (2024)
RAPID^3: Tri-Level Reinforced Acceleration Policies for Diffusion Transformer
von: Zhao, Wangbo, et al.
Veröffentlicht: (2025)
von: Zhao, Wangbo, et al.
Veröffentlicht: (2025)
DyDiT++: Diffusion Transformers with Timestep and Spatial Dynamics for Efficient Visual Generation
von: Zhao, Wangbo, et al.
Veröffentlicht: (2025)
von: Zhao, Wangbo, et al.
Veröffentlicht: (2025)
A Stitch in Time Saves Nine: Proactive Self-Refinement for Language Models
von: Han, Jinyi, et al.
Veröffentlicht: (2025)
von: Han, Jinyi, et al.
Veröffentlicht: (2025)
Recurrent Diffusion for Large-Scale Parameter Generation
von: Wang, Kai, et al.
Veröffentlicht: (2025)
von: Wang, Kai, et al.
Veröffentlicht: (2025)
K-Sort Eval: Efficient Preference Evaluation for Visual Generation via Corrected VLM-as-a-Judge
von: Li, Zhikai, et al.
Veröffentlicht: (2026)
von: Li, Zhikai, et al.
Veröffentlicht: (2026)
Few-Step Distillation for Text-to-Image Generation: A Practical Guide
von: Pu, Yifan, et al.
Veröffentlicht: (2025)
von: Pu, Yifan, et al.
Veröffentlicht: (2025)
BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
Conditional LoRA Parameter Generation
von: Jin, Xiaolong, et al.
Veröffentlicht: (2024)
von: Jin, Xiaolong, et al.
Veröffentlicht: (2024)
Speculative Thinking: Enhancing Small-Model Reasoning with Large Model Guidance at Inference Time
von: Yang, Wang, et al.
Veröffentlicht: (2025)
von: Yang, Wang, et al.
Veröffentlicht: (2025)
FracTrain: Fractionally Squeezing Bit Savings Both Temporally and Spatially for Efficient DNN Training
von: Fu, Yonggan, et al.
Veröffentlicht: (2020)
von: Fu, Yonggan, et al.
Veröffentlicht: (2020)
REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training
von: Wang, Ziqiao, et al.
Veröffentlicht: (2025)
von: Wang, Ziqiao, et al.
Veröffentlicht: (2025)
Unsupervised Learning for Class Distribution Mismatch
von: Du, Pan, et al.
Veröffentlicht: (2025)
von: Du, Pan, et al.
Veröffentlicht: (2025)
MentorCollab: Selective Large-to-Small Inference-Time Guidance for Efficient Reasoning
von: Wang, Haojin, et al.
Veröffentlicht: (2026)
von: Wang, Haojin, et al.
Veröffentlicht: (2026)
Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
von: Liang, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Liang, Zhiyuan, et al.
Veröffentlicht: (2025)
BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant Learning
von: Wang, Shengao, et al.
Veröffentlicht: (2025)
von: Wang, Shengao, et al.
Veröffentlicht: (2025)
Small-Large Collaboration: Training-efficient Concept Personalization for Large VLM using a Meta Personalized Small VLM
von: Yang, Sihan, et al.
Veröffentlicht: (2025)
von: Yang, Sihan, et al.
Veröffentlicht: (2025)
T-Stitch: Accelerating Sampling in Pre-Trained Diffusion Models with Trajectory Stitching
von: Pan, Zizheng, et al.
Veröffentlicht: (2024)
von: Pan, Zizheng, et al.
Veröffentlicht: (2024)
HybridStitch: Pixel and Timestep Level Model Stitching for Diffusion Acceleration
von: Sun, Desen, et al.
Veröffentlicht: (2026)
von: Sun, Desen, et al.
Veröffentlicht: (2026)
GSVA: Generalized Segmentation via Multimodal Large Language Models
von: Xia, Zhuofan, et al.
Veröffentlicht: (2023)
von: Xia, Zhuofan, et al.
Veröffentlicht: (2023)
DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution
von: Yue, Yang, et al.
Veröffentlicht: (2024)
von: Yue, Yang, et al.
Veröffentlicht: (2024)
Accelerating Parallel Sampling of Diffusion Models
von: Tang, Zhiwei, et al.
Veröffentlicht: (2024)
von: Tang, Zhiwei, et al.
Veröffentlicht: (2024)
HumanNOVA: Photorealistic, Universal and Rapid 3D Human Avatar Modeling from a Single Image
von: Hu, Hezhen, et al.
Veröffentlicht: (2026)
von: Hu, Hezhen, et al.
Veröffentlicht: (2026)
MyVLM: Personalizing VLMs for User-Specific Queries
von: Alaluf, Yuval, et al.
Veröffentlicht: (2024)
von: Alaluf, Yuval, et al.
Veröffentlicht: (2024)
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
von: Pani, Anupam, et al.
Veröffentlicht: (2025)
von: Pani, Anupam, et al.
Veröffentlicht: (2025)
GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets
von: Yao, Zhangyang, et al.
Veröffentlicht: (2026)
von: Yao, Zhangyang, et al.
Veröffentlicht: (2026)
MoNE: Replacing Redundant Experts with Lightweight Novices for Structured Pruning of MoE
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
Road Rage Reasoning with Vision-language Models (VLMs): Task Definition and Evaluation Dataset
von: Weng, Yibing, et al.
Veröffentlicht: (2025)
von: Weng, Yibing, et al.
Veröffentlicht: (2025)
EfficientTrain++: Generalized Curriculum Learning for Efficient Visual Backbone Training
von: Wang, Yulin, et al.
Veröffentlicht: (2024)
von: Wang, Yulin, et al.
Veröffentlicht: (2024)
Position: Weight Space Should Be a First-Class Generative AI Modality
von: Wang, Zhangyang, et al.
Veröffentlicht: (2026)
von: Wang, Zhangyang, et al.
Veröffentlicht: (2026)
Precision Construction of Chiral Optical Fields with Nine Controllable Degrees of Freedom
von: Duo Deng, et al.
Veröffentlicht: (2025)
von: Duo Deng, et al.
Veröffentlicht: (2025)
Ultraspherical Spectral Method for Block Copolymer Systems on Unit Disk
von: Luo, Wangbo, et al.
Veröffentlicht: (2024)
von: Luo, Wangbo, et al.
Veröffentlicht: (2024)
Double Fourier Sphere Methods with Low Rank Approximation for Block Copolymer Systems on Sphere
von: Luo, Wangbo, et al.
Veröffentlicht: (2025)
von: Luo, Wangbo, et al.
Veröffentlicht: (2025)
Asymptotically compatible schemes for nonlocal Ohta–Kawasaki model
von: Wangbo Luo, et al.
Veröffentlicht: (2024)
von: Wangbo Luo, et al.
Veröffentlicht: (2024)
CLGRPO: Reasoning Ability Enhancement for Small VLMs
von: Wang, Fanyi, et al.
Veröffentlicht: (2025)
von: Wang, Fanyi, et al.
Veröffentlicht: (2025)
World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
von: Zhang, Wanyue, et al.
Veröffentlicht: (2026)
von: Zhang, Wanyue, et al.
Veröffentlicht: (2026)
UniTTA: Unified Benchmark and Versatile Framework Towards Realistic Test-Time Adaptation
von: Du, Chaoqun, et al.
Veröffentlicht: (2024)
von: Du, Chaoqun, et al.
Veröffentlicht: (2024)
SimPro: A Simple Probabilistic Framework Towards Realistic Long-Tailed Semi-Supervised Learning
von: Du, Chaoqun, et al.
Veröffentlicht: (2024)
von: Du, Chaoqun, et al.
Veröffentlicht: (2024)
Enhance-A-Video: Better Generated Video for Free
von: Luo, Yang, et al.
Veröffentlicht: (2025)
von: Luo, Yang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Dynamic Diffusion Transformer
von: Zhao, Wangbo, et al.
Veröffentlicht: (2024) -
Dynamic Tuning Towards Parameter and Inference Efficiency for ViT Adaptation
von: Zhao, Wangbo, et al.
Veröffentlicht: (2024) -
RAPID^3: Tri-Level Reinforced Acceleration Policies for Diffusion Transformer
von: Zhao, Wangbo, et al.
Veröffentlicht: (2025) -
DyDiT++: Diffusion Transformers with Timestep and Spatial Dynamics for Efficient Visual Generation
von: Zhao, Wangbo, et al.
Veröffentlicht: (2025) -
A Stitch in Time Saves Nine: Proactive Self-Refinement for Language Models
von: Han, Jinyi, et al.
Veröffentlicht: (2025)