Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Zigeng, Ma, Xinyin, Fang, Gongfan, Wang, Xinchao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SlimSAM: 0.1% Data Makes Segment Anything Slim
von: Chen, Zigeng, et al.
Veröffentlicht: (2023)
von: Chen, Zigeng, et al.
Veröffentlicht: (2023)
AsyncDiff: Parallelizing Diffusion Models by Asynchronous Denoising
von: Chen, Zigeng, et al.
Veröffentlicht: (2024)
von: Chen, Zigeng, et al.
Veröffentlicht: (2024)
In-Video Instructions: Visual Signals as Generative Control
von: Fang, Gongfan, et al.
Veröffentlicht: (2025)
von: Fang, Gongfan, et al.
Veröffentlicht: (2025)
Remix-DiT: Mixing Diffusion Transformers for Multi-Expert Denoising
von: Fang, Gongfan, et al.
Veröffentlicht: (2024)
von: Fang, Gongfan, et al.
Veröffentlicht: (2024)
ConciseHint: Boosting Efficient Reasoning via Continuous Concise Hints during Generation
von: Tang, Siao, et al.
Veröffentlicht: (2025)
von: Tang, Siao, et al.
Veröffentlicht: (2025)
Q-ARVD: Quantizing Autoregressive Video Diffusion Models
von: Tang, Siao, et al.
Veröffentlicht: (2026)
von: Tang, Siao, et al.
Veröffentlicht: (2026)
Isomorphic Pruning for Vision Models
von: Fang, Gongfan, et al.
Veröffentlicht: (2024)
von: Fang, Gongfan, et al.
Veröffentlicht: (2024)
TinyFusion: Diffusion Transformers Learned Shallow
von: Fang, Gongfan, et al.
Veröffentlicht: (2024)
von: Fang, Gongfan, et al.
Veröffentlicht: (2024)
Learning-to-Cache: Accelerating Diffusion Transformer via Layer Caching
von: Ma, Xinyin, et al.
Veröffentlicht: (2024)
von: Ma, Xinyin, et al.
Veröffentlicht: (2024)
Dimple: Discrete Diffusion Multimodal Large Language Model with Parallel Decoding
von: Yu, Runpeng, et al.
Veröffentlicht: (2025)
von: Yu, Runpeng, et al.
Veröffentlicht: (2025)
Introducing Visual Perception Token into Multimodal Large Language Model
von: Yu, Runpeng, et al.
Veröffentlicht: (2025)
von: Yu, Runpeng, et al.
Veröffentlicht: (2025)
dParallel: Learnable Parallel Decoding for dLLMs
von: Chen, Zigeng, et al.
Veröffentlicht: (2025)
von: Chen, Zigeng, et al.
Veröffentlicht: (2025)
Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs
von: Lu, Haiquan, et al.
Veröffentlicht: (2026)
von: Lu, Haiquan, et al.
Veröffentlicht: (2026)
VeriThinker: Learning to Verify Makes Reasoning Model Efficient
von: Chen, Zigeng, et al.
Veröffentlicht: (2025)
von: Chen, Zigeng, et al.
Veröffentlicht: (2025)
Diversity-Guided MLP Reduction for Efficient Large Vision Transformers
von: Shen, Chengchao, et al.
Veröffentlicht: (2025)
von: Shen, Chengchao, et al.
Veröffentlicht: (2025)
DMax: Aggressive Parallel Decoding for dLLMs
von: Chen, Zigeng, et al.
Veröffentlicht: (2026)
von: Chen, Zigeng, et al.
Veröffentlicht: (2026)
dMoE: dLLMs with Learnable Block Experts
von: Feng, Sicheng, et al.
Veröffentlicht: (2026)
von: Feng, Sicheng, et al.
Veröffentlicht: (2026)
dVoting: Fast Voting for dLLMs
von: Feng, Sicheng, et al.
Veröffentlicht: (2026)
von: Feng, Sicheng, et al.
Veröffentlicht: (2026)
Heavy Labels Out! Dataset Distillation with Label Space Lightening
von: Yu, Ruonan, et al.
Veröffentlicht: (2024)
von: Yu, Ruonan, et al.
Veröffentlicht: (2024)
ViFeEdit: A Video-Free Tuner of Your Video Diffusion Transformer
von: Yu, Ruonan, et al.
Veröffentlicht: (2026)
von: Yu, Ruonan, et al.
Veröffentlicht: (2026)
OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
von: Li, Han, et al.
Veröffentlicht: (2025)
von: Li, Han, et al.
Veröffentlicht: (2025)
Language Model as Visual Explainer
von: Yang, Xingyi, et al.
Veröffentlicht: (2024)
von: Yang, Xingyi, et al.
Veröffentlicht: (2024)
PixelThink: Towards Efficient Chain-of-Pixel Reasoning
von: Wang, Song, et al.
Veröffentlicht: (2025)
von: Wang, Song, et al.
Veröffentlicht: (2025)
Efficient Reasoning Models: A Survey
von: Feng, Sicheng, et al.
Veröffentlicht: (2025)
von: Feng, Sicheng, et al.
Veröffentlicht: (2025)
Thinkless: LLM Learns When to Think
von: Fang, Gongfan, et al.
Veröffentlicht: (2025)
von: Fang, Gongfan, et al.
Veröffentlicht: (2025)
InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
von: Liu, Jinlai, et al.
Veröffentlicht: (2025)
von: Liu, Jinlai, et al.
Veröffentlicht: (2025)
Sample- and Parameter-Efficient Auto-Regressive Image Models
von: Amrani, Elad, et al.
Veröffentlicht: (2024)
von: Amrani, Elad, et al.
Veröffentlicht: (2024)
PointLoRA: Low-Rank Adaptation with Token Selection for Point Cloud Learning
von: Wang, Song, et al.
Veröffentlicht: (2025)
von: Wang, Song, et al.
Veröffentlicht: (2025)
VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling
von: Zhang, Qian, et al.
Veröffentlicht: (2024)
von: Zhang, Qian, et al.
Veröffentlicht: (2024)
Resurrect Mask AutoRegressive Modeling for Efficient and Scalable Image Generation
von: Xin, Yi, et al.
Veröffentlicht: (2025)
von: Xin, Yi, et al.
Veröffentlicht: (2025)
Rethinking Token Reduction for Large Vision-Language Models
von: Wang, Yi, et al.
Veröffentlicht: (2026)
von: Wang, Yi, et al.
Veröffentlicht: (2026)
LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?
von: Fang, Kechen, et al.
Veröffentlicht: (2026)
von: Fang, Kechen, et al.
Veröffentlicht: (2026)
dKV-Cache: The Cache for Diffusion Language Models
von: Ma, Xinyin, et al.
Veröffentlicht: (2025)
von: Ma, Xinyin, et al.
Veröffentlicht: (2025)
TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation
von: Chen, Zhekai, et al.
Veröffentlicht: (2025)
von: Chen, Zhekai, et al.
Veröffentlicht: (2025)
A Watermark for Auto-Regressive Image Generation Models
von: Wu, Yihan, et al.
Veröffentlicht: (2025)
von: Wu, Yihan, et al.
Veröffentlicht: (2025)
Efficient Gaussian Splatting for Monocular Dynamic Scene Rendering via Sparse Time-Variant Attribute Modeling
von: Kong, Hanyang, et al.
Veröffentlicht: (2025)
von: Kong, Hanyang, et al.
Veröffentlicht: (2025)
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning
von: li, Bonan, et al.
Veröffentlicht: (2025)
von: li, Bonan, et al.
Veröffentlicht: (2025)
EchoGen: Generating Visual Echoes in Any Scene via Feed-Forward Subject-Driven Auto-Regressive Model
von: Dong, Ruixiao, et al.
Veröffentlicht: (2025)
von: Dong, Ruixiao, et al.
Veröffentlicht: (2025)
AvatarPointillist: AutoRegressive 4D Gaussian Avatarization
von: Liu, Hongyu, et al.
Veröffentlicht: (2026)
von: Liu, Hongyu, et al.
Veröffentlicht: (2026)
AgentsCoMerge: Large Language Model Empowered Collaborative Decision Making for Ramp Merging
von: Hu, Senkang, et al.
Veröffentlicht: (2024)
von: Hu, Senkang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SlimSAM: 0.1% Data Makes Segment Anything Slim
von: Chen, Zigeng, et al.
Veröffentlicht: (2023) -
AsyncDiff: Parallelizing Diffusion Models by Asynchronous Denoising
von: Chen, Zigeng, et al.
Veröffentlicht: (2024) -
In-Video Instructions: Visual Signals as Generative Control
von: Fang, Gongfan, et al.
Veröffentlicht: (2025) -
Remix-DiT: Mixing Diffusion Transformers for Multi-Expert Denoising
von: Fang, Gongfan, et al.
Veröffentlicht: (2024) -
ConciseHint: Boosting Efficient Reasoning via Continuous Concise Hints during Generation
von: Tang, Siao, et al.
Veröffentlicht: (2025)