MARché: Fast Masked Autoregressive Image Generation with Cache-Aware Attention
Fuente:
arXiv
Salvato in:
| Autori principali: | Jiang, Chaoyi, Kim, Sungwoo, Gao, Lei, Zarch, Hossein Entezari, Ro, Won Woo, Annavaram, Murali |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
di: Jiang, Chaoyi, et al.
Pubblicazione: (2024)
di: Jiang, Chaoyi, et al.
Pubblicazione: (2024)
CADC: Encoding User-Item Interactions for Compressing Recommendation Model Training Data
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2024)
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2024)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2025)
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2025)
DEL: Context-Aware Dynamic Exit Layer for Efficient Self-Speculative Decoding
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2025)
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2025)
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
di: Ramachandran, Arun, et al.
Pubblicazione: (2025)
di: Ramachandran, Arun, et al.
Pubblicazione: (2025)
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing
di: Gao, Lei, et al.
Pubblicazione: (2025)
di: Gao, Lei, et al.
Pubblicazione: (2025)
Fast NF4 Dequantization Kernels for Large Language Model Inference
di: Qi, Xiangbo, et al.
Pubblicazione: (2026)
di: Qi, Xiangbo, et al.
Pubblicazione: (2026)
Differentially Private Retrieval-Augmented Generation
di: Tang, Tingting, et al.
Pubblicazione: (2026)
di: Tang, Tingting, et al.
Pubblicazione: (2026)
PR-MaGIC: Prompt Refinement Via Mask Decoder Gradient Flow For In-Context Segmentation
di: Lee, Minjae, et al.
Pubblicazione: (2026)
di: Lee, Minjae, et al.
Pubblicazione: (2026)
Motion-Aware Caching for Efficient Autoregressive Video Generation
di: Xu, Jing, et al.
Pubblicazione: (2026)
di: Xu, Jing, et al.
Pubblicazione: (2026)
Rethinking Training for De-biasing Text-to-Image Generation: Unlocking the Potential of Stable Diffusion
di: Kim, Eunji, et al.
Pubblicazione: (2024)
di: Kim, Eunji, et al.
Pubblicazione: (2024)
Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention
di: Samuel, Dvir, et al.
Pubblicazione: (2026)
di: Samuel, Dvir, et al.
Pubblicazione: (2026)
Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks
di: Zou, Siyu, et al.
Pubblicazione: (2024)
di: Zou, Siyu, et al.
Pubblicazione: (2024)
GECKO: Generative Language Model for English, Code and Korean
di: Oh, Sungwoo, et al.
Pubblicazione: (2024)
di: Oh, Sungwoo, et al.
Pubblicazione: (2024)
Edge Private Graph Neural Networks with Singular Value Perturbation
di: Tang, Tingting, et al.
Pubblicazione: (2024)
di: Tang, Tingting, et al.
Pubblicazione: (2024)
$A^3$: Attention-Aware Accurate KV Cache Fusion for Fast Large Language Model Serving
di: Zhou, Yuechi, et al.
Pubblicazione: (2025)
di: Zhou, Yuechi, et al.
Pubblicazione: (2025)
The Illusion of Equivalence: Systematic FP16 Divergence in KV-Cached Autoregressive Inference
di: Chodavarapu, Ranjith, et al.
Pubblicazione: (2026)
di: Chodavarapu, Ranjith, et al.
Pubblicazione: (2026)
Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
di: Zheng, Zirui, et al.
Pubblicazione: (2025)
di: Zheng, Zirui, et al.
Pubblicazione: (2025)
Direction-Aware Diagonal Autoregressive Image Generation
di: Xu, Yijia, et al.
Pubblicazione: (2025)
di: Xu, Yijia, et al.
Pubblicazione: (2025)
Fast Autoregressive Video Generation with Diagonal Decoding
di: Ye, Yang, et al.
Pubblicazione: (2025)
di: Ye, Yang, et al.
Pubblicazione: (2025)
Make It Efficient: Dynamic Sparse Attention for Autoregressive Image Generation
di: Xiang, Xunzhi, et al.
Pubblicazione: (2025)
di: Xiang, Xunzhi, et al.
Pubblicazione: (2025)
Transformable Gaussian Reward Function for Socially-Aware Navigation with Deep Reinforcement Learning
di: Kim, Jinyeob, et al.
Pubblicazione: (2024)
di: Kim, Jinyeob, et al.
Pubblicazione: (2024)
DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer
di: Wu, Yecheng, et al.
Pubblicazione: (2025)
di: Wu, Yecheng, et al.
Pubblicazione: (2025)
InstGenIE: Generative Image Editing Made Efficient with Mask-aware Caching and Scheduling
di: Jiang, Xiaoxiao, et al.
Pubblicazione: (2025)
di: Jiang, Xiaoxiao, et al.
Pubblicazione: (2025)
Autoregressive Omni-Aware Outpainting for Open-Vocabulary 360-Degree Image Generation
di: Lu, Zhuqiang, et al.
Pubblicazione: (2023)
di: Lu, Zhuqiang, et al.
Pubblicazione: (2023)
MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
di: Liu, Haozhe, et al.
Pubblicazione: (2024)
di: Liu, Haozhe, et al.
Pubblicazione: (2024)
CacheFormer: High Attention-Based Segment Caching
di: Singh, Sushant, et al.
Pubblicazione: (2025)
di: Singh, Sushant, et al.
Pubblicazione: (2025)
Diagnosing Korean-Language LLM Political Bias via Census-Grounded Agent Simulation
di: Kang, Sungwoo
Pubblicazione: (2026)
di: Kang, Sungwoo
Pubblicazione: (2026)
STAG-CN: Spatio-Temporal Apiary Graph Convolutional Network for Disease Onset Prediction in Beehive Sensor Networks
di: Kang, Sungwoo
Pubblicazione: (2026)
di: Kang, Sungwoo
Pubblicazione: (2026)
LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation
di: Ahn, Jinwoo, et al.
Pubblicazione: (2026)
di: Ahn, Jinwoo, et al.
Pubblicazione: (2026)
REPrune: Channel Pruning via Kernel Representative Selection
di: Park, Mincheol, et al.
Pubblicazione: (2024)
di: Park, Mincheol, et al.
Pubblicazione: (2024)
Polynomial Context-Truncation Sensitivity in Autoregressive Language Models: Sequential Wyner-Ziv Bounds for KV Cache Compression
di: Kim, Munsik
Pubblicazione: (2026)
di: Kim, Munsik
Pubblicazione: (2026)
Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
di: Park, Joonhyung, et al.
Pubblicazione: (2025)
di: Park, Joonhyung, et al.
Pubblicazione: (2025)
UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
di: Kang, Wonjun, et al.
Pubblicazione: (2025)
di: Kang, Wonjun, et al.
Pubblicazione: (2025)
Using Span Queries to Optimize for Cache and Attention Locality
di: Castro, Paul, et al.
Pubblicazione: (2025)
di: Castro, Paul, et al.
Pubblicazione: (2025)
QuickMerge++: Fast Token Merging with Autoregressive Prior
di: Liu, Dong, et al.
Pubblicazione: (2025)
di: Liu, Dong, et al.
Pubblicazione: (2025)
Enabling Autoregressive Models to Fill In Masked Tokens
di: Israel, Daniel, et al.
Pubblicazione: (2025)
di: Israel, Daniel, et al.
Pubblicazione: (2025)
Modern Deep Learning Approaches for Cricket Shot Classification: A Comprehensive Baseline Study
di: Kang, Sungwoo
Pubblicazione: (2025)
di: Kang, Sungwoo
Pubblicazione: (2025)
Autoregressive Image Generation with Linear Complexity: A Spatial-Aware Decay Perspective
di: Mao, Yuxin, et al.
Pubblicazione: (2025)
di: Mao, Yuxin, et al.
Pubblicazione: (2025)
RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse
di: Geng, Yingsheng, et al.
Pubblicazione: (2026)
di: Geng, Yingsheng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
di: Jiang, Chaoyi, et al.
Pubblicazione: (2024) -
CADC: Encoding User-Item Interactions for Compressing Recommendation Model Training Data
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2024) -
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2025) -
DEL: Context-Aware Dynamic Exit Layer for Efficient Self-Speculative Decoding
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2025) -
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
di: Ramachandran, Arun, et al.
Pubblicazione: (2025)