MARché: Fast Masked Autoregressive Image Generation with Cache-Aware Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Chaoyi, Kim, Sungwoo, Gao, Lei, Zarch, Hossein Entezari, Ro, Won Woo, Annavaram, Murali |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
von: Jiang, Chaoyi, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoyi, et al.
Veröffentlicht: (2024)
CADC: Encoding User-Item Interactions for Compressing Recommendation Model Training Data
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2024)
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2024)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2025)
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2025)
DEL: Context-Aware Dynamic Exit Layer for Efficient Self-Speculative Decoding
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2025)
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2025)
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
von: Ramachandran, Arun, et al.
Veröffentlicht: (2025)
von: Ramachandran, Arun, et al.
Veröffentlicht: (2025)
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing
von: Gao, Lei, et al.
Veröffentlicht: (2025)
von: Gao, Lei, et al.
Veröffentlicht: (2025)
Fast NF4 Dequantization Kernels for Large Language Model Inference
von: Qi, Xiangbo, et al.
Veröffentlicht: (2026)
von: Qi, Xiangbo, et al.
Veröffentlicht: (2026)
Differentially Private Retrieval-Augmented Generation
von: Tang, Tingting, et al.
Veröffentlicht: (2026)
von: Tang, Tingting, et al.
Veröffentlicht: (2026)
PR-MaGIC: Prompt Refinement Via Mask Decoder Gradient Flow For In-Context Segmentation
von: Lee, Minjae, et al.
Veröffentlicht: (2026)
von: Lee, Minjae, et al.
Veröffentlicht: (2026)
Motion-Aware Caching for Efficient Autoregressive Video Generation
von: Xu, Jing, et al.
Veröffentlicht: (2026)
von: Xu, Jing, et al.
Veröffentlicht: (2026)
Rethinking Training for De-biasing Text-to-Image Generation: Unlocking the Potential of Stable Diffusion
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention
von: Samuel, Dvir, et al.
Veröffentlicht: (2026)
von: Samuel, Dvir, et al.
Veröffentlicht: (2026)
Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks
von: Zou, Siyu, et al.
Veröffentlicht: (2024)
von: Zou, Siyu, et al.
Veröffentlicht: (2024)
GECKO: Generative Language Model for English, Code and Korean
von: Oh, Sungwoo, et al.
Veröffentlicht: (2024)
von: Oh, Sungwoo, et al.
Veröffentlicht: (2024)
Edge Private Graph Neural Networks with Singular Value Perturbation
von: Tang, Tingting, et al.
Veröffentlicht: (2024)
von: Tang, Tingting, et al.
Veröffentlicht: (2024)
$A^3$: Attention-Aware Accurate KV Cache Fusion for Fast Large Language Model Serving
von: Zhou, Yuechi, et al.
Veröffentlicht: (2025)
von: Zhou, Yuechi, et al.
Veröffentlicht: (2025)
The Illusion of Equivalence: Systematic FP16 Divergence in KV-Cached Autoregressive Inference
von: Chodavarapu, Ranjith, et al.
Veröffentlicht: (2026)
von: Chodavarapu, Ranjith, et al.
Veröffentlicht: (2026)
Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
von: Zheng, Zirui, et al.
Veröffentlicht: (2025)
von: Zheng, Zirui, et al.
Veröffentlicht: (2025)
Direction-Aware Diagonal Autoregressive Image Generation
von: Xu, Yijia, et al.
Veröffentlicht: (2025)
von: Xu, Yijia, et al.
Veröffentlicht: (2025)
Fast Autoregressive Video Generation with Diagonal Decoding
von: Ye, Yang, et al.
Veröffentlicht: (2025)
von: Ye, Yang, et al.
Veröffentlicht: (2025)
Make It Efficient: Dynamic Sparse Attention for Autoregressive Image Generation
von: Xiang, Xunzhi, et al.
Veröffentlicht: (2025)
von: Xiang, Xunzhi, et al.
Veröffentlicht: (2025)
Transformable Gaussian Reward Function for Socially-Aware Navigation with Deep Reinforcement Learning
von: Kim, Jinyeob, et al.
Veröffentlicht: (2024)
von: Kim, Jinyeob, et al.
Veröffentlicht: (2024)
DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer
von: Wu, Yecheng, et al.
Veröffentlicht: (2025)
von: Wu, Yecheng, et al.
Veröffentlicht: (2025)
InstGenIE: Generative Image Editing Made Efficient with Mask-aware Caching and Scheduling
von: Jiang, Xiaoxiao, et al.
Veröffentlicht: (2025)
von: Jiang, Xiaoxiao, et al.
Veröffentlicht: (2025)
Autoregressive Omni-Aware Outpainting for Open-Vocabulary 360-Degree Image Generation
von: Lu, Zhuqiang, et al.
Veröffentlicht: (2023)
von: Lu, Zhuqiang, et al.
Veröffentlicht: (2023)
MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
von: Liu, Haozhe, et al.
Veröffentlicht: (2024)
von: Liu, Haozhe, et al.
Veröffentlicht: (2024)
CacheFormer: High Attention-Based Segment Caching
von: Singh, Sushant, et al.
Veröffentlicht: (2025)
von: Singh, Sushant, et al.
Veröffentlicht: (2025)
Diagnosing Korean-Language LLM Political Bias via Census-Grounded Agent Simulation
von: Kang, Sungwoo
Veröffentlicht: (2026)
von: Kang, Sungwoo
Veröffentlicht: (2026)
STAG-CN: Spatio-Temporal Apiary Graph Convolutional Network for Disease Onset Prediction in Beehive Sensor Networks
von: Kang, Sungwoo
Veröffentlicht: (2026)
von: Kang, Sungwoo
Veröffentlicht: (2026)
LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation
von: Ahn, Jinwoo, et al.
Veröffentlicht: (2026)
von: Ahn, Jinwoo, et al.
Veröffentlicht: (2026)
REPrune: Channel Pruning via Kernel Representative Selection
von: Park, Mincheol, et al.
Veröffentlicht: (2024)
von: Park, Mincheol, et al.
Veröffentlicht: (2024)
Polynomial Context-Truncation Sensitivity in Autoregressive Language Models: Sequential Wyner-Ziv Bounds for KV Cache Compression
von: Kim, Munsik
Veröffentlicht: (2026)
von: Kim, Munsik
Veröffentlicht: (2026)
Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
von: Park, Joonhyung, et al.
Veröffentlicht: (2025)
von: Park, Joonhyung, et al.
Veröffentlicht: (2025)
UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
von: Kang, Wonjun, et al.
Veröffentlicht: (2025)
von: Kang, Wonjun, et al.
Veröffentlicht: (2025)
Using Span Queries to Optimize for Cache and Attention Locality
von: Castro, Paul, et al.
Veröffentlicht: (2025)
von: Castro, Paul, et al.
Veröffentlicht: (2025)
QuickMerge++: Fast Token Merging with Autoregressive Prior
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
Enabling Autoregressive Models to Fill In Masked Tokens
von: Israel, Daniel, et al.
Veröffentlicht: (2025)
von: Israel, Daniel, et al.
Veröffentlicht: (2025)
Modern Deep Learning Approaches for Cricket Shot Classification: A Comprehensive Baseline Study
von: Kang, Sungwoo
Veröffentlicht: (2025)
von: Kang, Sungwoo
Veröffentlicht: (2025)
Autoregressive Image Generation with Linear Complexity: A Spatial-Aware Decay Perspective
von: Mao, Yuxin, et al.
Veröffentlicht: (2025)
von: Mao, Yuxin, et al.
Veröffentlicht: (2025)
RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse
von: Geng, Yingsheng, et al.
Veröffentlicht: (2026)
von: Geng, Yingsheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
von: Jiang, Chaoyi, et al.
Veröffentlicht: (2024) -
CADC: Encoding User-Item Interactions for Compressing Recommendation Model Training Data
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2024) -
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2025) -
DEL: Context-Aware Dynamic Exit Layer for Efficient Self-Speculative Decoding
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2025) -
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
von: Ramachandran, Arun, et al.
Veröffentlicht: (2025)