Radar: Fast Long-Context Decoding for Any Transformer
Fuente:
arXiv
Salvato in:
| Autori principali: | Hao, Yongchang, Zhai, Mengyao, Hajimirsadeghi, Hossein, Hosseini, Sepidehsadat, Tung, Frederick |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Prompting-based Temporal Domain Generalization
di: Hosseini, Sepidehsadat, et al.
Pubblicazione: (2023)
di: Hosseini, Sepidehsadat, et al.
Pubblicazione: (2023)
FairNVT: Improving Fairness via Noise Injection in Vision Transformers
di: Tang, Qiaoyue, et al.
Pubblicazione: (2026)
di: Tang, Qiaoyue, et al.
Pubblicazione: (2026)
You Need Reasoning to Learn Reasoning: The Limitations of Label-Free RL in Weak Base Models
di: Roy, Shuvendu, et al.
Pubblicazione: (2025)
di: Roy, Shuvendu, et al.
Pubblicazione: (2025)
Tree Cross Attention
di: Feng, Leo, et al.
Pubblicazione: (2023)
di: Feng, Leo, et al.
Pubblicazione: (2023)
Were RNNs All We Needed?
di: Feng, Leo, et al.
Pubblicazione: (2024)
di: Feng, Leo, et al.
Pubblicazione: (2024)
Memory Efficient Neural Processes via Constant Memory Attention Block
di: Feng, Leo, et al.
Pubblicazione: (2023)
di: Feng, Leo, et al.
Pubblicazione: (2023)
Attention as an RNN
di: Feng, Leo, et al.
Pubblicazione: (2024)
di: Feng, Leo, et al.
Pubblicazione: (2024)
Cactus: Accelerating Auto-Regressive Decoding with Constrained Acceptance Speculative Sampling
di: Hao, Yongchang, et al.
Pubblicazione: (2026)
di: Hao, Yongchang, et al.
Pubblicazione: (2026)
TabReason: A Reinforcement Learning-Enhanced Reasoning LLM for Explainable Tabular Data Prediction
di: Xu, Tommy, et al.
Pubblicazione: (2025)
di: Xu, Tommy, et al.
Pubblicazione: (2025)
ContextPilot: Fast Long-Context Inference via Context Reuse
di: Jiang, Yinsicheng, et al.
Pubblicazione: (2025)
di: Jiang, Yinsicheng, et al.
Pubblicazione: (2025)
Forget Sharpness: Perturbed Forgetting of Model Biases Within SAM Dynamics
di: Vani, Ankit, et al.
Pubblicazione: (2024)
di: Vani, Ankit, et al.
Pubblicazione: (2024)
SPINT: Spatial Permutation-Invariant Neural Transformer for Consistent Intracortical Motor Decoding
di: Le, Trung, et al.
Pubblicazione: (2025)
di: Le, Trung, et al.
Pubblicazione: (2025)
Flora: Low-Rank Adapters Are Secretly Gradient Compressors
di: Hao, Yongchang, et al.
Pubblicazione: (2024)
di: Hao, Yongchang, et al.
Pubblicazione: (2024)
NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks
di: Hao, Yongchang, et al.
Pubblicazione: (2024)
di: Hao, Yongchang, et al.
Pubblicazione: (2024)
DEL: Context-Aware Dynamic Exit Layer for Efficient Self-Speculative Decoding
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2025)
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2025)
CADET: Context-Conditioned Ads CTR Prediction With a Decoder-Only Transformer
di: Pardoe, David, et al.
Pubblicazione: (2026)
di: Pardoe, David, et al.
Pubblicazione: (2026)
Core Context Aware Transformers for Long Context Language Modeling
di: Chen, Yaofo, et al.
Pubblicazione: (2024)
di: Chen, Yaofo, et al.
Pubblicazione: (2024)
Ginger: An Efficient Curvature Approximation with Linear Complexity for General Neural Networks
di: Hao, Yongchang, et al.
Pubblicazione: (2024)
di: Hao, Yongchang, et al.
Pubblicazione: (2024)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
di: Yang, Penghui, et al.
Pubblicazione: (2025)
di: Yang, Penghui, et al.
Pubblicazione: (2025)
Functional Interpolation for Relative Positions Improves Long Context Transformers
di: Li, Shanda, et al.
Pubblicazione: (2023)
di: Li, Shanda, et al.
Pubblicazione: (2023)
AnyLoss: Transforming Classification Metrics into Loss Functions
di: Han, Doheon, et al.
Pubblicazione: (2024)
di: Han, Doheon, et al.
Pubblicazione: (2024)
SpecPV: Improving Self-Speculative Decoding for Long-Context Generation via Partial Verification
di: Tan, Zhendong, et al.
Pubblicazione: (2025)
di: Tan, Zhendong, et al.
Pubblicazione: (2025)
Revolutionizing Traffic Management with AI-Powered Machine Vision: A Step Toward Smart Cities
di: DolatAbadi, Seyed Hossein Hosseini, et al.
Pubblicazione: (2025)
di: DolatAbadi, Seyed Hossein Hosseini, et al.
Pubblicazione: (2025)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
di: Jo, Dongwon, et al.
Pubblicazione: (2025)
di: Jo, Dongwon, et al.
Pubblicazione: (2025)
Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
di: Shyam, Vasudev, et al.
Pubblicazione: (2024)
di: Shyam, Vasudev, et al.
Pubblicazione: (2024)
Timer-XL: Long-Context Transformers for Unified Time Series Forecasting
di: Liu, Yong, et al.
Pubblicazione: (2024)
di: Liu, Yong, et al.
Pubblicazione: (2024)
Variational Linear Attention: Stable Associative Memory for Long-Context Transformers
di: Pandey, Vishal, et al.
Pubblicazione: (2026)
di: Pandey, Vishal, et al.
Pubblicazione: (2026)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2025)
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2025)
Efficient Solutions For An Intriguing Failure of LLMs: Long Context Window Does Not Mean LLMs Can Analyze Long Sequences Flawlessly
di: Hosseini, Peyman, et al.
Pubblicazione: (2024)
di: Hosseini, Peyman, et al.
Pubblicazione: (2024)
Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long Generation
di: Ren, Liliang, et al.
Pubblicazione: (2025)
di: Ren, Liliang, et al.
Pubblicazione: (2025)
Reviving Any-Subset Autoregressive Models with Principled Parallel Sampling and Speculative Decoding
di: Guo, Gabe, et al.
Pubblicazione: (2025)
di: Guo, Gabe, et al.
Pubblicazione: (2025)
Attention in Constant Time: Vashista Sparse Attention for Long-Context Decoding with Exponential Guarantees
di: Nobaub, Vashista
Pubblicazione: (2026)
di: Nobaub, Vashista
Pubblicazione: (2026)
Scaling Limits of Long-Context Transformers
di: Bruno, Giuseppe, et al.
Pubblicazione: (2026)
di: Bruno, Giuseppe, et al.
Pubblicazione: (2026)
NExT-GPT: Any-to-Any Multimodal LLM
di: Wu, Shengqiong, et al.
Pubblicazione: (2023)
di: Wu, Shengqiong, et al.
Pubblicazione: (2023)
Fast Inference via Hierarchical Speculative Decoding
di: Mohri, Clara, et al.
Pubblicazione: (2025)
di: Mohri, Clara, et al.
Pubblicazione: (2025)
Data Efficient Any Transformer-to-Mamba Distillation via Attention Bridge
di: Wang, Penghao, et al.
Pubblicazione: (2025)
di: Wang, Penghao, et al.
Pubblicazione: (2025)
Fast Inference with Kronecker-Sparse Matrices
di: Gonon, Antoine, et al.
Pubblicazione: (2024)
di: Gonon, Antoine, et al.
Pubblicazione: (2024)
Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers
di: Song, Woomin, et al.
Pubblicazione: (2025)
di: Song, Woomin, et al.
Pubblicazione: (2025)
Short Data, Long Context: Distilling Positional Knowledge in Transformers
di: Huber, Patrick, et al.
Pubblicazione: (2026)
di: Huber, Patrick, et al.
Pubblicazione: (2026)
MoCap2Radar: A Spatiotemporal Transformer for Synthesizing Micro-Doppler Radar Signatures from Motion Capture
di: Chen, Kevin, et al.
Pubblicazione: (2025)
di: Chen, Kevin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Prompting-based Temporal Domain Generalization
di: Hosseini, Sepidehsadat, et al.
Pubblicazione: (2023) -
FairNVT: Improving Fairness via Noise Injection in Vision Transformers
di: Tang, Qiaoyue, et al.
Pubblicazione: (2026) -
You Need Reasoning to Learn Reasoning: The Limitations of Label-Free RL in Weak Base Models
di: Roy, Shuvendu, et al.
Pubblicazione: (2025) -
Tree Cross Attention
di: Feng, Leo, et al.
Pubblicazione: (2023) -
Were RNNs All We Needed?
di: Feng, Leo, et al.
Pubblicazione: (2024)