LAWCAT: Efficient Distillation from Quadratic to Linear Attention with Convolution across Tokens for Long Context Modeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zeyu, Kundu, Souvik, Jiang, Lianghao, Li, Anni, Ronanki, Srikanth, Bodapati, Sravan, Datta, Gourav, Beerel, Peter A. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AFLoRA: Adaptive Freezing of Low Rank Adaptation in Parameter Efficient Fine-Tuning of Large Models
von: Liu, Zeyu, et al.
Veröffentlicht: (2024)
von: Liu, Zeyu, et al.
Veröffentlicht: (2024)
LMUFormer: Low Complexity Yet Powerful Spiking Model With Legendre Memory Units
von: Liu, Zeyu, et al.
Veröffentlicht: (2024)
von: Liu, Zeyu, et al.
Veröffentlicht: (2024)
MaskVD: Region Masking for Efficient Video Object Detection
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2024)
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2024)
Linearizing Models for Efficient yet Robust Private Inference
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2024)
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2024)
DCTX-Conformer: Dynamic context carry-over for low latency unified streaming and non-streaming Conformer ASR
von: Huybrechts, Goeric, et al.
Veröffentlicht: (2023)
von: Huybrechts, Goeric, et al.
Veröffentlicht: (2023)
Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers
von: Song, Woomin, et al.
Veröffentlicht: (2025)
von: Song, Woomin, et al.
Veröffentlicht: (2025)
Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
von: Du, Yufeng, et al.
Veröffentlicht: (2025)
von: Du, Yufeng, et al.
Veröffentlicht: (2025)
FixPix: Fixing Bad Pixels using Deep Learning
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2023)
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2023)
RedVTP: Training-Free Acceleration of Diffusion Vision-Language Models Inference via Masked Token-Guided Visual Token Pruning
von: Xu, Jingqi, et al.
Veröffentlicht: (2025)
von: Xu, Jingqi, et al.
Veröffentlicht: (2025)
RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably
von: Du, Yufeng, et al.
Veröffentlicht: (2026)
von: Du, Yufeng, et al.
Veröffentlicht: (2026)
On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention
von: Ro, Yeonju, et al.
Veröffentlicht: (2025)
von: Ro, Yeonju, et al.
Veröffentlicht: (2025)
Block Selective Reprogramming for On-device Training of Vision Transformers
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2024)
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2024)
Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark
von: Huybrechts, Goeric, et al.
Veröffentlicht: (2025)
von: Huybrechts, Goeric, et al.
Veröffentlicht: (2025)
Energy-Efficient & Real-Time Computer Vision with Intelligent Skipping via Reconfigurable CMOS Image Sensors
von: Kaiser, Md Abdullah-Al, et al.
Veröffentlicht: (2024)
von: Kaiser, Md Abdullah-Al, et al.
Veröffentlicht: (2024)
Toward High Performance, Programmable Extreme-Edge Intelligence for Neuromorphic Vision Sensors utilizing Magnetic Domain Wall Motion-based MTJ
von: Kaiser, Md Abdullah-Al, et al.
Veröffentlicht: (2024)
von: Kaiser, Md Abdullah-Al, et al.
Veröffentlicht: (2024)
Technology-Circuit-Algorithm Tri-Design for Processing-in-Pixel-in-Memory (P2M)
von: Kaiser, Md Abdullah-Al, et al.
Veröffentlicht: (2023)
von: Kaiser, Md Abdullah-Al, et al.
Veröffentlicht: (2023)
Mitigating Hallucinations in Vision-Language Models through Image-Guided Head Suppression
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2025)
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2025)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
von: Chen, Yingfa, et al.
Veröffentlicht: (2026)
von: Chen, Yingfa, et al.
Veröffentlicht: (2026)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2025)
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2025)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
von: Jo, Dongwon, et al.
Veröffentlicht: (2026)
von: Jo, Dongwon, et al.
Veröffentlicht: (2026)
Region Masking to Accelerate Video Processing on Neuromorphic Hardware
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2025)
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2025)
Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers
von: Horton, Mark, et al.
Veröffentlicht: (2025)
von: Horton, Mark, et al.
Veröffentlicht: (2025)
Short-Long Convolutions Help Hardware-Efficient Linear Attention to Focus on Long Sequences
von: Liu, Zicheng, et al.
Veröffentlicht: (2024)
von: Liu, Zicheng, et al.
Veröffentlicht: (2024)
ReDistill: Residual Encoded Distillation for Peak Memory Reduction of CNNs
von: Chen, Fang, et al.
Veröffentlicht: (2024)
von: Chen, Fang, et al.
Veröffentlicht: (2024)
Adaptive Video Understanding Agent: Enhancing efficiency with dynamic frame sampling and feedback-driven reasoning
von: Jeoung, Sullam, et al.
Veröffentlicht: (2024)
von: Jeoung, Sullam, et al.
Veröffentlicht: (2024)
ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language Models
von: Elangovan, Aparna, et al.
Veröffentlicht: (2024)
von: Elangovan, Aparna, et al.
Veröffentlicht: (2024)
Learning Scalable Temporal Representations in Spiking Neural Networks Without Labels
von: Zhou, Chengwei, et al.
Veröffentlicht: (2025)
von: Zhou, Chengwei, et al.
Veröffentlicht: (2025)
Facilitating Trustworthy Human-Agent Collaboration in LLM-based Multi-Agent System oriented Software Engineering
von: Ronanki, Krishna
Veröffentlicht: (2025)
von: Ronanki, Krishna
Veröffentlicht: (2025)
Assortment of Attention Heads: Accelerating Federated PEFT with Head Pruning and Strategic Client Selection
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025)
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025)
Unitary Quadratic Quantum Gravity in 4D
von: Kumar, K. Sravan, et al.
Veröffentlicht: (2026)
von: Kumar, K. Sravan, et al.
Veröffentlicht: (2026)
Sequential Editing for Lifelong Training of Speech Recognition Models
von: Kulshreshtha, Devang, et al.
Veröffentlicht: (2024)
von: Kulshreshtha, Devang, et al.
Veröffentlicht: (2024)
SeRA: Self-Reviewing and Alignment of Large Language Models using Implicit Reward Margins
von: Ko, Jongwoo, et al.
Veröffentlicht: (2024)
von: Ko, Jongwoo, et al.
Veröffentlicht: (2024)
Think Clearly: Improving Reasoning via Redundant Token Pruning
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
One Token Away from Collapse: The Fragility of Instruction-Tuned Helpfulness
von: Potraghloo, Erfan Baghaei, et al.
Veröffentlicht: (2026)
von: Potraghloo, Erfan Baghaei, et al.
Veröffentlicht: (2026)
FPAN: Mitigating Replication in Diffusion Models through the Fine-Grained Probabilistic Addition of Noise to Token Embeddings
von: Xu, Jingqi, et al.
Veröffentlicht: (2025)
von: Xu, Jingqi, et al.
Veröffentlicht: (2025)
CAT: Circular-Convolutional Attention for Sub-Quadratic Transformers
von: Yamada, Yoshihiro
Veröffentlicht: (2025)
von: Yamada, Yoshihiro
Veröffentlicht: (2025)
Branch Landing: Bloom Filter-Based Source Authorization for Forward-Edge CFI on RISC-V
von: Wu, You, et al.
Veröffentlicht: (2026)
von: Wu, You, et al.
Veröffentlicht: (2026)
In-Context Learning in Linear vs. Quadratic Attention Models: An Empirical Study on Regression Tasks
von: Goel, Ayush, et al.
Veröffentlicht: (2026)
von: Goel, Ayush, et al.
Veröffentlicht: (2026)
Opto-ViT: Architecting a Near-Sensor Region of Interest-Aware Vision Transformer Accelerator with Silicon Photonics
von: Morsali, Mehrdad, et al.
Veröffentlicht: (2025)
von: Morsali, Mehrdad, et al.
Veröffentlicht: (2025)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
von: MiniCPM Team, et al.
Veröffentlicht: (2026)
von: MiniCPM Team, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
AFLoRA: Adaptive Freezing of Low Rank Adaptation in Parameter Efficient Fine-Tuning of Large Models
von: Liu, Zeyu, et al.
Veröffentlicht: (2024) -
LMUFormer: Low Complexity Yet Powerful Spiking Model With Legendre Memory Units
von: Liu, Zeyu, et al.
Veröffentlicht: (2024) -
MaskVD: Region Masking for Efficient Video Object Detection
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2024) -
Linearizing Models for Efficient yet Robust Private Inference
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2024) -
DCTX-Conformer: Dynamic context carry-over for low latency unified streaming and non-streaming Conformer ASR
von: Huybrechts, Goeric, et al.
Veröffentlicht: (2023)