Efficient Large Language Model Inference with Neural Block Linearization
Fuente:
arXiv
Salvato in:
| Autori principali: | Erdogan, Mete, Tonin, Francesco, Cevher, Volkan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Linear Attention for Efficient Bidirectional Sequence Modeling
di: Afzal, Arshia, et al.
Pubblicazione: (2025)
di: Afzal, Arshia, et al.
Pubblicazione: (2025)
Membership Inference Attacks against Large Vision-Language Models
di: Li, Zhan, et al.
Pubblicazione: (2024)
di: Li, Zhan, et al.
Pubblicazione: (2024)
MaD-Mix: Multi-Modal Data Mixtures via Latent Space Coupling for Vision-Language Model Training
di: Xie, Wanyun, et al.
Pubblicazione: (2026)
di: Xie, Wanyun, et al.
Pubblicazione: (2026)
Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
di: Xie, Wanyun, et al.
Pubblicazione: (2025)
di: Xie, Wanyun, et al.
Pubblicazione: (2025)
MT-NAM: An Efficient and Adaptive Model for Epileptic Seizure Detection
di: Afzal, Arshia, et al.
Pubblicazione: (2025)
di: Afzal, Arshia, et al.
Pubblicazione: (2025)
Bridging the Bosphorus: Advancing Turkish Large Language Models through Strategies for Low-Resource Language Adaptation and Benchmarking
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2024)
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2024)
Error Broadcast and Decorrelation as a Potential Artificial and Natural Learning Mechanism
di: Erdogan, Mete, et al.
Pubblicazione: (2025)
di: Erdogan, Mete, et al.
Pubblicazione: (2025)
Score Broadcast and Decorrelation: A General Framework for Broadcast-Based Credit Assignment
di: Uzun, Mustafa, et al.
Pubblicazione: (2026)
di: Uzun, Mustafa, et al.
Pubblicazione: (2026)
Single-pass Detection of Jailbreaking Input in Large Language Models
di: Candogan, Leyla Naz, et al.
Pubblicazione: (2025)
di: Candogan, Leyla Naz, et al.
Pubblicazione: (2025)
Adversarial Training for Defense Against Label Poisoning Attacks
di: Bal, Melis Ilayda, et al.
Pubblicazione: (2025)
di: Bal, Melis Ilayda, et al.
Pubblicazione: (2025)
REST: Efficient and Accelerated EEG Seizure Analysis through Residual State Updates
di: Afzal, Arshia, et al.
Pubblicazione: (2024)
di: Afzal, Arshia, et al.
Pubblicazione: (2024)
SVD Based Least Squares for X-Ray Pneumonia Classification Using Deep Features
di: Erdogan, Mete, et al.
Pubblicazione: (2025)
di: Erdogan, Mete, et al.
Pubblicazione: (2025)
Inference Optimization of Foundation Models on AI Accelerators
di: Park, Youngsuk, et al.
Pubblicazione: (2024)
di: Park, Youngsuk, et al.
Pubblicazione: (2024)
Revisiting Character-level Adversarial Attacks for Language Models
di: Rocamora, Elias Abad, et al.
Pubblicazione: (2024)
di: Rocamora, Elias Abad, et al.
Pubblicazione: (2024)
Tangent Space Fine-Tuning for Directional Preference Alignment in Large Language Models
di: Erdogan, Mete
Pubblicazione: (2026)
di: Erdogan, Mete
Pubblicazione: (2026)
Ascent Fails to Forget
di: Mavrothalassitis, Ioannis, et al.
Pubblicazione: (2025)
di: Mavrothalassitis, Ioannis, et al.
Pubblicazione: (2025)
Certified Robustness Under Bounded Levenshtein Distance
di: Rocamora, Elias Abad, et al.
Pubblicazione: (2025)
di: Rocamora, Elias Abad, et al.
Pubblicazione: (2025)
Addressing Label Shift in Distributed Learning via Entropy Regularization
di: Wu, Zhiyuan, et al.
Pubblicazione: (2025)
di: Wu, Zhiyuan, et al.
Pubblicazione: (2025)
BlockBatch: Multi-Scale Consensus Decoding for Efficient Diffusion Language Model Inference
di: Wu, Xiaoyou, et al.
Pubblicazione: (2026)
di: Wu, Xiaoyou, et al.
Pubblicazione: (2026)
Efficient local linearity regularization to overcome catastrophic overfitting
di: Rocamora, Elias Abad, et al.
Pubblicazione: (2024)
di: Rocamora, Elias Abad, et al.
Pubblicazione: (2024)
Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
di: Cheng, Yixin, et al.
Pubblicazione: (2024)
di: Cheng, Yixin, et al.
Pubblicazione: (2024)
Accelerating Spectral Clustering under Fairness Constraints
di: Tonin, Francesco, et al.
Pubblicazione: (2025)
di: Tonin, Francesco, et al.
Pubblicazione: (2025)
Rate optimal learning of equilibria from data
di: Freihaut, Till, et al.
Pubblicazione: (2025)
di: Freihaut, Till, et al.
Pubblicazione: (2025)
BLAST: Block-Level Adaptive Structured Matrices for Efficient Deep Neural Network Inference
di: Lee, Changwoo, et al.
Pubblicazione: (2024)
di: Lee, Changwoo, et al.
Pubblicazione: (2024)
Efficient Compositional Multi-tasking for On-device Large Language Models
di: Bohdal, Ondrej, et al.
Pubblicazione: (2025)
di: Bohdal, Ondrej, et al.
Pubblicazione: (2025)
LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models
di: Elesedy, Hayder, et al.
Pubblicazione: (2024)
di: Elesedy, Hayder, et al.
Pubblicazione: (2024)
Robust NAS under adversarial training: benchmark, theory, and beyond
di: Wu, Yongtao, et al.
Pubblicazione: (2024)
di: Wu, Yongtao, et al.
Pubblicazione: (2024)
Generalized Gradient Norm Clipping & Non-Euclidean $(L_0,L_1)$-Smoothness
di: Pethick, Thomas, et al.
Pubblicazione: (2025)
di: Pethick, Thomas, et al.
Pubblicazione: (2025)
Efficient Low Rank Attention for Long-Context Inference in Large Language Models
di: Li, Tenghui, et al.
Pubblicazione: (2025)
di: Li, Tenghui, et al.
Pubblicazione: (2025)
ES-dLLM: Efficient Inference for Diffusion Large Language Models by Early-Skipping
di: Zhu, Zijian, et al.
Pubblicazione: (2026)
di: Zhu, Zijian, et al.
Pubblicazione: (2026)
Clustering-driven Memory Compression for On-device Large Language Models
di: Bohdal, Ondrej, et al.
Pubblicazione: (2026)
di: Bohdal, Ondrej, et al.
Pubblicazione: (2026)
Block Transformer: Global-to-Local Language Modeling for Fast Inference
di: Ho, Namgyu, et al.
Pubblicazione: (2024)
di: Ho, Namgyu, et al.
Pubblicazione: (2024)
NeuralMatrix: Compute the Entire Neural Networks with Linear Matrix Operations for Efficient Inference
di: Sun, Ruiqi, et al.
Pubblicazione: (2023)
di: Sun, Ruiqi, et al.
Pubblicazione: (2023)
Fast Inference for Augmented Large Language Models
di: Shahout, Rana, et al.
Pubblicazione: (2024)
di: Shahout, Rana, et al.
Pubblicazione: (2024)
An Efficient Inference Framework for Early-exit Large Language Models
di: Miao, Ruijie, et al.
Pubblicazione: (2024)
di: Miao, Ruijie, et al.
Pubblicazione: (2024)
Model-Distributed Inference for Large Language Models at the Edge
di: Macario, Davide, et al.
Pubblicazione: (2025)
di: Macario, Davide, et al.
Pubblicazione: (2025)
Efficient Inference Using Large Language Models with Limited Human Data: Fine-Tuning then Rectification
di: Wang, Lei, et al.
Pubblicazione: (2025)
di: Wang, Lei, et al.
Pubblicazione: (2025)
GRASP: Deterministic argument ranking in interaction graphs
di: Misra, Diganta, et al.
Pubblicazione: (2026)
di: Misra, Diganta, et al.
Pubblicazione: (2026)
Linearizing Models for Efficient yet Robust Private Inference
di: Sarkar, Sreetama, et al.
Pubblicazione: (2024)
di: Sarkar, Sreetama, et al.
Pubblicazione: (2024)
Linearization Explains Fine-Tuning in Large Language Models
di: Afzal, Zahra Rahimi, et al.
Pubblicazione: (2026)
di: Afzal, Zahra Rahimi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Linear Attention for Efficient Bidirectional Sequence Modeling
di: Afzal, Arshia, et al.
Pubblicazione: (2025) -
Membership Inference Attacks against Large Vision-Language Models
di: Li, Zhan, et al.
Pubblicazione: (2024) -
MaD-Mix: Multi-Modal Data Mixtures via Latent Space Coupling for Vision-Language Model Training
di: Xie, Wanyun, et al.
Pubblicazione: (2026) -
Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
di: Xie, Wanyun, et al.
Pubblicazione: (2025) -
MT-NAM: An Efficient and Adaptive Model for Epileptic Seizure Detection
di: Afzal, Arshia, et al.
Pubblicazione: (2025)