Systematic Evaluation of Optimization Techniques for Long-Context Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ahmed, Ammar, Di, Sheng, Cappello, Franck, Liu, Zirui, Han, Jingoo, Anwar, Ali |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
by: Liu, Zirui, et al.
Published: (2024)
by: Liu, Zirui, et al.
Published: (2024)
Priority Sampling of Large Language Models for Compilers
by: Grubisic, Dejan, et al.
Published: (2024)
by: Grubisic, Dejan, et al.
Published: (2024)
DeepCQ: General-Purpose Deep-Surrogate Framework for Lossy Compression Quality Prediction
by: Mumenin, Khondoker Mirazul, et al.
Published: (2025)
by: Mumenin, Khondoker Mirazul, et al.
Published: (2025)
Enhancing Inference Efficiency of Large Language Models: Investigating Optimization Strategies and Architectural Innovations
by: Tyukin, Georgy
Published: (2024)
by: Tyukin, Georgy
Published: (2024)
ModeSwitch-LLM: A Lightweight Phase-Aware Controller for Cross-Mode LLM Inference on a Single GPU
by: Sunesh, Aman, et al.
Published: (2026)
by: Sunesh, Aman, et al.
Published: (2026)
Data Efficacy for Language Model Training
by: Dai, Yalun, et al.
Published: (2025)
by: Dai, Yalun, et al.
Published: (2025)
Model Compression and Efficient Inference for Large Language Models: A Survey
by: Wang, Wenxiao, et al.
Published: (2024)
by: Wang, Wenxiao, et al.
Published: (2024)
OptiSeq: Ordering Examples On-The-Fly for In-Context Learning
by: Bhope, Rahul Atul, et al.
Published: (2025)
by: Bhope, Rahul Atul, et al.
Published: (2025)
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
by: Zheng, Wenhao, et al.
Published: (2025)
by: Zheng, Wenhao, et al.
Published: (2025)
Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading
by: Maurya, Avinash, et al.
Published: (2024)
by: Maurya, Avinash, et al.
Published: (2024)
Evaluating the Efficacy of Foundational Models: Advancing Benchmarking Practices to Enhance Fine-Tuning Decision-Making
by: Amujo, Oluyemi Enoch, et al.
Published: (2024)
by: Amujo, Oluyemi Enoch, et al.
Published: (2024)
FLuRKA: Fast and accurate unified Low-Rank & Kernel Attention
by: Gupta, Ahan, et al.
Published: (2023)
by: Gupta, Ahan, et al.
Published: (2023)
Block Sparse Flash Attention
by: Ohayon, Daniel, et al.
Published: (2025)
by: Ohayon, Daniel, et al.
Published: (2025)
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
by: Chen, Feiyang, et al.
Published: (2025)
by: Chen, Feiyang, et al.
Published: (2025)
REAM: Merging Improves Pruning of Experts in LLMs
by: Jha, Saurav, et al.
Published: (2026)
by: Jha, Saurav, et al.
Published: (2026)
Flex Attention: A Programming Model for Generating Optimized Attention Kernels
by: Dong, Juechu, et al.
Published: (2024)
by: Dong, Juechu, et al.
Published: (2024)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
by: Zandieh, Amir, et al.
Published: (2024)
by: Zandieh, Amir, et al.
Published: (2024)
EPIC: Efficient Position-Independent Caching for Serving Large Language Models
by: Hu, Junhao, et al.
Published: (2024)
by: Hu, Junhao, et al.
Published: (2024)
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
by: Zhao, Qihao, et al.
Published: (2024)
by: Zhao, Qihao, et al.
Published: (2024)
Regression Language Models for Code
by: Akhauri, Yash, et al.
Published: (2025)
by: Akhauri, Yash, et al.
Published: (2025)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
by: Yang, Shang, et al.
Published: (2025)
by: Yang, Shang, et al.
Published: (2025)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
by: Lin, Yujun, et al.
Published: (2024)
by: Lin, Yujun, et al.
Published: (2024)
OMPILOT: Harnessing Transformer Models for Auto Parallelization to Shared Memory Computing Paradigms
by: Bhattacharjee, Arijit, et al.
Published: (2025)
by: Bhattacharjee, Arijit, et al.
Published: (2025)
Morpheme Boundary Detection & Grammatical Feature Prediction for Gujarati : Dataset & Model
by: Baxi, Jatayu, et al.
Published: (2021)
by: Baxi, Jatayu, et al.
Published: (2021)
The Next 700 ML-Enabled Compiler Optimizations
by: VenkataKeerthy, S., et al.
Published: (2023)
by: VenkataKeerthy, S., et al.
Published: (2023)
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
by: Chiang, Hung-Yueh, et al.
Published: (2025)
by: Chiang, Hung-Yueh, et al.
Published: (2025)
LOOPerSet: A Large-Scale Dataset for Data-Driven Polyhedral Compiler Optimization
by: Merouani, Massinissa, et al.
Published: (2025)
by: Merouani, Massinissa, et al.
Published: (2025)
Reconstructing Biological Pathways by Applying Selective Incremental Learning to (Very) Small Language Models
by: Saha, Pranta, et al.
Published: (2025)
by: Saha, Pranta, et al.
Published: (2025)
LFED: A Literary Fiction Evaluation Dataset for Large Language Models
by: Yu, Linhao, et al.
Published: (2024)
by: Yu, Linhao, et al.
Published: (2024)
CDS4RAG: Cyclic Dual-Sequential Hyperparameter Optimization for RAG
by: Chen, Pengzhou, et al.
Published: (2026)
by: Chen, Pengzhou, et al.
Published: (2026)
Agentic Auto-Scheduling: An Experimental Study of LLM-Guided Loop Optimization
by: Merouani, Massinissa, et al.
Published: (2025)
by: Merouani, Massinissa, et al.
Published: (2025)
Energy-Aware LLMs: A step towards sustainable AI for downstream applications
by: Tran, Nguyen Phuc, et al.
Published: (2025)
by: Tran, Nguyen Phuc, et al.
Published: (2025)
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
by: Zhou, Zikai, et al.
Published: (2025)
by: Zhou, Zikai, et al.
Published: (2025)
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes
by: Hendria, Willy Fitra
Published: (2026)
by: Hendria, Willy Fitra
Published: (2026)
Bench360: Benchmarking Local LLM Inference from 360 Degrees
by: Stuhlmann, Linus, et al.
Published: (2025)
by: Stuhlmann, Linus, et al.
Published: (2025)
An energy-based comparative analysis of common approaches to text classification in the Legal domain
by: Gultekin, Sinan, et al.
Published: (2023)
by: Gultekin, Sinan, et al.
Published: (2023)
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
by: Israel, Daniel, et al.
Published: (2025)
by: Israel, Daniel, et al.
Published: (2025)
LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems
by: Wu, Siyu, et al.
Published: (2026)
by: Wu, Siyu, et al.
Published: (2026)
WCDT: Systematic WCET Optimization for Decision Tree Implementations
by: Hölscher, Nils, et al.
Published: (2025)
by: Hölscher, Nils, et al.
Published: (2025)
ProfilingAgent: Profiling-Guided Agentic Reasoning for Adaptive Model Optimization
by: Jafari, Sadegh, et al.
Published: (2025)
by: Jafari, Sadegh, et al.
Published: (2025)
Similar Items
-
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
by: Liu, Zirui, et al.
Published: (2024) -
Priority Sampling of Large Language Models for Compilers
by: Grubisic, Dejan, et al.
Published: (2024) -
DeepCQ: General-Purpose Deep-Surrogate Framework for Lossy Compression Quality Prediction
by: Mumenin, Khondoker Mirazul, et al.
Published: (2025) -
Enhancing Inference Efficiency of Large Language Models: Investigating Optimization Strategies and Architectural Innovations
by: Tyukin, Georgy
Published: (2024) -
ModeSwitch-LLM: A Lightweight Phase-Aware Controller for Cross-Mode LLM Inference on a Single GPU
by: Sunesh, Aman, et al.
Published: (2026)