Gespeichert in:
| Hauptverfasser: | Jha, Saurav, Hashemzadeh, Maryam, Pasand, Ali Saheb, Parviz, Ali, Lee, Min-Joong, Knyazev, Boris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2604.04356 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OPTIMA: Optimal One-shot Pruning for LLMs via Quadratic Programming Reconstruction
von: Mozaffari, Mohammad, et al.
Veröffentlicht: (2025)
von: Mozaffari, Mohammad, et al.
Veröffentlicht: (2025)
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
von: Yun, Vincent-Daniel, et al.
Veröffentlicht: (2026)
von: Yun, Vincent-Daniel, et al.
Veröffentlicht: (2026)
Scalable Graph Self-Supervised Learning
von: Pasand, Ali Saheb, et al.
Veröffentlicht: (2024)
von: Pasand, Ali Saheb, et al.
Veröffentlicht: (2024)
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
von: Israel, Daniel, et al.
Veröffentlicht: (2025)
von: Israel, Daniel, et al.
Veröffentlicht: (2025)
PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
von: Hourri, Younes, et al.
Veröffentlicht: (2025)
von: Hourri, Younes, et al.
Veröffentlicht: (2025)
Energy-Aware LLMs: A step towards sustainable AI for downstream applications
von: Tran, Nguyen Phuc, et al.
Veröffentlicht: (2025)
von: Tran, Nguyen Phuc, et al.
Veröffentlicht: (2025)
Throughput Optimization as a Strategic Lever in Large-Scale AI Systems: Evidence from Dataloader and Memory Profiling Innovations
von: Jha, Mayank
Veröffentlicht: (2026)
von: Jha, Mayank
Veröffentlicht: (2026)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
OMPILOT: Harnessing Transformer Models for Auto Parallelization to Shared Memory Computing Paradigms
von: Bhattacharjee, Arijit, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Arijit, et al.
Veröffentlicht: (2025)
Exchangeability in Neural Network and its Application to Dynamic Pruning
von: Pu, et al.
Veröffentlicht: (2025)
von: Pu, et al.
Veröffentlicht: (2025)
Stable Deep Reinforcement Learning via Isotropic Gaussian Representations
von: Pasand, Ali Saheb, et al.
Veröffentlicht: (2026)
von: Pasand, Ali Saheb, et al.
Veröffentlicht: (2026)
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
von: Chiang, Hung-Yueh, et al.
Veröffentlicht: (2025)
von: Chiang, Hung-Yueh, et al.
Veröffentlicht: (2025)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
von: Lin, Yujun, et al.
Veröffentlicht: (2024)
von: Lin, Yujun, et al.
Veröffentlicht: (2024)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
Evaluating the Efficacy of Foundational Models: Advancing Benchmarking Practices to Enhance Fine-Tuning Decision-Making
von: Amujo, Oluyemi Enoch, et al.
Veröffentlicht: (2024)
von: Amujo, Oluyemi Enoch, et al.
Veröffentlicht: (2024)
Data Efficacy for Language Model Training
von: Dai, Yalun, et al.
Veröffentlicht: (2025)
von: Dai, Yalun, et al.
Veröffentlicht: (2025)
Bench360: Benchmarking Local LLM Inference from 360 Degrees
von: Stuhlmann, Linus, et al.
Veröffentlicht: (2025)
von: Stuhlmann, Linus, et al.
Veröffentlicht: (2025)
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
von: Zheng, Wenhao, et al.
Veröffentlicht: (2025)
von: Zheng, Wenhao, et al.
Veröffentlicht: (2025)
Morpheme Boundary Detection & Grammatical Feature Prediction for Gujarati : Dataset & Model
von: Baxi, Jatayu, et al.
Veröffentlicht: (2021)
von: Baxi, Jatayu, et al.
Veröffentlicht: (2021)
An energy-based comparative analysis of common approaches to text classification in the Legal domain
von: Gultekin, Sinan, et al.
Veröffentlicht: (2023)
von: Gultekin, Sinan, et al.
Veröffentlicht: (2023)
ECCO: Evidence-Driven Causal Reasoning for Compiler Optimization
von: Pan, Haolin, et al.
Veröffentlicht: (2026)
von: Pan, Haolin, et al.
Veröffentlicht: (2026)
OptiSeq: Ordering Examples On-The-Fly for In-Context Learning
von: Bhope, Rahul Atul, et al.
Veröffentlicht: (2025)
von: Bhope, Rahul Atul, et al.
Veröffentlicht: (2025)
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
von: Zhao, Qihao, et al.
Veröffentlicht: (2024)
von: Zhao, Qihao, et al.
Veröffentlicht: (2024)
Model Compression and Efficient Inference for Large Language Models: A Survey
von: Wang, Wenxiao, et al.
Veröffentlicht: (2024)
von: Wang, Wenxiao, et al.
Veröffentlicht: (2024)
Enhancing Inference Efficiency of Large Language Models: Investigating Optimization Strategies and Architectural Innovations
von: Tyukin, Georgy
Veröffentlicht: (2024)
von: Tyukin, Georgy
Veröffentlicht: (2024)
SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
von: Mozaffari, Mohammad, et al.
Veröffentlicht: (2024)
von: Mozaffari, Mohammad, et al.
Veröffentlicht: (2024)
Generalizing Scaling Laws for Dense and Sparse Large Language Models
von: Hossain, Md Arafat, et al.
Veröffentlicht: (2025)
von: Hossain, Md Arafat, et al.
Veröffentlicht: (2025)
Egalitarian Gradient Descent: A Simple Approach to Accelerated Grokking
von: Pasand, Ali Saheb, et al.
Veröffentlicht: (2025)
von: Pasand, Ali Saheb, et al.
Veröffentlicht: (2025)
Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression
von: Ilin, Ivan, et al.
Veröffentlicht: (2025)
von: Ilin, Ivan, et al.
Veröffentlicht: (2025)
CodeRosetta: Pushing the Boundaries of Unsupervised Code Translation for Parallel Programming
von: TehraniJamsaz, Ali, et al.
Veröffentlicht: (2024)
von: TehraniJamsaz, Ali, et al.
Veröffentlicht: (2024)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
von: Shkolnik, Moran, et al.
Veröffentlicht: (2024)
von: Shkolnik, Moran, et al.
Veröffentlicht: (2024)
LLMs for Analog Circuit Design Continuum (ACDC)
von: Esfandiari, Yasaman, et al.
Veröffentlicht: (2025)
von: Esfandiari, Yasaman, et al.
Veröffentlicht: (2025)
Comment on paper: Position: Rethinking Post-Hoc Search-Based Neural Approaches for Solving Large-Scale Traveling Salesman Problems
von: Min, Yimeng
Veröffentlicht: (2024)
von: Min, Yimeng
Veröffentlicht: (2024)
\texttt{Range-Arithmetic}: Verifiable Deep Learning Inference on an Untrusted Party
von: Rahimi, Ali, et al.
Veröffentlicht: (2025)
von: Rahimi, Ali, et al.
Veröffentlicht: (2025)
MoEITS: A Green AI approach for simplifying MoE-LLMs
von: Balderas, Luis, et al.
Veröffentlicht: (2026)
von: Balderas, Luis, et al.
Veröffentlicht: (2026)
CDS4RAG: Cyclic Dual-Sequential Hyperparameter Optimization for RAG
von: Chen, Pengzhou, et al.
Veröffentlicht: (2026)
von: Chen, Pengzhou, et al.
Veröffentlicht: (2026)
LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems
von: Wu, Siyu, et al.
Veröffentlicht: (2026)
von: Wu, Siyu, et al.
Veröffentlicht: (2026)
Worst-Case Convergence Time of ML Algorithms via Extreme Value Theory
von: Tizpaz-Niari, Saeid, et al.
Veröffentlicht: (2024)
von: Tizpaz-Niari, Saeid, et al.
Veröffentlicht: (2024)
Regression Language Models for Code
von: Akhauri, Yash, et al.
Veröffentlicht: (2025)
von: Akhauri, Yash, et al.
Veröffentlicht: (2025)
Causal Unlearning in Collaborative Optimization: Exact and Approximate Influence Reversal under Adversarial Contributions
von: Mahdavi, Ali, et al.
Veröffentlicht: (2026)
von: Mahdavi, Ali, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
OPTIMA: Optimal One-shot Pruning for LLMs via Quadratic Programming Reconstruction
von: Mozaffari, Mohammad, et al.
Veröffentlicht: (2025) -
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
von: Yun, Vincent-Daniel, et al.
Veröffentlicht: (2026) -
Scalable Graph Self-Supervised Learning
von: Pasand, Ali Saheb, et al.
Veröffentlicht: (2024) -
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
von: Israel, Daniel, et al.
Veröffentlicht: (2025) -
PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
von: Hourri, Younes, et al.
Veröffentlicht: (2025)