GLASS: Global-Local Aggregation for Inference-time Sparsification of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Sattarifard, Amirmohsen, Lavasani, Sepehr, Zhang, Kunlin, Rajabpour, Amirhossein, Xu, Hanlin, Sun, Fengyu, Hassanpour, Negar, Gao, Chao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fantastic Multi-Task Gradient Updates and How to Find Them In a Cone
by: Hassanpour, Negar, et al.
Published: (2025)
by: Hassanpour, Negar, et al.
Published: (2025)
CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification
by: He, Junhui, et al.
Published: (2024)
by: He, Junhui, et al.
Published: (2024)
Fine-Tuning Causal LLMs for Text Classification: Embedding-Based vs. Instruction-Based Approaches
by: Yousefiramandi, Amirhossein, et al.
Published: (2025)
by: Yousefiramandi, Amirhossein, et al.
Published: (2025)
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
by: Gao, Zihan, et al.
Published: (2025)
by: Gao, Zihan, et al.
Published: (2025)
PixelMan: Consistent Object Editing with Diffusion Models via Pixel Manipulation and Generation
by: Jiang, Liyao, et al.
Published: (2024)
by: Jiang, Liyao, et al.
Published: (2024)
FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models
by: Weng, Zixuan, et al.
Published: (2026)
by: Weng, Zixuan, et al.
Published: (2026)
Location Not Found: Exposing Implicit Local and Global Biases in Multilingual LLMs
by: Mor-Lan, Guy, et al.
Published: (2026)
by: Mor-Lan, Guy, et al.
Published: (2026)
Block Transformer: Global-to-Local Language Modeling for Fast Inference
by: Ho, Namgyu, et al.
Published: (2024)
by: Ho, Namgyu, et al.
Published: (2024)
Diffusion Models: A Mathematical Introduction
by: Maleki, Sepehr, et al.
Published: (2025)
by: Maleki, Sepehr, et al.
Published: (2025)
Benchmarking Prompt Sensitivity in Large Language Models
by: Razavi, Amirhossein, et al.
Published: (2025)
by: Razavi, Amirhossein, et al.
Published: (2025)
Informed Routing in LLMs: Smarter Token-Level Computation for Faster Inference
by: Han, Chao, et al.
Published: (2025)
by: Han, Chao, et al.
Published: (2025)
StructRAG: Boosting Knowledge Intensive Reasoning of LLMs via Inference-time Hybrid Information Structurization
by: Li, Zhuoqun, et al.
Published: (2024)
by: Li, Zhuoqun, et al.
Published: (2024)
Structured Thinking Matters: Improving LLMs Generalization in Causal Inference Tasks
by: Sun, Wentao, et al.
Published: (2025)
by: Sun, Wentao, et al.
Published: (2025)
Nudging: Inference-time Alignment of LLMs via Guided Decoding
by: Fei, Yu, et al.
Published: (2024)
by: Fei, Yu, et al.
Published: (2024)
Regression-aware Inference with LLMs
by: Lukasik, Michal, et al.
Published: (2024)
by: Lukasik, Michal, et al.
Published: (2024)
CPTuning: Contrastive Prompt Tuning for Generative Relation Extraction
by: Duan, Jiaxin, et al.
Published: (2025)
by: Duan, Jiaxin, et al.
Published: (2025)
Persian-Phi: Efficient Cross-Lingual Adaptation of Compact LLMs via Curriculum Learning
by: Akhlaghi, Amir Mohammad, et al.
Published: (2025)
by: Akhlaghi, Amir Mohammad, et al.
Published: (2025)
Humans and LLMs Diverge on Probabilistic Inferences
by: Kamath, Gaurav, et al.
Published: (2026)
by: Kamath, Gaurav, et al.
Published: (2026)
Tandem Transformers for Inference Efficient LLMs
by: S, Aishwarya P, et al.
Published: (2024)
by: S, Aishwarya P, et al.
Published: (2024)
Draft-based Approximate Inference for LLMs
by: Galim, Kevin, et al.
Published: (2025)
by: Galim, Kevin, et al.
Published: (2025)
LLMs Are Prone to Fallacies in Causal Inference
by: Joshi, Nitish, et al.
Published: (2024)
by: Joshi, Nitish, et al.
Published: (2024)
Probing Association Biases in LLM Moderation Over-Sensitivity
by: Wang, Yuxin, et al.
Published: (2025)
by: Wang, Yuxin, et al.
Published: (2025)
Inference-time Alignment in Continuous Space
by: Yuan, Yige, et al.
Published: (2025)
by: Yuan, Yige, et al.
Published: (2025)
Not All Layers of LLMs Are Necessary During Inference
by: Fan, Siqi, et al.
Published: (2024)
by: Fan, Siqi, et al.
Published: (2024)
How Do LLMs Perform Two-Hop Reasoning in Context?
by: Guo, Tianyu, et al.
Published: (2025)
by: Guo, Tianyu, et al.
Published: (2025)
Localizing Persona Representations in LLMs
by: Cintas, Celia, et al.
Published: (2025)
by: Cintas, Celia, et al.
Published: (2025)
Fairness Evaluation and Inference Level Mitigation in LLMs
by: Nadeem, Afrozah, et al.
Published: (2025)
by: Nadeem, Afrozah, et al.
Published: (2025)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
by: Shen, Xuan, et al.
Published: (2023)
by: Shen, Xuan, et al.
Published: (2023)
Multi-Personality Generation of LLMs at Decoding-time
by: Chen, Rongxin, et al.
Published: (2025)
by: Chen, Rongxin, et al.
Published: (2025)
AggTruth: Contextual Hallucination Detection using Aggregated Attention Scores in LLMs
by: Matys, Piotr, et al.
Published: (2025)
by: Matys, Piotr, et al.
Published: (2025)
A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs
by: Srewa, Mahmoud, et al.
Published: (2025)
by: Srewa, Mahmoud, et al.
Published: (2025)
uTeBC-NLP at SemEval-2024 Task 9: Can LLMs be Lateral Thinkers?
by: Sadeghi, Pouya, et al.
Published: (2024)
by: Sadeghi, Pouya, et al.
Published: (2024)
Summarize Before You Speak with ARACH: A Training-Free Inference-Time Plug-In for Enhancing LLMs via Global Attention Reallocation
by: Wang, Jingtao, et al.
Published: (2026)
by: Wang, Jingtao, et al.
Published: (2026)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
by: Yadav, Prateek, et al.
Published: (2023)
by: Yadav, Prateek, et al.
Published: (2023)
Path-Consistency with Prefix Enhancement for Efficient Inference in LLMs
by: Zhu, Jiace, et al.
Published: (2024)
by: Zhu, Jiace, et al.
Published: (2024)
When Benchmarks Leak: Inference-Time Decontamination for LLMs
by: Chai, Jianzhe, et al.
Published: (2026)
by: Chai, Jianzhe, et al.
Published: (2026)
Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time
by: Chehade, Mohamad, et al.
Published: (2025)
by: Chehade, Mohamad, et al.
Published: (2025)
One Size Does Not Fit All: A Distribution-Aware Sparsification for More Precise Model Merging
by: Luo, Yingfeng, et al.
Published: (2025)
by: Luo, Yingfeng, et al.
Published: (2025)
Think Globally, Group Locally: Evaluating LLMs Using Multi-Lingual Word Grouping Games
by: Guerra-Solano, César, et al.
Published: (2025)
by: Guerra-Solano, César, et al.
Published: (2025)
Probing Multimodal Large Language Models for Global and Local Semantic Representations
by: Tao, Mingxu, et al.
Published: (2024)
by: Tao, Mingxu, et al.
Published: (2024)
Similar Items
-
Fantastic Multi-Task Gradient Updates and How to Find Them In a Cone
by: Hassanpour, Negar, et al.
Published: (2025) -
CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification
by: He, Junhui, et al.
Published: (2024) -
Fine-Tuning Causal LLMs for Text Classification: Embedding-Based vs. Instruction-Based Approaches
by: Yousefiramandi, Amirhossein, et al.
Published: (2025) -
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
by: Gao, Zihan, et al.
Published: (2025) -
PixelMan: Consistent Object Editing with Diffusion Models via Pixel Manipulation and Generation
by: Jiang, Liyao, et al.
Published: (2024)