Compression Laws for Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Sengupta, Ayan, Chaudhary, Siddhant, Chakraborty, Tanmoy |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
You Only Prune Once: Designing Calibration-Free Model Compression With Policy Learning
di: Sengupta, Ayan, et al.
Pubblicazione: (2025)
di: Sengupta, Ayan, et al.
Pubblicazione: (2025)
Value-Guided KV Compression for LLMs via Approximated CUR Decomposition
di: Sengupta, Ayan, et al.
Pubblicazione: (2025)
di: Sengupta, Ayan, et al.
Pubblicazione: (2025)
The Art of Scaling Test-Time Compute for Large Language Models
di: Agarwal, Aradhye, et al.
Pubblicazione: (2025)
di: Agarwal, Aradhye, et al.
Pubblicazione: (2025)
First Finish Search: Efficient Test-Time Scaling in Large Language Models
di: Agarwal, Aradhye, et al.
Pubblicazione: (2025)
di: Agarwal, Aradhye, et al.
Pubblicazione: (2025)
Understanding the Physics of Key-Value Cache Compression for LLMs through Attention Dynamics
di: Ananthanarayanan, Samhruth, et al.
Pubblicazione: (2026)
di: Ananthanarayanan, Samhruth, et al.
Pubblicazione: (2026)
How to Upscale Neural Networks with Scaling Law? A Survey and Practical Guidelines
di: Sengupta, Ayan, et al.
Pubblicazione: (2025)
di: Sengupta, Ayan, et al.
Pubblicazione: (2025)
Step-by-Step Unmasking for Parameter-Efficient Fine-tuning of Large Language Models
di: Agarwal, Aradhye, et al.
Pubblicazione: (2024)
di: Agarwal, Aradhye, et al.
Pubblicazione: (2024)
Persona-aware Generative Model for Code-mixed Language
di: Sengupta, Ayan, et al.
Pubblicazione: (2023)
di: Sengupta, Ayan, et al.
Pubblicazione: (2023)
Position: Enough of Scaling LLMs! Lets Focus on Downscaling
di: Goel, Yash, et al.
Pubblicazione: (2025)
di: Goel, Yash, et al.
Pubblicazione: (2025)
From Images to Words: Efficient Cross-Modal Knowledge Distillation to Language Models from Black-box Teachers
di: Sengupta, Ayan, et al.
Pubblicazione: (2026)
di: Sengupta, Ayan, et al.
Pubblicazione: (2026)
On the Generalization vs Fidelity Paradox in Knowledge Distillation
di: Ramesh, Suhas Kamasetty, et al.
Pubblicazione: (2025)
di: Ramesh, Suhas Kamasetty, et al.
Pubblicazione: (2025)
Latent Performance Profiling of Large Language Models
di: Chakraborty, Tanmoy, et al.
Pubblicazione: (2026)
di: Chakraborty, Tanmoy, et al.
Pubblicazione: (2026)
Information Anxiety in Large Language Models
di: Bajpai, Prasoon, et al.
Pubblicazione: (2024)
di: Bajpai, Prasoon, et al.
Pubblicazione: (2024)
MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing
di: Agarwal, Siddhant, et al.
Pubblicazione: (2024)
di: Agarwal, Siddhant, et al.
Pubblicazione: (2024)
Exposing Long-Tail Safety Failures in Large Language Models through Efficient Diverse Response Sampling
di: Hajra, Suvadeep, et al.
Pubblicazione: (2026)
di: Hajra, Suvadeep, et al.
Pubblicazione: (2026)
Temporally Consistent Factuality Probing for Large Language Models
di: Bajpai, Ashutosh, et al.
Pubblicazione: (2024)
di: Bajpai, Ashutosh, et al.
Pubblicazione: (2024)
Knowledge Planning in Large Language Models for Domain-Aligned Counseling Summarization
di: Srivastava, Aseem, et al.
Pubblicazione: (2024)
di: Srivastava, Aseem, et al.
Pubblicazione: (2024)
FLAME: Self-Supervised Low-Resource Taxonomy Expansion using Large Language Models
di: Mishra, Sahil, et al.
Pubblicazione: (2024)
di: Mishra, Sahil, et al.
Pubblicazione: (2024)
Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models
di: Nandi, Palash, et al.
Pubblicazione: (2025)
di: Nandi, Palash, et al.
Pubblicazione: (2025)
Mechanistic Behavior Editing of Language Models
di: Singh, Joykirat, et al.
Pubblicazione: (2024)
di: Singh, Joykirat, et al.
Pubblicazione: (2024)
Decoding Memes: Benchmarking Narrative Role Classification across Multilingual and Multimodal Models
di: Sharma, Shivam, et al.
Pubblicazione: (2025)
di: Sharma, Shivam, et al.
Pubblicazione: (2025)
Multilingual Needle in a Haystack: Investigating Long-Context Behavior of Multilingual Large Language Models
di: Hengle, Amey, et al.
Pubblicazione: (2024)
di: Hengle, Amey, et al.
Pubblicazione: (2024)
Robust and Efficient Fine-tuning of LLMs with Bayesian Reparameterization of Low-Rank Adaptation
di: Sengupta, Ayan, et al.
Pubblicazione: (2024)
di: Sengupta, Ayan, et al.
Pubblicazione: (2024)
HIDE and Seek: Detecting Hallucinations in Language Models via Decoupled Representations
di: Chatterjee, Anwoy, et al.
Pubblicazione: (2025)
di: Chatterjee, Anwoy, et al.
Pubblicazione: (2025)
coTherapist: A Behavior-Aligned Small Language Model to Support Mental Healthcare Experts
di: Adhikary, Prottay Kumar, et al.
Pubblicazione: (2026)
di: Adhikary, Prottay Kumar, et al.
Pubblicazione: (2026)
$\texttt{LM}^\texttt{2}$: A Simple Society of Language Models Solves Complex Reasoning
di: Juneja, Gurusha, et al.
Pubblicazione: (2024)
di: Juneja, Gurusha, et al.
Pubblicazione: (2024)
Distribution-Aware Companding Quantization of Large Language Models
di: Radhakrishnan, Athul, et al.
Pubblicazione: (2026)
di: Radhakrishnan, Athul, et al.
Pubblicazione: (2026)
Language Models can Exploit Cross-Task In-context Learning for Data-Scarce Novel Tasks
di: Chatterjee, Anwoy, et al.
Pubblicazione: (2024)
di: Chatterjee, Anwoy, et al.
Pubblicazione: (2024)
Multilingual Test-Time Scaling via Initial Thought Transfer
di: Bajpai, Prasoon, et al.
Pubblicazione: (2025)
di: Bajpai, Prasoon, et al.
Pubblicazione: (2025)
Harmonizing Code-mixed Conversations: Personality-assisted Code-mixed Response Generation in Dialogues
di: Kumar, Shivani, et al.
Pubblicazione: (2024)
di: Kumar, Shivani, et al.
Pubblicazione: (2024)
Self-Evolved Preference Optimization for Enhancing Mathematical Reasoning in Small Language Models
di: Singh, Joykirat, et al.
Pubblicazione: (2025)
di: Singh, Joykirat, et al.
Pubblicazione: (2025)
POSIX: A Prompt Sensitivity Index For Large Language Models
di: Chatterjee, Anwoy, et al.
Pubblicazione: (2024)
di: Chatterjee, Anwoy, et al.
Pubblicazione: (2024)
Multilingual LLMs Inherently Reward In-Language Time-Sensitive Semantic Alignment for Low-Resource Languages
di: Bajpai, Ashutosh, et al.
Pubblicazione: (2024)
di: Bajpai, Ashutosh, et al.
Pubblicazione: (2024)
Small Language Models Fine-tuned to Coordinate Larger Language Models improve Complex Reasoning
di: Juneja, Gurusha, et al.
Pubblicazione: (2023)
di: Juneja, Gurusha, et al.
Pubblicazione: (2023)
Are Large Language Models In-Context Personalized Summarizers? Get an iCOPERNICUS Test Done!
di: Patel, Divya, et al.
Pubblicazione: (2024)
di: Patel, Divya, et al.
Pubblicazione: (2024)
Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models
di: Sahnan, Dhruv, et al.
Pubblicazione: (2026)
di: Sahnan, Dhruv, et al.
Pubblicazione: (2026)
On the Brittle Foundations of ReAct Prompting for Agentic Large Language Models
di: Verma, Mudit, et al.
Pubblicazione: (2024)
di: Verma, Mudit, et al.
Pubblicazione: (2024)
Towards Privacy-aware Mental Health AI Models: Advances, Challenges, and Opportunities
di: Mandal, Aishik, et al.
Pubblicazione: (2025)
di: Mandal, Aishik, et al.
Pubblicazione: (2025)
Multilingual Language Models Encode Script Over Linguistic Structure
di: Verma, Aastha A K, et al.
Pubblicazione: (2026)
di: Verma, Aastha A K, et al.
Pubblicazione: (2026)
Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?
di: Bajpai, Ashutosh, et al.
Pubblicazione: (2025)
di: Bajpai, Ashutosh, et al.
Pubblicazione: (2025)
Documenti analoghi
-
You Only Prune Once: Designing Calibration-Free Model Compression With Policy Learning
di: Sengupta, Ayan, et al.
Pubblicazione: (2025) -
Value-Guided KV Compression for LLMs via Approximated CUR Decomposition
di: Sengupta, Ayan, et al.
Pubblicazione: (2025) -
The Art of Scaling Test-Time Compute for Large Language Models
di: Agarwal, Aradhye, et al.
Pubblicazione: (2025) -
First Finish Search: Efficient Test-Time Scaling in Large Language Models
di: Agarwal, Aradhye, et al.
Pubblicazione: (2025) -
Understanding the Physics of Key-Value Cache Compression for LLMs through Attention Dynamics
di: Ananthanarayanan, Samhruth, et al.
Pubblicazione: (2026)