Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques
Fuente:
arXiv
Guardado en:
| Autores principales: | Girija, Sanjay Surendranath, Kapoor, Shashank, Arora, Lakshit, Pradhan, Dipen, Raj, Aman, Shetgaonkar, Ankit |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Explainable Artificial Intelligence Techniques for Software Development Lifecycle: A Phase-specific Survey
por: Arora, Lakshit, et al.
Publicado: (2025)
por: Arora, Lakshit, et al.
Publicado: (2025)
AI and Generative AI Transforming Disaster Management: A Survey of Damage Assessment and Response Techniques
por: Raj, Aman, et al.
Publicado: (2025)
por: Raj, Aman, et al.
Publicado: (2025)
Adversarial Attacks in Multimodal Systems: A Practitioner's Survey
por: Kapoor, Shashank, et al.
Publicado: (2025)
por: Kapoor, Shashank, et al.
Publicado: (2025)
Opportunities and Applications of GenAI in Smart Cities: A User-Centric Survey
por: Shetgaonkar, Ankit, et al.
Publicado: (2025)
por: Shetgaonkar, Ankit, et al.
Publicado: (2025)
Mitigating Clinician Information Overload: Generative AI for Integrated EHR and RPM Data Analysis
por: Shetgaonkar, Ankit, et al.
Publicado: (2025)
por: Shetgaonkar, Ankit, et al.
Publicado: (2025)
A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models
por: Khoshnoodi, Mahsa, et al.
Publicado: (2024)
por: Khoshnoodi, Mahsa, et al.
Publicado: (2024)
Pre-training LLMs using human-like development data corpus
por: Bhardwaj, Khushi, et al.
Publicado: (2023)
por: Bhardwaj, Khushi, et al.
Publicado: (2023)
From Intentions to Techniques: A Comprehensive Taxonomy and Challenges in Text Watermarking for Large Language Models
por: Lalai, Harsh Nishant, et al.
Publicado: (2024)
por: Lalai, Harsh Nishant, et al.
Publicado: (2024)
A Survey on Prompting Techniques in LLMs
por: Bhandari, Prabin
Publicado: (2023)
por: Bhandari, Prabin
Publicado: (2023)
Can LLMs Augment Low-Resource Reading Comprehension Datasets? Opportunities and Challenges
por: Samuel, Vinay, et al.
Publicado: (2023)
por: Samuel, Vinay, et al.
Publicado: (2023)
The potential -- and the pitfalls -- of using pre-trained language models as cognitive science theories
por: Shah, Raj Sanjay, et al.
Publicado: (2025)
por: Shah, Raj Sanjay, et al.
Publicado: (2025)
The World According to LLMs: How Geographic Origin Influences LLMs' Entity Deduction Capabilities
por: Lalai, Harsh Nishant, et al.
Publicado: (2025)
por: Lalai, Harsh Nishant, et al.
Publicado: (2025)
A Systematic Survey of Automatic Prompt Optimization Techniques
por: Ramnath, Kiran, et al.
Publicado: (2025)
por: Ramnath, Kiran, et al.
Publicado: (2025)
A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications
por: Sahoo, Pranab, et al.
Publicado: (2024)
por: Sahoo, Pranab, et al.
Publicado: (2024)
How Well Do Deep Learning Models Capture Human Concepts? The Case of the Typicality Effect
por: Vemuri, Siddhartha K., et al.
Publicado: (2024)
por: Vemuri, Siddhartha K., et al.
Publicado: (2024)
The Chameleon Nature of LLMs: Quantifying Multi-Turn Stance Instability in Search-Enabled Language Models
por: Ratnakar, Shivam, et al.
Publicado: (2025)
por: Ratnakar, Shivam, et al.
Publicado: (2025)
Steering LLMs for Formal Theorem Proving
por: Kirtania, Shashank, et al.
Publicado: (2025)
por: Kirtania, Shashank, et al.
Publicado: (2025)
A Survey on Model Compression for Large Language Models
por: Zhu, Xunyu, et al.
Publicado: (2023)
por: Zhu, Xunyu, et al.
Publicado: (2023)
Can LLMs Detect Their Confabulations? Estimating Reliability in Uncertainty-Aware Language Models
por: Zhou, Tianyi, et al.
Publicado: (2025)
por: Zhou, Tianyi, et al.
Publicado: (2025)
Generative Data Augmentation using LLMs improves Distributional Robustness in Question Answering
por: Chowdhury, Arijit Ghosh, et al.
Publicado: (2023)
por: Chowdhury, Arijit Ghosh, et al.
Publicado: (2023)
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
por: Yadav, Ankit, et al.
Publicado: (2024)
por: Yadav, Ankit, et al.
Publicado: (2024)
Constraining Sequential Model Editing with Editing Anchor Compression
por: Xu, Hao-Xiang, et al.
Publicado: (2025)
por: Xu, Hao-Xiang, et al.
Publicado: (2025)
Evaluation of Language Models in the Medical Context Under Resource-Constrained Settings
por: Posada, Andrea, et al.
Publicado: (2024)
por: Posada, Andrea, et al.
Publicado: (2024)
Optimizing Length Compression in Large Reasoning Models
por: Cheng, Zhengxiang, et al.
Publicado: (2025)
por: Cheng, Zhengxiang, et al.
Publicado: (2025)
The Point of No Return: Counterfactual Localization of Deceptive Commitment in Language-Model Reasoning
por: Merrill, Scott, et al.
Publicado: (2026)
por: Merrill, Scott, et al.
Publicado: (2026)
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
por: Hong, Junyuan, et al.
Publicado: (2024)
por: Hong, Junyuan, et al.
Publicado: (2024)
A Comprehensive Survey of Machine Unlearning Techniques for Large Language Models
por: Geng, Jiahui, et al.
Publicado: (2025)
por: Geng, Jiahui, et al.
Publicado: (2025)
Evaluating the Effectiveness of Data Augmentation for Emotion Classification in Low-Resource Settings
por: Arora, Aashish, et al.
Publicado: (2024)
por: Arora, Aashish, et al.
Publicado: (2024)
YesBut: A High-Quality Annotated Multimodal Dataset for evaluating Satire Comprehension capability of Vision-Language Models
por: Nandy, Abhilash, et al.
Publicado: (2024)
por: Nandy, Abhilash, et al.
Publicado: (2024)
How Culturally Aware are Vision-Language Models?
por: Burda-Lassen, Olena, et al.
Publicado: (2024)
por: Burda-Lassen, Olena, et al.
Publicado: (2024)
TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning Tasks
por: Kapoor, Vansh, et al.
Publicado: (2026)
por: Kapoor, Vansh, et al.
Publicado: (2026)
A Survey of Generative Categories and Techniques in Multimodal Generative Models
por: Han, Longzhen, et al.
Publicado: (2025)
por: Han, Longzhen, et al.
Publicado: (2025)
LLMs for Translation: Historical, Low-Resourced Languages and Contemporary AI Models
por: Tekgurler, Merve
Publicado: (2025)
por: Tekgurler, Merve
Publicado: (2025)
FedPT: Federated Proxy-Tuning of Large Language Models on Resource-Constrained Edge Devices
por: Gao, Zhidong, et al.
Publicado: (2024)
por: Gao, Zhidong, et al.
Publicado: (2024)
Difficulty Estimation and Simplification of French Text Using LLMs
por: Jamet, Henri, et al.
Publicado: (2024)
por: Jamet, Henri, et al.
Publicado: (2024)
CycleDistill: Bootstrapping Machine Translation using LLMs with Cyclical Distillation
por: Halder, Deepon, et al.
Publicado: (2025)
por: Halder, Deepon, et al.
Publicado: (2025)
Mamba in Vision: A Comprehensive Survey of Techniques and Applications
por: Rahman, Md Maklachur, et al.
Publicado: (2024)
por: Rahman, Md Maklachur, et al.
Publicado: (2024)
Cross-Domain Content Generation with Domain-Specific Small Language Models
por: Maloo, Ankit, et al.
Publicado: (2024)
por: Maloo, Ankit, et al.
Publicado: (2024)
Benchmarking LLMs for Pairwise Causal Discovery in Biomedical and Multi-Domain Contexts
por: Anuyah, Sydney, et al.
Publicado: (2026)
por: Anuyah, Sydney, et al.
Publicado: (2026)
TokenSkip: Controllable Chain-of-Thought Compression in LLMs
por: Xia, Heming, et al.
Publicado: (2025)
por: Xia, Heming, et al.
Publicado: (2025)
Ejemplares similares
-
Explainable Artificial Intelligence Techniques for Software Development Lifecycle: A Phase-specific Survey
por: Arora, Lakshit, et al.
Publicado: (2025) -
AI and Generative AI Transforming Disaster Management: A Survey of Damage Assessment and Response Techniques
por: Raj, Aman, et al.
Publicado: (2025) -
Adversarial Attacks in Multimodal Systems: A Practitioner's Survey
por: Kapoor, Shashank, et al.
Publicado: (2025) -
Opportunities and Applications of GenAI in Smart Cities: A User-Centric Survey
por: Shetgaonkar, Ankit, et al.
Publicado: (2025) -
Mitigating Clinician Information Overload: Generative AI for Integrated EHR and RPM Data Analysis
por: Shetgaonkar, Ankit, et al.
Publicado: (2025)