Faster and Lighter LLMs: A Survey on Current Challenges and Way Forward
Fuente:
arXiv
Salvato in:
| Autori principali: | Chavan, Arnav, Magazine, Raghav, Kushwaha, Shubham, Debbah, Mérouane, Gupta, Deepak |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Surgical Feature-Space Decomposition of LLMs: Why, When and How?
di: Chavan, Arnav, et al.
Pubblicazione: (2024)
di: Chavan, Arnav, et al.
Pubblicazione: (2024)
Beyond Uniform Scaling: Exploring Depth Heterogeneity in Neural Architectures
di: T, Akash Guna R., et al.
Pubblicazione: (2024)
di: T, Akash Guna R., et al.
Pubblicazione: (2024)
PromptWizard: Task-Aware Prompt Optimization Framework
di: Agarwal, Eshaan, et al.
Pubblicazione: (2024)
di: Agarwal, Eshaan, et al.
Pubblicazione: (2024)
Sociodemographic Bias in Language Models: A Survey and Forward Path
di: Gupta, Vipul, et al.
Pubblicazione: (2023)
di: Gupta, Vipul, et al.
Pubblicazione: (2023)
How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
di: Seddik, Mohamed El Amine, et al.
Pubblicazione: (2024)
di: Seddik, Mohamed El Amine, et al.
Pubblicazione: (2024)
CORAL: Learning Consistent Representations across Multi-step Training with Lighter Speculative Drafter
di: Weng, Yepeng, et al.
Pubblicazione: (2025)
di: Weng, Yepeng, et al.
Pubblicazione: (2025)
RLHF: A comprehensive Survey for Cultural, Multimodal and Low Latency Alignment Methods
di: Sharma, Raghav, et al.
Pubblicazione: (2025)
di: Sharma, Raghav, et al.
Pubblicazione: (2025)
DOT-MoE: Differentiable Optimal Transport for MoEfication
di: Bamba, Udbhav, et al.
Pubblicazione: (2026)
di: Bamba, Udbhav, et al.
Pubblicazione: (2026)
Sparser, Faster, Lighter Transformer Language Models
di: Cetin, Edoardo, et al.
Pubblicazione: (2026)
di: Cetin, Edoardo, et al.
Pubblicazione: (2026)
MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological Generation
di: Agarwal, Mehul, et al.
Pubblicazione: (2026)
di: Agarwal, Mehul, et al.
Pubblicazione: (2026)
Faster Cascades via Speculative Decoding
di: Narasimhan, Harikrishna, et al.
Pubblicazione: (2024)
di: Narasimhan, Harikrishna, et al.
Pubblicazione: (2024)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
di: Shen, Xuan, et al.
Pubblicazione: (2023)
di: Shen, Xuan, et al.
Pubblicazione: (2023)
Multilingual Prompt Engineering in Large Language Models: A Survey Across NLP Tasks
di: Vatsal, Shubham, et al.
Pubblicazione: (2025)
di: Vatsal, Shubham, et al.
Pubblicazione: (2025)
LoRMA: Low-Rank Multiplicative Adaptation for LLMs
di: Bihany, Harsh, et al.
Pubblicazione: (2025)
di: Bihany, Harsh, et al.
Pubblicazione: (2025)
CXMArena: Unified Dataset to benchmark performance in realistic CXM Scenarios
di: Garg, Raghav, et al.
Pubblicazione: (2025)
di: Garg, Raghav, et al.
Pubblicazione: (2025)
S2D: Selective Spectral Decay for Quantization-Friendly Conditioning of Neural Activations
di: Chavan, Arnav, et al.
Pubblicazione: (2026)
di: Chavan, Arnav, et al.
Pubblicazione: (2026)
IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs
di: Manikantan, Kawshik, et al.
Pubblicazione: (2024)
di: Manikantan, Kawshik, et al.
Pubblicazione: (2024)
3D-Properties: Identifying Challenges in DPO and Charting a Path Forward
di: Yan, Yuzi, et al.
Pubblicazione: (2024)
di: Yan, Yuzi, et al.
Pubblicazione: (2024)
Reasoning Beyond Limits: Advances and Open Problems for LLMs
di: Ferrag, Mohamed Amine, et al.
Pubblicazione: (2025)
di: Ferrag, Mohamed Amine, et al.
Pubblicazione: (2025)
CrisisKAN: Knowledge-infused and Explainable Multimodal Attention Network for Crisis Event Classification
di: Gupta, Shubham, et al.
Pubblicazione: (2024)
di: Gupta, Shubham, et al.
Pubblicazione: (2024)
Reinforcement Learning Enhanced LLMs: A Survey
di: Wang, Shuhe, et al.
Pubblicazione: (2024)
di: Wang, Shuhe, et al.
Pubblicazione: (2024)
RLSF: Fine-tuning LLMs via Symbolic Feedback
di: Jha, Piyush, et al.
Pubblicazione: (2024)
di: Jha, Piyush, et al.
Pubblicazione: (2024)
Towards Scalable Automated Alignment of LLMs: A Survey
di: Cao, Boxi, et al.
Pubblicazione: (2024)
di: Cao, Boxi, et al.
Pubblicazione: (2024)
A Surprising Failure? Multimodal LLMs and the NLVR Challenge
di: Wu, Anne, et al.
Pubblicazione: (2024)
di: Wu, Anne, et al.
Pubblicazione: (2024)
LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report
di: Zhao, Justin, et al.
Pubblicazione: (2024)
di: Zhao, Justin, et al.
Pubblicazione: (2024)
Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case Study
di: Ponkshe, Kaustubh, et al.
Pubblicazione: (2025)
di: Ponkshe, Kaustubh, et al.
Pubblicazione: (2025)
Exploring Design Choices for Building Language-Specific LLMs
di: Tejaswi, Atula, et al.
Pubblicazione: (2024)
di: Tejaswi, Atula, et al.
Pubblicazione: (2024)
Large Language Models for Time Series: A Survey
di: Zhang, Xiyuan, et al.
Pubblicazione: (2024)
di: Zhang, Xiyuan, et al.
Pubblicazione: (2024)
Generation and De-Identification of Indian Clinical Discharge Summaries using LLMs
di: Singh, Sanjeet, et al.
Pubblicazione: (2024)
di: Singh, Sanjeet, et al.
Pubblicazione: (2024)
Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques
di: Girija, Sanjay Surendranath, et al.
Pubblicazione: (2025)
di: Girija, Sanjay Surendranath, et al.
Pubblicazione: (2025)
Self-Speculative Biased Decoding for Faster Re-Translation
di: Zeng, Linxiao, et al.
Pubblicazione: (2025)
di: Zeng, Linxiao, et al.
Pubblicazione: (2025)
Scaling Agentic Capabilities, Not Context: Efficient Reinforcement Finetuning for Large Toolspaces
di: Gupta, Karan, et al.
Pubblicazione: (2026)
di: Gupta, Karan, et al.
Pubblicazione: (2026)
ABBA-Adapters: Efficient and Expressive Fine-Tuning of Foundation Models
di: Singhal, Raghav, et al.
Pubblicazione: (2025)
di: Singhal, Raghav, et al.
Pubblicazione: (2025)
When Every Token Counts: Optimal Segmentation for Low-Resource Language Models
di: Raj, Bharath, et al.
Pubblicazione: (2024)
di: Raj, Bharath, et al.
Pubblicazione: (2024)
DENIAHL: In-Context Features Influence LLM Needle-In-A-Haystack Abilities
di: Dai, Hui, et al.
Pubblicazione: (2024)
di: Dai, Hui, et al.
Pubblicazione: (2024)
NyayaMind- A Framework for Transparent Legal Reasoning and Judgment Prediction in the Indian Legal System
di: Shukla, Parjanya Aditya, et al.
Pubblicazione: (2026)
di: Shukla, Parjanya Aditya, et al.
Pubblicazione: (2026)
Structured Legal Document Generation in India: A Model-Agnostic Wrapper Approach with VidhikDastaavej
di: Nigam, Shubham Kumar, et al.
Pubblicazione: (2025)
di: Nigam, Shubham Kumar, et al.
Pubblicazione: (2025)
Can GPT Redefine Medical Understanding? Evaluating GPT on Biomedical Machine Reading Comprehension
di: Vatsal, Shubham, et al.
Pubblicazione: (2024)
di: Vatsal, Shubham, et al.
Pubblicazione: (2024)
ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challenges
di: Qian, Cheng, et al.
Pubblicazione: (2025)
di: Qian, Cheng, et al.
Pubblicazione: (2025)
A Case Study Exploring the Current Landscape of Synthetic Medical Record Generation with Commercial LLMs
di: Lin, Yihan, et al.
Pubblicazione: (2025)
di: Lin, Yihan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Surgical Feature-Space Decomposition of LLMs: Why, When and How?
di: Chavan, Arnav, et al.
Pubblicazione: (2024) -
Beyond Uniform Scaling: Exploring Depth Heterogeneity in Neural Architectures
di: T, Akash Guna R., et al.
Pubblicazione: (2024) -
PromptWizard: Task-Aware Prompt Optimization Framework
di: Agarwal, Eshaan, et al.
Pubblicazione: (2024) -
Sociodemographic Bias in Language Models: A Survey and Forward Path
di: Gupta, Vipul, et al.
Pubblicazione: (2023) -
How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
di: Seddik, Mohamed El Amine, et al.
Pubblicazione: (2024)