Efficient Strategy for Improving Large Language Model (LLM) Capabilities
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Gutiérrez, Julián Camilo Velandia |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Computational Economics in Large Language Models: Exploring Model Behavior and Incentive Design under Resource Constraints
par: Reddy, Sandeep, et autres
Publié: (2025)
par: Reddy, Sandeep, et autres
Publié: (2025)
Cost-Aware Model Selection for Text Classification: Multi-Objective Trade-offs Between Fine-Tuned Encoders and LLM Prompting in Production
par: Gonzalez, Alberto Andres Valdes
Publié: (2026)
par: Gonzalez, Alberto Andres Valdes
Publié: (2026)
Adaptive Activation Cancellation for Hallucination Mitigation in Large Language Models
par: Yocam, Eric, et autres
Publié: (2026)
par: Yocam, Eric, et autres
Publié: (2026)
Listwise Direct Preference Optimization with Multi-Dimensional Preference Mixing
par: Sun, Yuhui, et autres
Publié: (2025)
par: Sun, Yuhui, et autres
Publié: (2025)
Less is More: Learning Graph Tasks with Just LLMs
par: Shirai, Sola, et autres
Publié: (2025)
par: Shirai, Sola, et autres
Publié: (2025)
Beyond Subtokens: A Rich Character Embedding for Low-resource and Morphologically Complex Languages
par: Schneider, Felix, et autres
Publié: (2026)
par: Schneider, Felix, et autres
Publié: (2026)
Exploring Model Invariance with Discrete Search for Ultra-Low-Bit Quantization
par: Wen, Yuqiao, et autres
Publié: (2025)
par: Wen, Yuqiao, et autres
Publié: (2025)
When is dataset cartography ineffective? Using training dynamics does not improve robustness against Adversarial SQuAD
par: Mandal, Paul K.
Publié: (2025)
par: Mandal, Paul K.
Publié: (2025)
ADALog: Adaptive Unsupervised Anomaly detection in Logs with Self-attention Masked Language Model
par: Pospieszny, Przemek, et autres
Publié: (2025)
par: Pospieszny, Przemek, et autres
Publié: (2025)
EBBS: An Ensemble with Bi-Level Beam Search for Zero-Shot Machine Translation
par: Wen, Yuqiao, et autres
Publié: (2024)
par: Wen, Yuqiao, et autres
Publié: (2024)
Adversarially Probing Cross-Family Sound Symbolism in 27 Languages
par: Sharma, Anika, et autres
Publié: (2025)
par: Sharma, Anika, et autres
Publié: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
par: Fadli, Samih
Publié: (2025)
par: Fadli, Samih
Publié: (2025)
Neural Attention: A Novel Mechanism for Enhanced Expressive Power in Transformer Models
par: DiGiugno, Andrew, et autres
Publié: (2025)
par: DiGiugno, Andrew, et autres
Publié: (2025)
Large Language Models Are Not Strong Abstract Reasoners
par: Gendron, Gaël, et autres
Publié: (2023)
par: Gendron, Gaël, et autres
Publié: (2023)
Improving Consistency in Large Language Models through Chain of Guidance
par: Raj, Harsh, et autres
Publié: (2025)
par: Raj, Harsh, et autres
Publié: (2025)
A Comparative Study of Feature Selection in Tsetlin Machines
par: Halenka, Vojtech, et autres
Publié: (2025)
par: Halenka, Vojtech, et autres
Publié: (2025)
LLM Performance Predictors: Learning When to Escalate in Hybrid Human-AI Moderation Systems
par: Bachar, Or, et autres
Publié: (2026)
par: Bachar, Or, et autres
Publié: (2026)
A Multi-Encoder Frozen-Decoder Approach for Fine-Tuning Large Language Models
par: Dhole, Kaustubh D.
Publié: (2025)
par: Dhole, Kaustubh D.
Publié: (2025)
Measuring Intent Comprehension in LLMs
par: Kunievsky, Nadav, et autres
Publié: (2025)
par: Kunievsky, Nadav, et autres
Publié: (2025)
Memory Bank Compression for Continual Adaptation of Large Language Models
par: Katraouras, Thomas, et autres
Publié: (2026)
par: Katraouras, Thomas, et autres
Publié: (2026)
Let Your Graph Do the Talking: Encoding Structured Data for LLMs
par: Perozzi, Bryan, et autres
Publié: (2024)
par: Perozzi, Bryan, et autres
Publié: (2024)
Cytoarchitecture in Words: Weakly Supervised Vision-Language Modeling for Human Brain Microscopy
par: Sutton, Matthew, et autres
Publié: (2026)
par: Sutton, Matthew, et autres
Publié: (2026)
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
par: Shravan, Rohan
Publié: (2026)
par: Shravan, Rohan
Publié: (2026)
Applied Explainability for Large Language Models: A Comparative Study
par: Kancharla, Venkata Abhinandan
Publié: (2026)
par: Kancharla, Venkata Abhinandan
Publié: (2026)
Language Models Are Implicitly Continuous
par: Marro, Samuele, et autres
Publié: (2025)
par: Marro, Samuele, et autres
Publié: (2025)
Combining Language and Topic Models for Hierarchical Text Classification
par: Toit, Jaco du, et autres
Publié: (2025)
par: Toit, Jaco du, et autres
Publié: (2025)
Sample-Efficient Language Model for Hinglish Conversational AI
par: Singh, Sakshi, et autres
Publié: (2025)
par: Singh, Sakshi, et autres
Publié: (2025)
Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey
par: Huang, Yunpeng, et autres
Publié: (2023)
par: Huang, Yunpeng, et autres
Publié: (2023)
LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
par: Hashemi, Helia, et autres
Publié: (2024)
par: Hashemi, Helia, et autres
Publié: (2024)
Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding
par: Zhu, Jiajun, et autres
Publié: (2025)
par: Zhu, Jiajun, et autres
Publié: (2025)
propella-1: Multi-Property Document Annotation for LLM Data Curation at Scale
par: Idahl, Maximilian, et autres
Publié: (2026)
par: Idahl, Maximilian, et autres
Publié: (2026)
PersonalLLM: Tailoring LLMs to Individual Preferences
par: Zollo, Thomas P., et autres
Publié: (2024)
par: Zollo, Thomas P., et autres
Publié: (2024)
LLM Vocabulary Compression for Low-Compute Environments
par: Vennam, Sreeram, et autres
Publié: (2024)
par: Vennam, Sreeram, et autres
Publié: (2024)
How Language Models Process Out-of-Distribution Inputs: A Two-Pathway Framework
par: Saghir, Hamidreza
Publié: (2026)
par: Saghir, Hamidreza
Publié: (2026)
Quantization-Robust LLM Unlearning via Low-Rank Adaptation
par: Abitante, João Vitor Boer, et autres
Publié: (2026)
par: Abitante, João Vitor Boer, et autres
Publié: (2026)
Neural Activation Patterns Across Language Model Architectures: A Comprehensive Analysis of Cognitive Task Performance
par: Naser-Moghadasi, Mahdi, et autres
Publié: (2026)
par: Naser-Moghadasi, Mahdi, et autres
Publié: (2026)
HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools
par: Garg, Aashna, et autres
Publié: (2026)
par: Garg, Aashna, et autres
Publié: (2026)
Future Token Prediction -- Causal Language Modelling with Per-Token Semantic State Vector for Multi-Token Prediction
par: Walker, Nicholas
Publié: (2024)
par: Walker, Nicholas
Publié: (2024)
Your Pretrained Model Tells the Difficulty Itself: A Self-Adaptive Curriculum Learning Paradigm for Natural Language Understanding
par: Feng, Qi, et autres
Publié: (2025)
par: Feng, Qi, et autres
Publié: (2025)
Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders
par: Patel, Het, et autres
Publié: (2026)
par: Patel, Het, et autres
Publié: (2026)
Documents similaires
-
Computational Economics in Large Language Models: Exploring Model Behavior and Incentive Design under Resource Constraints
par: Reddy, Sandeep, et autres
Publié: (2025) -
Cost-Aware Model Selection for Text Classification: Multi-Objective Trade-offs Between Fine-Tuned Encoders and LLM Prompting in Production
par: Gonzalez, Alberto Andres Valdes
Publié: (2026) -
Adaptive Activation Cancellation for Hallucination Mitigation in Large Language Models
par: Yocam, Eric, et autres
Publié: (2026) -
Listwise Direct Preference Optimization with Multi-Dimensional Preference Mixing
par: Sun, Yuhui, et autres
Publié: (2025) -
Less is More: Learning Graph Tasks with Just LLMs
par: Shirai, Sola, et autres
Publié: (2025)