Fine-tuning large language models for domain adaptation: Exploration of training strategies, scaling, model merging and synergistic capabilities
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Wei, Luu, Rachel K., Buehler, Markus J. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Agentic Deep Graph Reasoning Yields Self-Organizing Knowledge Networks
von: Buehler, Markus J.
Veröffentlicht: (2025)
von: Buehler, Markus J.
Veröffentlicht: (2025)
In-situ graph reasoning and knowledge expansion using Graph-PReFLexOR
von: Buehler, Markus J.
Veröffentlicht: (2025)
von: Buehler, Markus J.
Veröffentlicht: (2025)
Fine-tuning of lightweight large language models for sentiment classification on heterogeneous financial textual data
von: Amorin, Alvaro Paredes, et al.
Veröffentlicht: (2025)
von: Amorin, Alvaro Paredes, et al.
Veröffentlicht: (2025)
ZNO-Eval: Benchmarking reasoning capabilities of large language models in Ukrainian
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
Graph-Aware Isomorphic Attention for Adaptive Dynamics in Transformers
von: Buehler, Markus J.
Veröffentlicht: (2025)
von: Buehler, Markus J.
Veröffentlicht: (2025)
StatLLaMA: Multi-Stage training for domain-optimized statistical large language models
von: Zeng, Jing-Yi, et al.
Veröffentlicht: (2025)
von: Zeng, Jing-Yi, et al.
Veröffentlicht: (2025)
BeamPERL: Parameter-Efficient RL with Verifiable Rewards Specializes Compact LLMs for Structured Beam Mechanics Reasoning
von: Hage, Tarjei Paule, et al.
Veröffentlicht: (2026)
von: Hage, Tarjei Paule, et al.
Veröffentlicht: (2026)
Higher-Order Knowledge Representations for Agentic Scientific Reasoning
von: Stewart, Isabella A., et al.
Veröffentlicht: (2026)
von: Stewart, Isabella A., et al.
Veröffentlicht: (2026)
ProtAgents: Protein discovery via large language model multi-agent collaborations combining physics and machine learning
von: Ghafarollahi, A., et al.
Veröffentlicht: (2024)
von: Ghafarollahi, A., et al.
Veröffentlicht: (2024)
Reshaping MOFs text mining with a dynamic multi-agents framework of large language model
von: Lin, Zuhong, et al.
Veröffentlicht: (2025)
von: Lin, Zuhong, et al.
Veröffentlicht: (2025)
PRefLexOR: Preference-based Recursive Language Modeling for Exploratory Optimization of Reasoning and Agentic Thinking
von: Buehler, Markus J.
Veröffentlicht: (2024)
von: Buehler, Markus J.
Veröffentlicht: (2024)
How does fine-tuning improve sensorimotor representations in large language models?
von: Wu, Minghua, et al.
Veröffentlicht: (2026)
von: Wu, Minghua, et al.
Veröffentlicht: (2026)
Zero-shot cross-lingual transfer in instruction tuning of large language models
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2024)
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2024)
Learning the rules of peptide self-assembly through data mining with large language models
von: Yang, Zhenze, et al.
Veröffentlicht: (2024)
von: Yang, Zhenze, et al.
Veröffentlicht: (2024)
Accelerating Scientific Discovery with Generative Knowledge Extraction, Graph-Based Representation, and Multimodal Intelligent Graph Reasoning
von: Buehler, Markus J.
Veröffentlicht: (2024)
von: Buehler, Markus J.
Veröffentlicht: (2024)
Facilitating large language model Russian adaptation with Learned Embedding Propagation
von: Tikhomirov, Mikhail, et al.
Veröffentlicht: (2024)
von: Tikhomirov, Mikhail, et al.
Veröffentlicht: (2024)
Evolution of meta's llama models and parameter-efficient fine-tuning of large language models: a survey
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2025)
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2025)
Evidence of interrelated cognitive-like capabilities in large language models: Indications of artificial general intelligence or achievement?
von: Ilić, David, et al.
Veröffentlicht: (2023)
von: Ilić, David, et al.
Veröffentlicht: (2023)
ClinicalGPT-R1: Pushing reasoning capability of generalist disease diagnosis with large language model
von: Lan, Wuyang, et al.
Veröffentlicht: (2025)
von: Lan, Wuyang, et al.
Veröffentlicht: (2025)
TelcoLM: collecting data, adapting, and benchmarking language models for the telecommunication domain
von: Barboule, Camille, et al.
Veröffentlicht: (2024)
von: Barboule, Camille, et al.
Veröffentlicht: (2024)
Bringing legal knowledge to the public by constructing a legal question bank using large-scale pre-trained language model
von: Yuan, Mingruo, et al.
Veröffentlicht: (2025)
von: Yuan, Mingruo, et al.
Veröffentlicht: (2025)
Self-Revising Discovery Systems for Science: A Categorical Framework for Agentic Artificial Intelligence
von: Wang, Fiona Y., et al.
Veröffentlicht: (2026)
von: Wang, Fiona Y., et al.
Veröffentlicht: (2026)
SciAgents: Automating scientific discovery through multi-agent intelligent graph reasoning
von: Ghafarollahi, Alireza, et al.
Veröffentlicht: (2024)
von: Ghafarollahi, Alireza, et al.
Veröffentlicht: (2024)
Emergent effects of scaling on the functional hierarchies within large language models
von: Bogdan, Paul C.
Veröffentlicht: (2025)
von: Bogdan, Paul C.
Veröffentlicht: (2025)
Auxiliary task demands mask the capabilities of smaller language models
von: Hu, Jennifer, et al.
Veröffentlicht: (2024)
von: Hu, Jennifer, et al.
Veröffentlicht: (2024)
Dissociating language and thought in large language models
von: Mahowald, Kyle, et al.
Veröffentlicht: (2023)
von: Mahowald, Kyle, et al.
Veröffentlicht: (2023)
Large language models in healthcare and medical domain: A review
von: Nazi, Zabir Al, et al.
Veröffentlicht: (2023)
von: Nazi, Zabir Al, et al.
Veröffentlicht: (2023)
Fine-tuning multilingual language models in Twitter/X sentiment analysis: a study on Eastern-European V4 languages
von: Filip, Tomáš, et al.
Veröffentlicht: (2024)
von: Filip, Tomáš, et al.
Veröffentlicht: (2024)
WizardLM: Empowering large pre-trained language models to follow complex instructions
von: Xu, Can, et al.
Veröffentlicht: (2023)
von: Xu, Can, et al.
Veröffentlicht: (2023)
Metaphor identification using large language models: A comparison of RAG, prompt engineering, and fine-tuning
von: Fuoli, Matteo, et al.
Veröffentlicht: (2025)
von: Fuoli, Matteo, et al.
Veröffentlicht: (2025)
Post-training makes large language models less human-like
von: Binz, Marcel, et al.
Veröffentlicht: (2026)
von: Binz, Marcel, et al.
Veröffentlicht: (2026)
Generative Artificial Intelligence Extracts Structure-Function Relationships from Plants for New Materials
von: Luu, Rachel K., et al.
Veröffentlicht: (2025)
von: Luu, Rachel K., et al.
Veröffentlicht: (2025)
Assessment and manipulation of latent constructs in pre-trained language models using psychometric scales
von: Reuben, Maor, et al.
Veröffentlicht: (2024)
von: Reuben, Maor, et al.
Veröffentlicht: (2024)
DepressLLM: Interpretable domain-adapted language model for depression detection from real-world narratives
von: Moon, Sehwan, et al.
Veröffentlicht: (2025)
von: Moon, Sehwan, et al.
Veröffentlicht: (2025)
On the attribution of confidence to large language models
von: Keeling, Geoff, et al.
Veröffentlicht: (2024)
von: Keeling, Geoff, et al.
Veröffentlicht: (2024)
A dataset and benchmark for hospital course summarization with adapted large language models
von: Aali, Asad, et al.
Veröffentlicht: (2024)
von: Aali, Asad, et al.
Veröffentlicht: (2024)
Can large language models assist choice modelling? Insights into prompting strategies and current models capabilities
von: Sfeir, Georges, et al.
Veröffentlicht: (2025)
von: Sfeir, Georges, et al.
Veröffentlicht: (2025)
Evaluating large language models in medical applications: a survey
von: Chen, Xiaolan, et al.
Veröffentlicht: (2024)
von: Chen, Xiaolan, et al.
Veröffentlicht: (2024)
Long-form factuality in large language models
von: Wei, Jerry, et al.
Veröffentlicht: (2024)
von: Wei, Jerry, et al.
Veröffentlicht: (2024)
Disentangling generalization and memorization in large language models using chess
von: Pleiss, Leonard S., et al.
Veröffentlicht: (2026)
von: Pleiss, Leonard S., et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Agentic Deep Graph Reasoning Yields Self-Organizing Knowledge Networks
von: Buehler, Markus J.
Veröffentlicht: (2025) -
In-situ graph reasoning and knowledge expansion using Graph-PReFLexOR
von: Buehler, Markus J.
Veröffentlicht: (2025) -
Fine-tuning of lightweight large language models for sentiment classification on heterogeneous financial textual data
von: Amorin, Alvaro Paredes, et al.
Veröffentlicht: (2025) -
ZNO-Eval: Benchmarking reasoning capabilities of large language models in Ukrainian
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025) -
Graph-Aware Isomorphic Attention for Adaptive Dynamics in Transformers
von: Buehler, Markus J.
Veröffentlicht: (2025)