Merging Continual Pretraining Models for Domain-Specialized LLMs: A Case Study in Finance
Fuente:
arXiv
Salvato in:
| Autori principali: | Ueda, Kentaro, Portet, François, Suwa, Hirohiko, Yasumoto, Keiichi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Construction of Instruction-tuned LLMs for Finance without Instruction Data Using Continual Pretraining and Model Merging
di: Hirano, Masanori, et al.
Pubblicazione: (2024)
di: Hirano, Masanori, et al.
Pubblicazione: (2024)
Construction of Domain-specified Japanese Large Language Model for Finance through Continual Pre-training
di: Hirano, Masanori, et al.
Pubblicazione: (2024)
di: Hirano, Masanori, et al.
Pubblicazione: (2024)
MedMeta: A Benchmark for LLMs in Synthesizing Meta-Analysis Conclusion from Medical Studies
di: Ha, Huy Hoang, et al.
Pubblicazione: (2026)
di: Ha, Huy Hoang, et al.
Pubblicazione: (2026)
Automated Clinical Report Generation for Remote Cognitive Remediation: Comparing Knowledge-Engineered Templates and LLMs in Low-Resource Settings
di: Zhou, Yongxin, et al.
Pubblicazione: (2026)
di: Zhou, Yongxin, et al.
Pubblicazione: (2026)
Is Biomedical Specialization Still Worth It? Insights from Domain-Adaptive Language Modelling with a New French Health Corpus
di: Mannion, Aidan, et al.
Pubblicazione: (2026)
di: Mannion, Aidan, et al.
Pubblicazione: (2026)
Evaluation of Deontic Conditional Reasoning in Large Language Models: The Case of Wason's Selection Task
di: Abe, Hirohiko, et al.
Pubblicazione: (2026)
di: Abe, Hirohiko, et al.
Pubblicazione: (2026)
Can GPT models Follow Human Summarization Guidelines? A Study for Targeted Communication Goals
di: Zhou, Yongxin, et al.
Pubblicazione: (2023)
di: Zhou, Yongxin, et al.
Pubblicazione: (2023)
THaLLE-ThaiLLM: Domain-Specialized Small LLMs for Finance and Thai -- Technical Report
di: Labs, KBTG, et al.
Pubblicazione: (2026)
di: Labs, KBTG, et al.
Pubblicazione: (2026)
PSentScore: Evaluating Sentiment Polarity in Dialogue Summarization
di: Zhou, Yongxin, et al.
Pubblicazione: (2023)
di: Zhou, Yongxin, et al.
Pubblicazione: (2023)
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
di: Méloux, Maxime, et al.
Pubblicazione: (2025)
di: Méloux, Maxime, et al.
Pubblicazione: (2025)
Enabling Vibration-Based Gesture Recognition on Everyday Furniture via Energy-Efficient FPGA Implementation of 1D Convolutional Networks
di: Shibata, Koki, et al.
Pubblicazione: (2025)
di: Shibata, Koki, et al.
Pubblicazione: (2025)
Where Do Self-Supervised Speech Models Become Unfair?
di: Herron, Felix, et al.
Pubblicazione: (2026)
di: Herron, Felix, et al.
Pubblicazione: (2026)
What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search
di: Zhang, Xinhao, et al.
Pubblicazione: (2026)
di: Zhang, Xinhao, et al.
Pubblicazione: (2026)
Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models
di: Herron, Felix, et al.
Pubblicazione: (2026)
di: Herron, Felix, et al.
Pubblicazione: (2026)
Normative Reasoning in Large Language Models: A Comparative Benchmark from Logical and Modal Perspectives
di: Ozeki, Kentaro, et al.
Pubblicazione: (2025)
di: Ozeki, Kentaro, et al.
Pubblicazione: (2025)
ixi-GEN: Efficient Industrial sLLMs through Domain Adaptive Continual Pretraining
di: Kim, Seonwu, et al.
Pubblicazione: (2025)
di: Kim, Seonwu, et al.
Pubblicazione: (2025)
Checkpoint Merging via Bayesian Optimization in LLM Pretraining
di: Liu, Deyuan, et al.
Pubblicazione: (2024)
di: Liu, Deyuan, et al.
Pubblicazione: (2024)
Abductive Reasoning with Syllogistic Forms in Large Language Models
di: Abe, Hirohiko, et al.
Pubblicazione: (2026)
di: Abe, Hirohiko, et al.
Pubblicazione: (2026)
Syntactic Learnability of Echo State Neural Language Models at Scale
di: Ueda, Ryo, et al.
Pubblicazione: (2025)
di: Ueda, Ryo, et al.
Pubblicazione: (2025)
Efficient Domain-adaptive Continual Pretraining for the Process Industry in the German Language
di: Zhukova, Anastasia, et al.
Pubblicazione: (2025)
di: Zhukova, Anastasia, et al.
Pubblicazione: (2025)
Channel Merging: Preserving Specialization for Merged Experts
di: Zhang, Mingyang, et al.
Pubblicazione: (2024)
di: Zhang, Mingyang, et al.
Pubblicazione: (2024)
Pretraining and Updates of Domain-Specific LLM: A Case Study in the Japanese Business Domain
di: Takahashi, Kosuke, et al.
Pubblicazione: (2024)
di: Takahashi, Kosuke, et al.
Pubblicazione: (2024)
Exploring Reasoning Biases in Large Language Models Through Syllogism: Insights from the NeuBAROCO Dataset
di: Ozeki, Kentaro, et al.
Pubblicazione: (2024)
di: Ozeki, Kentaro, et al.
Pubblicazione: (2024)
Beyond Fine-tuning: Unleashing the Potential of Continuous Pretraining for Clinical LLMs
di: Christophe, Clément, et al.
Pubblicazione: (2024)
di: Christophe, Clément, et al.
Pubblicazione: (2024)
ELO: Efficient Layer-Specific Optimization for Continual Pretraining of Multilingual LLMs
di: Yoo, HanGyeol, et al.
Pubblicazione: (2026)
di: Yoo, HanGyeol, et al.
Pubblicazione: (2026)
FinanceMath: Knowledge-Intensive Math Reasoning in Finance Domains
di: Zhao, Yilun, et al.
Pubblicazione: (2023)
di: Zhao, Yilun, et al.
Pubblicazione: (2023)
Continual Pretraining on Encrypted Synthetic Data for Privacy-Preserving LLMs
di: Liu, Honghao, et al.
Pubblicazione: (2026)
di: Liu, Honghao, et al.
Pubblicazione: (2026)
Responsible Benchmarking of Fairness for Automatic Speech Recognition
di: Herron, Felix, et al.
Pubblicazione: (2026)
di: Herron, Felix, et al.
Pubblicazione: (2026)
Compressing Language Models for Specialized Domains
di: Williams, Miles, et al.
Pubblicazione: (2025)
di: Williams, Miles, et al.
Pubblicazione: (2025)
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
di: Itzhak, Itay, et al.
Pubblicazione: (2025)
di: Itzhak, Itay, et al.
Pubblicazione: (2025)
Assessing the Political Fairness of Multilingual LLMs: A Case Study based on a 21-way Multiparallel EuroParl Dataset
di: Lerner, Paul, et al.
Pubblicazione: (2025)
di: Lerner, Paul, et al.
Pubblicazione: (2025)
Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks
di: Wang, Zheng, et al.
Pubblicazione: (2024)
di: Wang, Zheng, et al.
Pubblicazione: (2024)
IKnow: Instruction-Knowledge-Aware Continual Pretraining for Effective Domain Adaptation
di: Zhang, Tianyi, et al.
Pubblicazione: (2025)
di: Zhang, Tianyi, et al.
Pubblicazione: (2025)
The Thinking Spectrum: An Empirical Study of Tunable Reasoning in LLMs through Model Merging
di: Lan, Xiaochong, et al.
Pubblicazione: (2025)
di: Lan, Xiaochong, et al.
Pubblicazione: (2025)
Rethinking Multilingual Continual Pretraining: Data Mixing for Adapting LLMs Across Languages and Resources
di: Li, Zihao, et al.
Pubblicazione: (2025)
di: Li, Zihao, et al.
Pubblicazione: (2025)
Evaluating the Effectiveness of Linguistic Knowledge in Pretrained Language Models: A Case Study of Universal Dependencies
di: Li, Wenxi
Pubblicazione: (2025)
di: Li, Wenxi
Pubblicazione: (2025)
Ingest-And-Ground: Dispelling Hallucinations from Continually-Pretrained LLMs with RAG
di: Fang, Chenhao, et al.
Pubblicazione: (2024)
di: Fang, Chenhao, et al.
Pubblicazione: (2024)
DogeRM: Equipping Reward Models with Domain Knowledge through Model Merging
di: Lin, Tzu-Han, et al.
Pubblicazione: (2024)
di: Lin, Tzu-Han, et al.
Pubblicazione: (2024)
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
di: Méloux, Maxime, et al.
Pubblicazione: (2025)
di: Méloux, Maxime, et al.
Pubblicazione: (2025)
How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion
di: Seth, Agrima, et al.
Pubblicazione: (2025)
di: Seth, Agrima, et al.
Pubblicazione: (2025)
Documenti analoghi
-
The Construction of Instruction-tuned LLMs for Finance without Instruction Data Using Continual Pretraining and Model Merging
di: Hirano, Masanori, et al.
Pubblicazione: (2024) -
Construction of Domain-specified Japanese Large Language Model for Finance through Continual Pre-training
di: Hirano, Masanori, et al.
Pubblicazione: (2024) -
MedMeta: A Benchmark for LLMs in Synthesizing Meta-Analysis Conclusion from Medical Studies
di: Ha, Huy Hoang, et al.
Pubblicazione: (2026) -
Automated Clinical Report Generation for Remote Cognitive Remediation: Comparing Knowledge-Engineered Templates and LLMs in Low-Resource Settings
di: Zhou, Yongxin, et al.
Pubblicazione: (2026) -
Is Biomedical Specialization Still Worth It? Insights from Domain-Adaptive Language Modelling with a New French Health Corpus
di: Mannion, Aidan, et al.
Pubblicazione: (2026)