Merging Continual Pretraining Models for Domain-Specialized LLMs: A Case Study in Finance
Fuente:
arXiv
Guardado en:
| Autores principales: | Ueda, Kentaro, Portet, François, Suwa, Hirohiko, Yasumoto, Keiichi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Construction of Instruction-tuned LLMs for Finance without Instruction Data Using Continual Pretraining and Model Merging
por: Hirano, Masanori, et al.
Publicado: (2024)
por: Hirano, Masanori, et al.
Publicado: (2024)
Construction of Domain-specified Japanese Large Language Model for Finance through Continual Pre-training
por: Hirano, Masanori, et al.
Publicado: (2024)
por: Hirano, Masanori, et al.
Publicado: (2024)
MedMeta: A Benchmark for LLMs in Synthesizing Meta-Analysis Conclusion from Medical Studies
por: Ha, Huy Hoang, et al.
Publicado: (2026)
por: Ha, Huy Hoang, et al.
Publicado: (2026)
Automated Clinical Report Generation for Remote Cognitive Remediation: Comparing Knowledge-Engineered Templates and LLMs in Low-Resource Settings
por: Zhou, Yongxin, et al.
Publicado: (2026)
por: Zhou, Yongxin, et al.
Publicado: (2026)
Is Biomedical Specialization Still Worth It? Insights from Domain-Adaptive Language Modelling with a New French Health Corpus
por: Mannion, Aidan, et al.
Publicado: (2026)
por: Mannion, Aidan, et al.
Publicado: (2026)
Evaluation of Deontic Conditional Reasoning in Large Language Models: The Case of Wason's Selection Task
por: Abe, Hirohiko, et al.
Publicado: (2026)
por: Abe, Hirohiko, et al.
Publicado: (2026)
Can GPT models Follow Human Summarization Guidelines? A Study for Targeted Communication Goals
por: Zhou, Yongxin, et al.
Publicado: (2023)
por: Zhou, Yongxin, et al.
Publicado: (2023)
THaLLE-ThaiLLM: Domain-Specialized Small LLMs for Finance and Thai -- Technical Report
por: Labs, KBTG, et al.
Publicado: (2026)
por: Labs, KBTG, et al.
Publicado: (2026)
PSentScore: Evaluating Sentiment Polarity in Dialogue Summarization
por: Zhou, Yongxin, et al.
Publicado: (2023)
por: Zhou, Yongxin, et al.
Publicado: (2023)
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
por: Méloux, Maxime, et al.
Publicado: (2025)
por: Méloux, Maxime, et al.
Publicado: (2025)
Enabling Vibration-Based Gesture Recognition on Everyday Furniture via Energy-Efficient FPGA Implementation of 1D Convolutional Networks
por: Shibata, Koki, et al.
Publicado: (2025)
por: Shibata, Koki, et al.
Publicado: (2025)
Where Do Self-Supervised Speech Models Become Unfair?
por: Herron, Felix, et al.
Publicado: (2026)
por: Herron, Felix, et al.
Publicado: (2026)
What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search
por: Zhang, Xinhao, et al.
Publicado: (2026)
por: Zhang, Xinhao, et al.
Publicado: (2026)
Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models
por: Herron, Felix, et al.
Publicado: (2026)
por: Herron, Felix, et al.
Publicado: (2026)
Normative Reasoning in Large Language Models: A Comparative Benchmark from Logical and Modal Perspectives
por: Ozeki, Kentaro, et al.
Publicado: (2025)
por: Ozeki, Kentaro, et al.
Publicado: (2025)
ixi-GEN: Efficient Industrial sLLMs through Domain Adaptive Continual Pretraining
por: Kim, Seonwu, et al.
Publicado: (2025)
por: Kim, Seonwu, et al.
Publicado: (2025)
Checkpoint Merging via Bayesian Optimization in LLM Pretraining
por: Liu, Deyuan, et al.
Publicado: (2024)
por: Liu, Deyuan, et al.
Publicado: (2024)
Abductive Reasoning with Syllogistic Forms in Large Language Models
por: Abe, Hirohiko, et al.
Publicado: (2026)
por: Abe, Hirohiko, et al.
Publicado: (2026)
Syntactic Learnability of Echo State Neural Language Models at Scale
por: Ueda, Ryo, et al.
Publicado: (2025)
por: Ueda, Ryo, et al.
Publicado: (2025)
Efficient Domain-adaptive Continual Pretraining for the Process Industry in the German Language
por: Zhukova, Anastasia, et al.
Publicado: (2025)
por: Zhukova, Anastasia, et al.
Publicado: (2025)
Channel Merging: Preserving Specialization for Merged Experts
por: Zhang, Mingyang, et al.
Publicado: (2024)
por: Zhang, Mingyang, et al.
Publicado: (2024)
Pretraining and Updates of Domain-Specific LLM: A Case Study in the Japanese Business Domain
por: Takahashi, Kosuke, et al.
Publicado: (2024)
por: Takahashi, Kosuke, et al.
Publicado: (2024)
Exploring Reasoning Biases in Large Language Models Through Syllogism: Insights from the NeuBAROCO Dataset
por: Ozeki, Kentaro, et al.
Publicado: (2024)
por: Ozeki, Kentaro, et al.
Publicado: (2024)
Beyond Fine-tuning: Unleashing the Potential of Continuous Pretraining for Clinical LLMs
por: Christophe, Clément, et al.
Publicado: (2024)
por: Christophe, Clément, et al.
Publicado: (2024)
ELO: Efficient Layer-Specific Optimization for Continual Pretraining of Multilingual LLMs
por: Yoo, HanGyeol, et al.
Publicado: (2026)
por: Yoo, HanGyeol, et al.
Publicado: (2026)
FinanceMath: Knowledge-Intensive Math Reasoning in Finance Domains
por: Zhao, Yilun, et al.
Publicado: (2023)
por: Zhao, Yilun, et al.
Publicado: (2023)
Continual Pretraining on Encrypted Synthetic Data for Privacy-Preserving LLMs
por: Liu, Honghao, et al.
Publicado: (2026)
por: Liu, Honghao, et al.
Publicado: (2026)
Responsible Benchmarking of Fairness for Automatic Speech Recognition
por: Herron, Felix, et al.
Publicado: (2026)
por: Herron, Felix, et al.
Publicado: (2026)
Compressing Language Models for Specialized Domains
por: Williams, Miles, et al.
Publicado: (2025)
por: Williams, Miles, et al.
Publicado: (2025)
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
por: Itzhak, Itay, et al.
Publicado: (2025)
por: Itzhak, Itay, et al.
Publicado: (2025)
Assessing the Political Fairness of Multilingual LLMs: A Case Study based on a 21-way Multiparallel EuroParl Dataset
por: Lerner, Paul, et al.
Publicado: (2025)
por: Lerner, Paul, et al.
Publicado: (2025)
Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks
por: Wang, Zheng, et al.
Publicado: (2024)
por: Wang, Zheng, et al.
Publicado: (2024)
IKnow: Instruction-Knowledge-Aware Continual Pretraining for Effective Domain Adaptation
por: Zhang, Tianyi, et al.
Publicado: (2025)
por: Zhang, Tianyi, et al.
Publicado: (2025)
The Thinking Spectrum: An Empirical Study of Tunable Reasoning in LLMs through Model Merging
por: Lan, Xiaochong, et al.
Publicado: (2025)
por: Lan, Xiaochong, et al.
Publicado: (2025)
Rethinking Multilingual Continual Pretraining: Data Mixing for Adapting LLMs Across Languages and Resources
por: Li, Zihao, et al.
Publicado: (2025)
por: Li, Zihao, et al.
Publicado: (2025)
Evaluating the Effectiveness of Linguistic Knowledge in Pretrained Language Models: A Case Study of Universal Dependencies
por: Li, Wenxi
Publicado: (2025)
por: Li, Wenxi
Publicado: (2025)
Ingest-And-Ground: Dispelling Hallucinations from Continually-Pretrained LLMs with RAG
por: Fang, Chenhao, et al.
Publicado: (2024)
por: Fang, Chenhao, et al.
Publicado: (2024)
DogeRM: Equipping Reward Models with Domain Knowledge through Model Merging
por: Lin, Tzu-Han, et al.
Publicado: (2024)
por: Lin, Tzu-Han, et al.
Publicado: (2024)
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
por: Méloux, Maxime, et al.
Publicado: (2025)
por: Méloux, Maxime, et al.
Publicado: (2025)
How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion
por: Seth, Agrima, et al.
Publicado: (2025)
por: Seth, Agrima, et al.
Publicado: (2025)
Ejemplares similares
-
The Construction of Instruction-tuned LLMs for Finance without Instruction Data Using Continual Pretraining and Model Merging
por: Hirano, Masanori, et al.
Publicado: (2024) -
Construction of Domain-specified Japanese Large Language Model for Finance through Continual Pre-training
por: Hirano, Masanori, et al.
Publicado: (2024) -
MedMeta: A Benchmark for LLMs in Synthesizing Meta-Analysis Conclusion from Medical Studies
por: Ha, Huy Hoang, et al.
Publicado: (2026) -
Automated Clinical Report Generation for Remote Cognitive Remediation: Comparing Knowledge-Engineered Templates and LLMs in Low-Resource Settings
por: Zhou, Yongxin, et al.
Publicado: (2026) -
Is Biomedical Specialization Still Worth It? Insights from Domain-Adaptive Language Modelling with a New French Health Corpus
por: Mannion, Aidan, et al.
Publicado: (2026)