StatLLaMA: Multi-Stage training for domain-optimized statistical large language models
Fuente:
arXiv
Saved in:
| Main Authors: | Zeng, Jing-Yi, Huang, Guan-Hua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
by: Xia, Mengzhou, et al.
Published: (2023)
by: Xia, Mengzhou, et al.
Published: (2023)
Me LLaMA: Foundation Large Language Models for Medical Applications
by: Xie, Qianqian, et al.
Published: (2024)
by: Xie, Qianqian, et al.
Published: (2024)
LLaMA Beyond English: An Empirical Study on Language Capability Transfer
by: Zhao, Jun, et al.
Published: (2024)
by: Zhao, Jun, et al.
Published: (2024)
How Vocabulary Sharing Facilitates Multilingualism in LLaMA?
by: Yuan, Fei, et al.
Published: (2023)
by: Yuan, Fei, et al.
Published: (2023)
Fine-tuning large language models for domain adaptation: Exploration of training strategies, scaling, model merging and synergistic capabilities
by: Lu, Wei, et al.
Published: (2024)
by: Lu, Wei, et al.
Published: (2024)
LLaMAs Have Feelings Too: Unveiling Sentiment and Emotion Representations in LLaMA Models Through Probing
by: Di Palma, Dario, et al.
Published: (2025)
by: Di Palma, Dario, et al.
Published: (2025)
BanglaLlama: LLaMA for Bangla Language
by: Zehady, Abdullah Khan, et al.
Published: (2024)
by: Zehady, Abdullah Khan, et al.
Published: (2024)
Multi-round jailbreak attack on large language models
by: Zhou, Yihua, et al.
Published: (2024)
by: Zhou, Yihua, et al.
Published: (2024)
WizardLM: Empowering large pre-trained language models to follow complex instructions
by: Xu, Can, et al.
Published: (2023)
by: Xu, Can, et al.
Published: (2023)
ChatGPT vs Gemini vs LLaMA on Multilingual Sentiment Analysis
by: Buscemi, Alessio, et al.
Published: (2024)
by: Buscemi, Alessio, et al.
Published: (2024)
LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning
by: Zhang, Di, et al.
Published: (2024)
by: Zhang, Di, et al.
Published: (2024)
Dissociating language and thought in large language models
by: Mahowald, Kyle, et al.
Published: (2023)
by: Mahowald, Kyle, et al.
Published: (2023)
Post-training makes large language models less human-like
by: Binz, Marcel, et al.
Published: (2026)
by: Binz, Marcel, et al.
Published: (2026)
The impact of fine tuning in LLaMA on hallucinations for named entity extraction in legal documentation
by: Vargas, Francisco, et al.
Published: (2025)
by: Vargas, Francisco, et al.
Published: (2025)
Streamlining evidence based clinical recommendations with large language models
by: Li, Dubai, et al.
Published: (2025)
by: Li, Dubai, et al.
Published: (2025)
Empowering Smaller Models: Tuning LLaMA and Gemma with Chain-of-Thought for Ukrainian Exam Tasks
by: Syromiatnikov, Mykyta, et al.
Published: (2025)
by: Syromiatnikov, Mykyta, et al.
Published: (2025)
On the attribution of confidence to large language models
by: Keeling, Geoff, et al.
Published: (2024)
by: Keeling, Geoff, et al.
Published: (2024)
MusiScene: Leveraging MU-LLaMA for Scene Imagination and Enhanced Video Background Music Generation
by: Izzati, Fathinah, et al.
Published: (2025)
by: Izzati, Fathinah, et al.
Published: (2025)
Metamorphic Testing for Fairness Evaluation in Large Language Models: Identifying Intersectional Bias in LLaMA and GPT
by: Reddy, Harishwar, et al.
Published: (2025)
by: Reddy, Harishwar, et al.
Published: (2025)
LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement
by: Kang, Boyi, et al.
Published: (2025)
by: Kang, Boyi, et al.
Published: (2025)
Advanced Natural-based interaction for the ITAlian language: LLaMAntino-3-ANITA
by: Polignano, Marco, et al.
Published: (2024)
by: Polignano, Marco, et al.
Published: (2024)
Resource-Efficient Fine-Tuning of LLaMA-3.2-3B for Medical Chain-of-Thought Reasoning
by: Mansha, Imran
Published: (2025)
by: Mansha, Imran
Published: (2025)
Does Tone Change the Answer? Evaluating Prompt Politeness Effects on Modern LLMs: GPT, Gemini, and LLaMA
by: Cai, Hanyu, et al.
Published: (2025)
by: Cai, Hanyu, et al.
Published: (2025)
Principled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4
by: Bsharat, Sondos Mahmoud, et al.
Published: (2023)
by: Bsharat, Sondos Mahmoud, et al.
Published: (2023)
Long-form factuality in large language models
by: Wei, Jerry, et al.
Published: (2024)
by: Wei, Jerry, et al.
Published: (2024)
Uncovering inequalities in new knowledge learning by large language models across different languages
by: Wang, Chenglong, et al.
Published: (2025)
by: Wang, Chenglong, et al.
Published: (2025)
Bringing legal knowledge to the public by constructing a legal question bank using large-scale pre-trained language model
by: Yuan, Mingruo, et al.
Published: (2025)
by: Yuan, Mingruo, et al.
Published: (2025)
Large language models in healthcare and medical domain: A review
by: Nazi, Zabir Al, et al.
Published: (2023)
by: Nazi, Zabir Al, et al.
Published: (2023)
Quantifying non deterministic drift in large language models
by: Nicholson, Claire
Published: (2026)
by: Nicholson, Claire
Published: (2026)
Can large language models build causal graphs?
by: Long, Stephanie, et al.
Published: (2023)
by: Long, Stephanie, et al.
Published: (2023)
Response: Emergent analogical reasoning in large language models
by: Hodel, Damian, et al.
Published: (2023)
by: Hodel, Damian, et al.
Published: (2023)
The 20 questions game to distinguish large language models
by: Richardeau, Gurvan, et al.
Published: (2024)
by: Richardeau, Gurvan, et al.
Published: (2024)
Representation in large language models
by: Yetman, Cameron
Published: (2025)
by: Yetman, Cameron
Published: (2025)
MindScope: Exploring cognitive biases in large language models through Multi-Agent Systems
by: Xie, Zhentao, et al.
Published: (2024)
by: Xie, Zhentao, et al.
Published: (2024)
LLaMA-Excitor: General Instruction Tuning via Indirect Feature Interaction
by: Zou, Bo, et al.
Published: (2024)
by: Zou, Bo, et al.
Published: (2024)
LLaMA-E: Empowering E-commerce Authoring with Object-Interleaved Instruction Following
by: Shi, Kaize, et al.
Published: (2023)
by: Shi, Kaize, et al.
Published: (2023)
Failure of contextual invariance in large language models
by: Kumar, Sagar, et al.
Published: (2026)
by: Kumar, Sagar, et al.
Published: (2026)
Emission-GPT: A domain-specific language model agent for knowledge retrieval, emission inventory and data analysis
by: Ye, Jiashu, et al.
Published: (2025)
by: Ye, Jiashu, et al.
Published: (2025)
A survey of textual cyber abuse detection using cutting-edge language models and large language models
by: Diaz-Garcia, Jose A., et al.
Published: (2025)
by: Diaz-Garcia, Jose A., et al.
Published: (2025)
Re-evaluating Theory of Mind evaluation in large language models
by: Hu, Jennifer, et al.
Published: (2025)
by: Hu, Jennifer, et al.
Published: (2025)
Similar Items
-
Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
by: Xia, Mengzhou, et al.
Published: (2023) -
Me LLaMA: Foundation Large Language Models for Medical Applications
by: Xie, Qianqian, et al.
Published: (2024) -
LLaMA Beyond English: An Empirical Study on Language Capability Transfer
by: Zhao, Jun, et al.
Published: (2024) -
How Vocabulary Sharing Facilitates Multilingualism in LLaMA?
by: Yuan, Fei, et al.
Published: (2023) -
Fine-tuning large language models for domain adaptation: Exploration of training strategies, scaling, model merging and synergistic capabilities
by: Lu, Wei, et al.
Published: (2024)