Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary Expansion
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Jianqing, Huang, Huang, Lin, Zhihang, Liang, Juhao, Tang, Zhengyang, Almubarak, Khalid, Alharthik, Abdulmohsen, An, Bang, He, Juncai, Wu, Xiangbo, Yu, Fei, Chen, Junying, Ma, Zhuoheng, Du, Yuhao, Zhang, He, Alghamdi, Emad A., Zhang, Lian, Sun, Ruoyu, Li, Haizhou, Wang, Benyou, Xu, Jinchao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Alignment at Pre-training! Towards Native Alignment for Arabic LLMs
by: Liang, Juhao, et al.
Published: (2024)
by: Liang, Juhao, et al.
Published: (2024)
AceGPT, Localizing Large Language Models in Arabic
by: Huang, Huang, et al.
Published: (2023)
by: Huang, Huang, et al.
Published: (2023)
Online Training of Large Language Models: Learn while chatting
by: Liang, Juhao, et al.
Published: (2024)
by: Liang, Juhao, et al.
Published: (2024)
MgNet: A Unified Framework of Multigrid and Convolutional Neural Network
by: He, Juncai, et al.
Published: (2019)
by: He, Juncai, et al.
Published: (2019)
Deep Neural Networks and Finite Elements of Any Order on Arbitrary Dimensions
by: He, Juncai, et al.
Published: (2023)
by: He, Juncai, et al.
Published: (2023)
Mind the Gap: A Review of Arabic Post-Training Datasets and Their Limitations
by: Alkhowaiter, Mohammed, et al.
Published: (2025)
by: Alkhowaiter, Mohammed, et al.
Published: (2025)
Self-Composing Neural Operators with Depth and Accuracy Scaling via Adaptive Train-and-Unroll Approach
by: He, Juncai, et al.
Published: (2025)
by: He, Juncai, et al.
Published: (2025)
Expressivity and Approximation Properties of Deep Neural Networks with ReLU$^k$ Activation
by: He, Juncai, et al.
Published: (2023)
by: He, Juncai, et al.
Published: (2023)
MgNO: Efficient Parameterization of Linear Operators via Multigrid
by: He, Juncai, et al.
Published: (2023)
by: He, Juncai, et al.
Published: (2023)
Smurfs: Multi-Agent System using Context-Efficient DFSDT for Tool Planning
by: Chen, Junzhi, et al.
Published: (2024)
by: Chen, Junzhi, et al.
Published: (2024)
What makes video‐based academic lectures difficult for language learners to comprehend? The role of multimodal complexity
by: Emad A. Alghamdi
Published: (2024)
by: Emad A. Alghamdi
Published: (2024)
AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic
by: Alghamdi, Emad A., et al.
Published: (2024)
by: Alghamdi, Emad A., et al.
Published: (2024)
ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic
by: Koto, Fajri, et al.
Published: (2024)
by: Koto, Fajri, et al.
Published: (2024)
EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
by: Zhang, Yuhao, et al.
Published: (2025)
by: Zhang, Yuhao, et al.
Published: (2025)
Efficiently Democratizing Medical LLMs for 50 Languages via a Mixture of Language Family Experts
by: Zheng, Guorui, et al.
Published: (2024)
by: Zheng, Guorui, et al.
Published: (2024)
Deeper or Wider: A Perspective from Optimal Generalization Error with Sobolev Loss
by: Yang, Yahong, et al.
Published: (2024)
by: Yang, Yahong, et al.
Published: (2024)
Deep Neural Networks with General Activations: Super-Convergence in Sobolev Norms
by: Yang, Yahong, et al.
Published: (2025)
by: Yang, Yahong, et al.
Published: (2025)
Roadmap towards Superhuman Speech Understanding using Large Language Models
by: Bu, Fan, et al.
Published: (2024)
by: Bu, Fan, et al.
Published: (2024)
Soundwave: Less is More for Speech-Text Alignment in LLMs
by: Zhang, Yuhao, et al.
Published: (2025)
by: Zhang, Yuhao, et al.
Published: (2025)
Structure Disruption: Subverting Malicious Diffusion-Based Inpainting via Self-Attention Query Perturbation
by: He, Yuhao, et al.
Published: (2025)
by: He, Yuhao, et al.
Published: (2025)
Davenport-Heilbronn Function Ratio Properties and Non-Trivial Zeros Study
by: Liu, Tao, et al.
Published: (2025)
by: Liu, Tao, et al.
Published: (2025)
CIDAR: Culturally Relevant Instruction Dataset For Arabic
by: Alyafeai, Zaid, et al.
Published: (2024)
by: Alyafeai, Zaid, et al.
Published: (2024)
BALSAM: A Platform for Benchmarking Arabic Large Language Models
by: Al-Matham, Rawan, et al.
Published: (2025)
by: Al-Matham, Rawan, et al.
Published: (2025)
From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test
by: Dai, Xunlian, et al.
Published: (2025)
by: Dai, Xunlian, et al.
Published: (2025)
FedSLoP: Memory-Efficient Federated Learning with Low-Rank Gradient Projection
by: He, Yutong, et al.
Published: (2026)
by: He, Yutong, et al.
Published: (2026)
Divergence-free Linearized Neural Networks: Integral Representation and Optimal Approximation Rates
by: He, Juncai, et al.
Published: (2026)
by: He, Juncai, et al.
Published: (2026)
Interdependence of Venture Capital Screening Criteria: A Decision Support Framework
by: Norah Almubarak, et al.
Published: (2026)
by: Norah Almubarak, et al.
Published: (2026)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
by: Tang, Zhengyang, et al.
Published: (2024)
by: Tang, Zhengyang, et al.
Published: (2024)
MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
by: Du, Yuhao, et al.
Published: (2025)
by: Du, Yuhao, et al.
Published: (2025)
Ti-Audio: The First Multi-Dialectal End-to-End Speech LLM for Tibetan
by: Wang, Jialing, et al.
Published: (2026)
by: Wang, Jialing, et al.
Published: (2026)
Neural Flow Operators can Approximate any Operator: Abstract Frameworks and Universal Approximations
by: Chen, Shuang, et al.
Published: (2026)
by: Chen, Shuang, et al.
Published: (2026)
ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
by: Chen, Guiming Hardy, et al.
Published: (2024)
by: Chen, Guiming Hardy, et al.
Published: (2024)
Towards Output-Optimal Uniform Sampling and Approximate Counting for Join-Project Queries
by: Hu, Xiao, et al.
Published: (2026)
by: Hu, Xiao, et al.
Published: (2026)
DIPS: Optimal Dynamic Index for Poisson $\boldsymbolπ$ps Sampling
by: Huang, Jinchao, et al.
Published: (2024)
by: Huang, Jinchao, et al.
Published: (2024)
Non-degenerate Ground State of the Spin-Boson Model under Abelian Diagonalization
by: Liu, Tao, et al.
Published: (2025)
by: Liu, Tao, et al.
Published: (2025)
CSLRConformer: A Data-Centric Conformer Approach for Continuous Arabic Sign Language Recognition on the Isharah Datase
by: Elden, Fatimah Mohamed Emad
Published: (2025)
by: Elden, Fatimah Mohamed Emad
Published: (2025)
Quantifying Self-diagnostic Atomic Knowledge in Chinese Medical Foundation Model: A Computational Analysis
by: Fan, Yaxin, et al.
Published: (2023)
by: Fan, Yaxin, et al.
Published: (2023)
Beyond Binary: Towards Fine-Grained LLM-Generated Text Detection via Role Recognition and Involvement Measurement
by: Cheng, Zihao, et al.
Published: (2024)
by: Cheng, Zihao, et al.
Published: (2024)
A Novel Strategy to Prepare Hierarchically Porous Phenolic Resin‐Based Carbon Aerogel With Excellent Adsorption Performance
by: Qiaomu Zhang, et al.
Published: (2024)
by: Qiaomu Zhang, et al.
Published: (2024)
Self-Evolving Critique Abilities in Large Language Models
by: Tang, Zhengyang, et al.
Published: (2025)
by: Tang, Zhengyang, et al.
Published: (2025)
Similar Items
-
Alignment at Pre-training! Towards Native Alignment for Arabic LLMs
by: Liang, Juhao, et al.
Published: (2024) -
AceGPT, Localizing Large Language Models in Arabic
by: Huang, Huang, et al.
Published: (2023) -
Online Training of Large Language Models: Learn while chatting
by: Liang, Juhao, et al.
Published: (2024) -
MgNet: A Unified Framework of Multigrid and Convolutional Neural Network
by: He, Juncai, et al.
Published: (2019) -
Deep Neural Networks and Finite Elements of Any Order on Arbitrary Dimensions
by: He, Juncai, et al.
Published: (2023)