Toward LLMs Beyond English-Centric Development
Fuente:
arXiv
Saved in:
| Main Authors: | Takase, Sho, Honda, Ukyo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Revisiting the Capacity Gap in Chain-of-Thought Distillation from a Practical Perspective
by: Kajitsuka, Tokio, et al.
Published: (2026)
by: Kajitsuka, Tokio, et al.
Published: (2026)
Does Self-Consistency Improve the Recall of Encyclopedic Knowledge?
by: Hoshino, Sho, et al.
Published: (2026)
by: Hoshino, Sho, et al.
Published: (2026)
A Single Linear Layer Yields Task-Adapted Low-Rank Matrices
by: Kim, Hwichan, et al.
Published: (2024)
by: Kim, Hwichan, et al.
Published: (2024)
Exploring Explanations Improves the Robustness of In-Context Learning
by: Honda, Ukyo, et al.
Published: (2025)
by: Honda, Ukyo, et al.
Published: (2025)
Annotation-Efficient Language Model Alignment via Diverse and Representative Response Texts
by: Jinnai, Yuu, et al.
Published: (2024)
by: Jinnai, Yuu, et al.
Published: (2024)
Distilling Many-Shot In-Context Learning into a Cheat Sheet
by: Honda, Ukyo, et al.
Published: (2025)
by: Honda, Ukyo, et al.
Published: (2025)
FaithCAMERA: Construction of a Faithful Dataset for Ad Text Generation
by: Kato, Akihiko, et al.
Published: (2024)
by: Kato, Akihiko, et al.
Published: (2024)
Reinforcement Learning for Edit-Based Non-Autoregressive Neural Machine Translation
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Exploring the Relationship Between Diversity and Quality in Ad Text Generation
by: Aoki, Yoichi, et al.
Published: (2025)
by: Aoki, Yoichi, et al.
Published: (2025)
On the True Distribution Approximation of Minimum Bayes-Risk Decoding
by: Ohashi, Atsumoto, et al.
Published: (2024)
by: Ohashi, Atsumoto, et al.
Published: (2024)
Generating Diverse and High-Quality Texts by Minimum Bayes Risk Decoding
by: Jinnai, Yuu, et al.
Published: (2024)
by: Jinnai, Yuu, et al.
Published: (2024)
Not Eliminate but Aggregate: Post-Hoc Control over Mixture-of-Experts to Address Shortcut Shifts in Natural Language Understanding
by: Honda, Ukyo, et al.
Published: (2024)
by: Honda, Ukyo, et al.
Published: (2024)
Natural Fingerprints of Large Language Models
by: Suzuki, Teppei, et al.
Published: (2025)
by: Suzuki, Teppei, et al.
Published: (2025)
Self-Translate-Train: Enhancing Cross-Lingual Transfer of Large Language Models via Inherent Capability
by: Ri, Ryokan, et al.
Published: (2024)
by: Ri, Ryokan, et al.
Published: (2024)
Model-Based Minimum Bayes Risk Decoding for Text Generation
by: Jinnai, Yuu, et al.
Published: (2023)
by: Jinnai, Yuu, et al.
Published: (2023)
Beyond English-Centric Training: How Reinforcement Learning Improves Cross-Lingual Reasoning in LLMs
by: Huang, Shulin, et al.
Published: (2025)
by: Huang, Shulin, et al.
Published: (2025)
Beyond English-Centric LLMs: What Language Do Multilingual Language Models Think in?
by: Zhong, Chengzhi, et al.
Published: (2024)
by: Zhong, Chengzhi, et al.
Published: (2024)
Large Vocabulary Size Improves Large Language Models
by: Takase, Sho, et al.
Published: (2024)
by: Takase, Sho, et al.
Published: (2024)
Scaling Laws for Upcycling Mixture-of-Experts Language Models
by: Liew, Seng Pei, et al.
Published: (2025)
by: Liew, Seng Pei, et al.
Published: (2025)
Spike No More: Stabilizing the Pre-training of Large Language Models
by: Takase, Sho, et al.
Published: (2023)
by: Takase, Sho, et al.
Published: (2023)
Efficient Construction of Model Family through Progressive Training Using Model Expansion
by: Yano, Kazuki, et al.
Published: (2025)
by: Yano, Kazuki, et al.
Published: (2025)
LLMs Beyond English: Scaling the Multilingual Capability of LLMs with Cross-Lingual Feedback
by: Lai, Wen, et al.
Published: (2024)
by: Lai, Wen, et al.
Published: (2024)
MEXA: Multilingual Evaluation of English-Centric LLMs via Cross-Lingual Alignment
by: Kargaran, Amir Hossein, et al.
Published: (2024)
by: Kargaran, Amir Hossein, et al.
Published: (2024)
From English-Centric to Effective Bilingual: LLMs with Custom Tokenizers for Underrepresented Languages
by: Kiulian, Artur, et al.
Published: (2024)
by: Kiulian, Artur, et al.
Published: (2024)
Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-Tuning
by: Yano, Kazuki, et al.
Published: (2026)
by: Yano, Kazuki, et al.
Published: (2026)
Beyond English: The Impact of Prompt Translation Strategies across Languages and Tasks in Multilingual LLMs
by: Mondshine, Itai, et al.
Published: (2025)
by: Mondshine, Itai, et al.
Published: (2025)
Enhancing Non-English Capabilities of English-Centric Large Language Models through Deep Supervision Fine-Tuning
by: Huo, Wenshuai, et al.
Published: (2025)
by: Huo, Wenshuai, et al.
Published: (2025)
Optimizing Korean-Centric LLMs via Token Pruning
by: Kim, Hoyeol, et al.
Published: (2026)
by: Kim, Hoyeol, et al.
Published: (2026)
Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models
by: Nie, Ercong, et al.
Published: (2025)
by: Nie, Ercong, et al.
Published: (2025)
A Survey on Human-Centric LLMs
by: Wang, Jing Yi, et al.
Published: (2024)
by: Wang, Jing Yi, et al.
Published: (2024)
Answer-Centric or Reasoning-Driven? Uncovering the Latent Memory Anchor in LLMs
by: Wu, Yang, et al.
Published: (2025)
by: Wu, Yang, et al.
Published: (2025)
Measuring Affinity between Attention-Head Weight Subspaces via the Projection Kernel
by: Yamagiwa, Hiroaki, et al.
Published: (2026)
by: Yamagiwa, Hiroaki, et al.
Published: (2026)
Language Model Maps for Prompt-Response Distributions via Log-Likelihood Vectors
by: Takase, Yusuke, et al.
Published: (2026)
by: Takase, Yusuke, et al.
Published: (2026)
Axis Tour: Word Tour Determines the Order of Axes in ICA-transformed Embeddings
by: Yamagiwa, Hiroaki, et al.
Published: (2024)
by: Yamagiwa, Hiroaki, et al.
Published: (2024)
RMTBench: Benchmarking LLMs Through Multi-Turn User-Centric Role-Playing
by: Xiang, Hao, et al.
Published: (2025)
by: Xiang, Hao, et al.
Published: (2025)
Desert Camels and Oil Sheikhs: Arab-Centric Red Teaming of Frontier LLMs
by: Saeed, Muhammed, et al.
Published: (2024)
by: Saeed, Muhammed, et al.
Published: (2024)
Generalized Category Discovery in Event-Centric Contexts: Latent Pattern Mining with LLMs
by: Luo, Yi, et al.
Published: (2025)
by: Luo, Yi, et al.
Published: (2025)
Which English Do LLMs Prefer? Triangulating Structural Bias Towards American English in Foundation Models
by: Nayeem, Mir Tafseer, et al.
Published: (2026)
by: Nayeem, Mir Tafseer, et al.
Published: (2026)
CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning
by: Ibrahim, George, et al.
Published: (2025)
by: Ibrahim, George, et al.
Published: (2025)
Enabling Doctor-Centric Medical AI with LLMs through Workflow-Aligned Tasks and Benchmarks
by: Xie, Wenya, et al.
Published: (2025)
by: Xie, Wenya, et al.
Published: (2025)
Similar Items
-
Revisiting the Capacity Gap in Chain-of-Thought Distillation from a Practical Perspective
by: Kajitsuka, Tokio, et al.
Published: (2026) -
Does Self-Consistency Improve the Recall of Encyclopedic Knowledge?
by: Hoshino, Sho, et al.
Published: (2026) -
A Single Linear Layer Yields Task-Adapted Low-Rank Matrices
by: Kim, Hwichan, et al.
Published: (2024) -
Exploring Explanations Improves the Robustness of In-Context Learning
by: Honda, Ukyo, et al.
Published: (2025) -
Annotation-Efficient Language Model Alignment via Diverse and Representative Response Texts
by: Jinnai, Yuu, et al.
Published: (2024)