Investigating the Multilingual Calibration Effects of Language Model Instruction-Tuning
Fuente:
arXiv
Guardado en:
| Autores principales: | Huang, Jerry, Lu, Peng, Zeng, Qiuhao, Iwasawa, Yusuke, Matsuo, Yutaka, Chandar, Sarath, Marrese-Taylor, Edison, Li, Irene |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
When Instructions Multiply: Measuring and Estimating LLM Capabilities of Multiple Instructions Following
por: Harada, Keno, et al.
Publicado: (2025)
por: Harada, Keno, et al.
Publicado: (2025)
MKG-Rank: Enhancing Large Language Models with Knowledge Graph for Multilingual Medical Question Answering
por: Li, Feiyang, et al.
Publicado: (2025)
por: Li, Feiyang, et al.
Publicado: (2025)
On the Multilingual Ability of Decoder-based Pre-trained Language Models: Finding and Controlling Language-Specific Neurons
por: Kojima, Takeshi, et al.
Publicado: (2024)
por: Kojima, Takeshi, et al.
Publicado: (2024)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
por: Gambardella, Andrew, et al.
Publicado: (2024)
por: Gambardella, Andrew, et al.
Publicado: (2024)
Image-Text Relation Prediction for Multilingual Tweets
por: Rikters, Matīss, et al.
Publicado: (2025)
por: Rikters, Matīss, et al.
Publicado: (2025)
Large Language Models as Theory of Mind Aware Generative Agents with Counterfactual Reflection
por: Yang, Bo, et al.
Publicado: (2025)
por: Yang, Bo, et al.
Publicado: (2025)
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
por: Gambardella, Andrew, et al.
Publicado: (2025)
por: Gambardella, Andrew, et al.
Publicado: (2025)
Multilingual Definition Modeling
por: Marrese-Taylor, Edison, et al.
Publicado: (2025)
por: Marrese-Taylor, Edison, et al.
Publicado: (2025)
Calibrated Language Models and How to Find Them with Label Smoothing
por: Huang, Jerry, et al.
Publicado: (2025)
por: Huang, Jerry, et al.
Publicado: (2025)
Automated Refinement of Essay Scoring Rubrics for Language Models via Reflect-and-Revise
por: Harada, Keno, et al.
Publicado: (2025)
por: Harada, Keno, et al.
Publicado: (2025)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
por: Huang, Jerry, et al.
Publicado: (2024)
por: Huang, Jerry, et al.
Publicado: (2024)
EpiK-Eval: Evaluation for Language Models as Epistemic Models
por: Prato, Gabriele, et al.
Publicado: (2023)
por: Prato, Gabriele, et al.
Publicado: (2023)
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
por: Cao, Qi, et al.
Publicado: (2026)
por: Cao, Qi, et al.
Publicado: (2026)
Probabilistic Calibration Is a Trainable Capability in Language Models
por: Baldelli, Davide, et al.
Publicado: (2026)
por: Baldelli, Davide, et al.
Publicado: (2026)
KG-Rank: Enhancing Large Language Models for Medical QA with Knowledge Graphs and Ranking Techniques
por: Yang, Rui, et al.
Publicado: (2024)
por: Yang, Rui, et al.
Publicado: (2024)
Do Large Language Models Know How Much They Know?
por: Prato, Gabriele, et al.
Publicado: (2025)
por: Prato, Gabriele, et al.
Publicado: (2025)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
por: Minegishi, Gouki, et al.
Publicado: (2025)
por: Minegishi, Gouki, et al.
Publicado: (2025)
Answer When Needed, Forget When Not: Language Models Pretend to Forget via In-Context Knowledge Unlearning
por: Takashiro, Shota, et al.
Publicado: (2024)
por: Takashiro, Shota, et al.
Publicado: (2024)
Towards Practical Tool Usage for Continually Learning LLMs
por: Huang, Jerry, et al.
Publicado: (2024)
por: Huang, Jerry, et al.
Publicado: (2024)
Faithfulness Measurable Masked Language Models
por: Madsen, Andreas, et al.
Publicado: (2023)
por: Madsen, Andreas, et al.
Publicado: (2023)
Dynamic Injection of Entity Knowledge into Dense Retrievers
por: Yamada, Ikuya, et al.
Publicado: (2025)
por: Yamada, Ikuya, et al.
Publicado: (2025)
MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation
por: Xuan, Weihao, et al.
Publicado: (2025)
por: Xuan, Weihao, et al.
Publicado: (2025)
Beyond In-Distribution Success: Scaling Curves of CoT Granularity for Language Model Generalization
por: Wang, Ru, et al.
Publicado: (2025)
por: Wang, Ru, et al.
Publicado: (2025)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
por: Takashiro, Shota, et al.
Publicado: (2026)
por: Takashiro, Shota, et al.
Publicado: (2026)
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
por: Minegishi, Gouki, et al.
Publicado: (2025)
por: Minegishi, Gouki, et al.
Publicado: (2025)
Which Programming Language and What Features at Pre-training Stage Affect Downstream Logical Inference Performance?
por: Uchiyama, Fumiya, et al.
Publicado: (2024)
por: Uchiyama, Fumiya, et al.
Publicado: (2024)
Are self-explanations from Large Language Models faithful?
por: Madsen, Andreas, et al.
Publicado: (2024)
por: Madsen, Andreas, et al.
Publicado: (2024)
Investigating Multilingual Instruction-Tuning: Do Polyglot Models Demand for Multilingual Instructions?
por: Weber, Alexander Arno, et al.
Publicado: (2024)
por: Weber, Alexander Arno, et al.
Publicado: (2024)
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
por: Prato, Gabriele, et al.
Publicado: (2025)
por: Prato, Gabriele, et al.
Publicado: (2025)
Annotations for Exploring Food Tweets From Multiple Aspects
por: Rikters, Matīss, et al.
Publicado: (2024)
por: Rikters, Matīss, et al.
Publicado: (2024)
Self-Harmony: Learning to Harmonize Self-Supervision and Self-Play in Test-Time Reinforcement Learning
por: Wang, Ru, et al.
Publicado: (2025)
por: Wang, Ru, et al.
Publicado: (2025)
On the Role of Unobserved Sequences on Sample-based Uncertainty Quantification for LLMs
por: Kunitomo-Jacquin, Lucie, et al.
Publicado: (2025)
por: Kunitomo-Jacquin, Lucie, et al.
Publicado: (2025)
Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models
por: Gu, Xiaojie, et al.
Publicado: (2026)
por: Gu, Xiaojie, et al.
Publicado: (2026)
LLMs Can't Play Hangman: On the Necessity of a Private Working Memory for Language Agents
por: Baldelli, Davide, et al.
Publicado: (2026)
por: Baldelli, Davide, et al.
Publicado: (2026)
Bridging Lottery Ticket and Grokking: Understanding Grokking from Inner Structure of Networks
por: Minegishi, Gouki, et al.
Publicado: (2023)
por: Minegishi, Gouki, et al.
Publicado: (2023)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
por: Huang, Jerry, et al.
Publicado: (2024)
por: Huang, Jerry, et al.
Publicado: (2024)
From Chains to Graphs: Self-Structured Reasoning for General-Domain LLMs
por: Chen, Yingjian, et al.
Publicado: (2026)
por: Chen, Yingjian, et al.
Publicado: (2026)
Investigating Instruction Tuning Large Language Models on Graphs
por: Zhu, Kerui, et al.
Publicado: (2024)
por: Zhu, Kerui, et al.
Publicado: (2024)
mCoT: Multilingual Instruction Tuning for Reasoning Consistency in Language Models
por: Lai, Huiyuan, et al.
Publicado: (2024)
por: Lai, Huiyuan, et al.
Publicado: (2024)
JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models
por: Liu, Junyu, et al.
Publicado: (2026)
por: Liu, Junyu, et al.
Publicado: (2026)
Ejemplares similares
-
When Instructions Multiply: Measuring and Estimating LLM Capabilities of Multiple Instructions Following
por: Harada, Keno, et al.
Publicado: (2025) -
MKG-Rank: Enhancing Large Language Models with Knowledge Graph for Multilingual Medical Question Answering
por: Li, Feiyang, et al.
Publicado: (2025) -
On the Multilingual Ability of Decoder-based Pre-trained Language Models: Finding and Controlling Language-Specific Neurons
por: Kojima, Takeshi, et al.
Publicado: (2024) -
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
por: Gambardella, Andrew, et al.
Publicado: (2024) -
Image-Text Relation Prediction for Multilingual Tweets
por: Rikters, Matīss, et al.
Publicado: (2025)