Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set
Fuente:
arXiv
Saved in:
| Main Authors: | Eichin, Florian, Liu, Yang Janet, Plank, Barbara, Hedderich, Michael A. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Copy First, Translate Later: Interpreting Translation Dynamics in Multilingual Pretraining
by: Körner, Felicia, et al.
Published: (2026)
by: Körner, Felicia, et al.
Published: (2026)
What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns
by: Hedderich, Michael A., et al.
Published: (2025)
by: Hedderich, Michael A., et al.
Published: (2025)
MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs
by: Zhao, Raoyuan, et al.
Published: (2025)
by: Zhao, Raoyuan, et al.
Published: (2025)
Semantic Component Analysis: Introducing Multi-Topic Distributions to Clustering-Based Topic Modeling
by: Eichin, Florian, et al.
Published: (2024)
by: Eichin, Florian, et al.
Published: (2024)
Understanding When Tree of Thoughts Succeeds: Larger Models Excel in Generation, Not Discrimination
by: Chen, Qiqi, et al.
Published: (2024)
by: Chen, Qiqi, et al.
Published: (2024)
ExPLAIND: Unifying Model, Data, and Training Attribution to Study Model Behavior
by: Eichin, Florian, et al.
Published: (2025)
by: Eichin, Florian, et al.
Published: (2025)
Threading the Needle: Reweaving Chain-of-Thought Reasoning to Explain Human Label Variation
by: Chen, Beiduo, et al.
Published: (2025)
by: Chen, Beiduo, et al.
Published: (2025)
Which course? Discourse! Teaching Discourse and Generation in the Era of LLMs
by: Li, Junyi Jessy, et al.
Published: (2026)
by: Li, Junyi Jessy, et al.
Published: (2026)
Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration
by: Shim, Ryan Soh-Eun, et al.
Published: (2026)
by: Shim, Ryan Soh-Eun, et al.
Published: (2026)
Crosslingual On-Policy Self-Distillation for Multilingual Reasoning
by: Liu, Yihong, et al.
Published: (2026)
by: Liu, Yihong, et al.
Published: (2026)
Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners
by: Liu, Yihong, et al.
Published: (2026)
by: Liu, Yihong, et al.
Published: (2026)
A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning: Performance, Consistency, and Faithfulness Across Languages
by: Zhao, Raoyuan, et al.
Published: (2025)
by: Zhao, Raoyuan, et al.
Published: (2025)
Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing
by: Zhao, Raoyuan, et al.
Published: (2025)
by: Zhao, Raoyuan, et al.
Published: (2025)
References Matter: Investigating the Impact of Reference Set Variation on Summarization Evaluation
by: Casola, Silvia, et al.
Published: (2025)
by: Casola, Silvia, et al.
Published: (2025)
Evaluating Robustness of Large Language Models Against Multilingual Typographical Errors
by: Zhao, Raoyuan, et al.
Published: (2025)
by: Zhao, Raoyuan, et al.
Published: (2025)
The Potential and Challenges of Evaluating Attitudes, Opinions, and Values in Large Language Models
by: Ma, Bolei, et al.
Published: (2024)
by: Ma, Bolei, et al.
Published: (2024)
Guardians of Discourse: Evaluating LLMs on Multilingual Offensive Language Detection
by: He, Jianfei, et al.
Published: (2024)
by: He, Jianfei, et al.
Published: (2024)
A Rose by Any Other Name: LLM-Generated Explanations Are Good Proxies for Human Explanations to Collect Label Distributions on NLI
by: Chen, Beiduo, et al.
Published: (2024)
by: Chen, Beiduo, et al.
Published: (2024)
SteerEval: Inference-time Interventions Strengthen Multilingual Generalization in Neural Summarization Metrics
by: Casola, Silvia, et al.
Published: (2026)
by: Casola, Silvia, et al.
Published: (2026)
Do LLMs Give Psychometrically Plausible Responses in Educational Assessments?
by: Säuberli, Andreas, et al.
Published: (2025)
by: Säuberli, Andreas, et al.
Published: (2025)
Interpreting Predictive Probabilities: Model Confidence or Human Label Variation?
by: Baan, Joris, et al.
Published: (2024)
by: Baan, Joris, et al.
Published: (2024)
RAcQUEt: Unveiling the Dangers of Overlooked Referential Ambiguity in Visual LLMs
by: Testoni, Alberto, et al.
Published: (2024)
by: Testoni, Alberto, et al.
Published: (2024)
From Noise to Signal to Selbstzweck: Reframing Human Label Variation in the Era of Post-training in NLP
by: Xu, Shanshan, et al.
Published: (2025)
by: Xu, Shanshan, et al.
Published: (2025)
Different Tastes of Entities: Investigating Human Label Variation in Named Entity Annotations
by: Peng, Siyao, et al.
Published: (2024)
by: Peng, Siyao, et al.
Published: (2024)
Beyond Input Understanding: Diagnosing Multilingual Mathematical Reasoning with Directed Acyclic Trace Graphs
by: Zhang, Jiaqiao, et al.
Published: (2026)
by: Zhang, Jiaqiao, et al.
Published: (2026)
ReverseMath: Answer Inversion for Scalable and Verifiable Mathematical Problem Generation
by: Zhao, Raoyuan, et al.
Published: (2026)
by: Zhao, Raoyuan, et al.
Published: (2026)
When Meanings Meet: Investigating the Emergence and Quality of Shared Concept Spaces during Multilingual Language Model Training
by: Körner, Felicia, et al.
Published: (2026)
by: Körner, Felicia, et al.
Published: (2026)
Tracing Multilingual Factual Knowledge Acquisition in Pretraining
by: Liu, Yihong, et al.
Published: (2025)
by: Liu, Yihong, et al.
Published: (2025)
EVADE: LLM-Based Explanation Generation and Validation for Error Detection in NLI
by: Zuo, Longfei, et al.
Published: (2025)
by: Zuo, Longfei, et al.
Published: (2025)
Probing the Feasibility of Multilingual Speaker Anonymization
by: Meyer, Sarina, et al.
Published: (2024)
by: Meyer, Sarina, et al.
Published: (2024)
VariErr NLI: Separating Annotation Error from Human Label Variation
by: Weber-Genzel, Leon, et al.
Published: (2024)
by: Weber-Genzel, Leon, et al.
Published: (2024)
Unlocking Latent Discourse Translation in LLMs Through Quality-Aware Decoding
by: Mohammed, Wafaa, et al.
Published: (2025)
by: Mohammed, Wafaa, et al.
Published: (2025)
Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
Crossing Domains without Labels: Distant Supervision for Term Extraction
by: Senger, Elena, et al.
Published: (2025)
by: Senger, Elena, et al.
Published: (2025)
I Came, I Saw, I Explained: Benchmarking Multimodal LLMs on Figurative Meaning in Memes
by: Zhou, Shijia, et al.
Published: (2026)
by: Zhou, Shijia, et al.
Published: (2026)
Decoupling the Effect of Chain-of-Thought Reasoning: A Human Label Variation Perspective
by: Chen, Beiduo, et al.
Published: (2026)
by: Chen, Beiduo, et al.
Published: (2026)
Turning English-centric LLMs Into Polyglots: How Much Multilinguality Is Needed?
by: Kew, Tannon, et al.
Published: (2023)
by: Kew, Tannon, et al.
Published: (2023)
Probing Gender Bias in Multilingual LLMs: A Case Study of Stereotypes in Persian
by: Kalhor, Ghazal, et al.
Published: (2025)
by: Kalhor, Ghazal, et al.
Published: (2025)
LiTEx: A Linguistic Taxonomy of Explanations for Understanding Within-Label Variation in Natural Language Inference
by: Hong, Pingjun, et al.
Published: (2025)
by: Hong, Pingjun, et al.
Published: (2025)
Through the Lens of Split Vote: Exploring Disagreement, Difficulty and Calibration in Legal Case Outcome Classification
by: Xu, Shanshan, et al.
Published: (2024)
by: Xu, Shanshan, et al.
Published: (2024)
Similar Items
-
Copy First, Translate Later: Interpreting Translation Dynamics in Multilingual Pretraining
by: Körner, Felicia, et al.
Published: (2026) -
What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns
by: Hedderich, Michael A., et al.
Published: (2025) -
MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs
by: Zhao, Raoyuan, et al.
Published: (2025) -
Semantic Component Analysis: Introducing Multi-Topic Distributions to Clustering-Based Topic Modeling
by: Eichin, Florian, et al.
Published: (2024) -
Understanding When Tree of Thoughts Succeeds: Larger Models Excel in Generation, Not Discrimination
by: Chen, Qiqi, et al.
Published: (2024)