Multilingual Embedding Probes Fail to Generalize Across Learner Corpora
Fuente:
arXiv
Saved in:
| Main Authors: | Lyngbaek, Laurits, Kristensen-McLachlan, Ross Deans |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Alignment Drift in CEFR-prompted LLMs for Interactive Spanish Tutoring
by: Almasi, Mina, et al.
Published: (2025)
by: Almasi, Mina, et al.
Published: (2025)
I only read it for the plot! Maturity Ratings Affect Fanfiction Style and Community Engagement
by: Jacobsen, Mia, et al.
Published: (2025)
by: Jacobsen, Mia, et al.
Published: (2025)
Science is Exploration: Computational Frontiers for Conceptual Metaphor Theory
by: Hicke, Rebecca M. M., et al.
Published: (2024)
by: Hicke, Rebecca M. M., et al.
Published: (2024)
Context is Key(NMF): Modelling Topical Information Dynamics in Chinese Diaspora Media
by: Kristensen-McLachlan, Ross Deans, et al.
Published: (2024)
by: Kristensen-McLachlan, Ross Deans, et al.
Published: (2024)
Are Chatbots Reliable Text Annotators? Sometimes
by: Kristensen-McLachlan, Ross Deans, et al.
Published: (2023)
by: Kristensen-McLachlan, Ross Deans, et al.
Published: (2023)
Says Who? Effective Zero-Shot Annotation of Focalization
by: Hicke, Rebecca M. M., et al.
Published: (2024)
by: Hicke, Rebecca M. M., et al.
Published: (2024)
Attention Flows: Tracing LLM Conceptual Engagement via Story Summaries
by: Hicke, Rebecca M. M., et al.
Published: (2026)
by: Hicke, Rebecca M. M., et al.
Published: (2026)
Continuous sentiment scores for literary and multilingual contexts
by: Lyngbaek, Laurits, et al.
Published: (2025)
by: Lyngbaek, Laurits, et al.
Published: (2025)
Is Sentiment Banana-Shaped? Exploring the Geometry and Portability of Sentiment Concept Vectors
by: Lyngbaek, Laurits, et al.
Published: (2026)
by: Lyngbaek, Laurits, et al.
Published: (2026)
Active Learning for Multilingual Fingerspelling Corpora
by: Wang, Shuai, et al.
Published: (2023)
by: Wang, Shuai, et al.
Published: (2023)
Multilingual and Explainable Text Detoxification with Parallel Corpora
by: Dementieva, Daryna, et al.
Published: (2024)
by: Dementieva, Daryna, et al.
Published: (2024)
AcrosticSleuth: Probabilistic Identification and Ranking of Acrostics in Multilingual Corpora
by: Fedchin, Aleksandr, et al.
Published: (2024)
by: Fedchin, Aleksandr, et al.
Published: (2024)
A Recipe of Parallel Corpora Exploitation for Multilingual Large Language Models
by: Lin, Peiqin, et al.
Published: (2024)
by: Lin, Peiqin, et al.
Published: (2024)
CAMEO: Collection of Multilingual Emotional Speech Corpora
by: Christop, Iwona, et al.
Published: (2025)
by: Christop, Iwona, et al.
Published: (2025)
Building Corpora for Single-Channel Speech Separation Across Multiple Domains
by: Maciejewski, Matthew, et al.
Published: (2018)
by: Maciejewski, Matthew, et al.
Published: (2018)
A Survey on Multilingual Large Language Models: Corpora, Alignment, and Bias
by: Xu, Yuemei, et al.
Published: (2024)
by: Xu, Yuemei, et al.
Published: (2024)
False Sense of Security: Why Probing-based Malicious Input Detection Fails to Generalize
by: Wang, Cheng, et al.
Published: (2025)
by: Wang, Cheng, et al.
Published: (2025)
From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corpora
by: Shen, Yingli, et al.
Published: (2025)
by: Shen, Yingli, et al.
Published: (2025)
Characterizing the Effects of Translation on Intertextuality using Multilingual Embedding Spaces
by: McGovern, Hope, et al.
Published: (2025)
by: McGovern, Hope, et al.
Published: (2025)
GhanaNLP Parallel Corpora: Comprehensive Multilingual Resources for Low-Resource Ghanaian Languages
by: Gyamfi, Lawrence Adu, et al.
Published: (2026)
by: Gyamfi, Lawrence Adu, et al.
Published: (2026)
Entity Insertion in Multilingual Linked Corpora: The Case of Wikipedia
by: Feith, Tomás, et al.
Published: (2024)
by: Feith, Tomás, et al.
Published: (2024)
CodePivot: Bootstrapping Multilingual Transpilation in LLMs via Reinforcement Learning without Parallel Corpora
by: Li, Shangyu, et al.
Published: (2026)
by: Li, Shangyu, et al.
Published: (2026)
AI Brown and AI Koditex: LLM-Generated Corpora Comparable to Traditional Corpora of English and Czech Texts
by: Milička, Jiří, et al.
Published: (2025)
by: Milička, Jiří, et al.
Published: (2025)
Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set
by: Eichin, Florian, et al.
Published: (2025)
by: Eichin, Florian, et al.
Published: (2025)
Attributing Culture-Conditioned Generations to Pretraining Corpora
by: Li, Huihan, et al.
Published: (2024)
by: Li, Huihan, et al.
Published: (2024)
Do Language Models Care About Text Quality? Evaluating Web-Crawled Corpora Across 11 Languages
by: van Noord, Rik, et al.
Published: (2024)
by: van Noord, Rik, et al.
Published: (2024)
A Multilingual Perspective on Probing Gender Bias
by: Stańczak, Karolina
Published: (2024)
by: Stańczak, Karolina
Published: (2024)
Which Feedback Works for Whom? Differential Effects of LLM-Generated Feedback Elements Across Learner Profiles
by: Furuhashi, Momoka, et al.
Published: (2026)
by: Furuhashi, Momoka, et al.
Published: (2026)
Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic Evaluations
by: Agarwal, Ananth, et al.
Published: (2025)
by: Agarwal, Ananth, et al.
Published: (2025)
Adapting Multilingual Embedding Models to Historical Luxembourgish
by: Michail, Andrianos, et al.
Published: (2025)
by: Michail, Andrianos, et al.
Published: (2025)
Leveraging LLM For Synchronizing Information Across Multilingual Tables
by: Khincha, Siddharth, et al.
Published: (2025)
by: Khincha, Siddharth, et al.
Published: (2025)
Semiparametric Latent Topic Modeling on Consumer-Generated Corpora
by: Dayta, Dominic B., et al.
Published: (2021)
by: Dayta, Dominic B., et al.
Published: (2021)
Examining Multilingual Embedding Models Cross-Lingually Through LLM-Generated Adversarial Examples
by: Michail, Andrianos, et al.
Published: (2025)
by: Michail, Andrianos, et al.
Published: (2025)
Measuring the Effect of Disfluency in Multilingual Knowledge Probing Benchmarks
by: Semenov, Kirill, et al.
Published: (2025)
by: Semenov, Kirill, et al.
Published: (2025)
Getting More from Less: Large Language Models are Good Spontaneous Multilingual Learners
by: Zhang, Shimao, et al.
Published: (2024)
by: Zhang, Shimao, et al.
Published: (2024)
SSLfmm: An R Package for Semi-Supervised Learning with a Mixed-Missingness Mechanism in Finite Mixture Models
by: McLachlan, Geoffrey J., et al.
Published: (2025)
by: McLachlan, Geoffrey J., et al.
Published: (2025)
Data Caricatures: On the Representation of African American Language in Pretraining Corpora
by: Deas, Nicholas, et al.
Published: (2025)
by: Deas, Nicholas, et al.
Published: (2025)
Findings of the Second BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
by: Hu, Michael Y., et al.
Published: (2024)
by: Hu, Michael Y., et al.
Published: (2024)
Validating and Exploring Large Geographic Corpora
by: Dunn, Jonathan
Published: (2024)
by: Dunn, Jonathan
Published: (2024)
Identifying Emerging Concepts in Large Corpora
by: Ma, Sibo, et al.
Published: (2025)
by: Ma, Sibo, et al.
Published: (2025)
Similar Items
-
Alignment Drift in CEFR-prompted LLMs for Interactive Spanish Tutoring
by: Almasi, Mina, et al.
Published: (2025) -
I only read it for the plot! Maturity Ratings Affect Fanfiction Style and Community Engagement
by: Jacobsen, Mia, et al.
Published: (2025) -
Science is Exploration: Computational Frontiers for Conceptual Metaphor Theory
by: Hicke, Rebecca M. M., et al.
Published: (2024) -
Context is Key(NMF): Modelling Topical Information Dynamics in Chinese Diaspora Media
by: Kristensen-McLachlan, Ross Deans, et al.
Published: (2024) -
Are Chatbots Reliable Text Annotators? Sometimes
by: Kristensen-McLachlan, Ross Deans, et al.
Published: (2023)