How Much Noise Can BERT Handle? Insights from Multilingual Sentence Difficulty Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Khallaf, Nouran, Sharoff, Serge |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification
von: Hilasaca, Kenji, et al.
Veröffentlicht: (2026)
von: Hilasaca, Kenji, et al.
Veröffentlicht: (2026)
To Predict or Not to Predict? Towards reliable uncertainty estimation in the presence of noise
von: Khallaf, Nouran, et al.
Veröffentlicht: (2026)
von: Khallaf, Nouran, et al.
Veröffentlicht: (2026)
Reading Between the Lines: A dataset and a study on why some texts are tougher than others
von: Khallaf, Nouran, et al.
Veröffentlicht: (2025)
von: Khallaf, Nouran, et al.
Veröffentlicht: (2025)
Can LLM Reasoning Be Trusted? A Comparative Study: Using Human Benchmarking on Statistical Tasks
von: Nagarkar, Crish, et al.
Veröffentlicht: (2026)
von: Nagarkar, Crish, et al.
Veröffentlicht: (2026)
A Multilingual Human Annotated Corpus of Original and Easy-to-Read Texts to Support Access to Democratic Participatory Processes
von: Bott, Stefan, et al.
Veröffentlicht: (2026)
von: Bott, Stefan, et al.
Veröffentlicht: (2026)
Controlling Out-of-Domain Gaps in LLMs for Genre Classification and Generated Text Detection
von: Roussinov, Dmitri, et al.
Veröffentlicht: (2024)
von: Roussinov, Dmitri, et al.
Veröffentlicht: (2024)
Almost Clinical: Linguistic properties of synthetic electronic health records
von: Sharoff, Serge, et al.
Veröffentlicht: (2026)
von: Sharoff, Serge, et al.
Veröffentlicht: (2026)
Multilingual Sentence-T5: Scalable Sentence Encoders for Multilingual Applications
von: Yano, Chihiro, et al.
Veröffentlicht: (2024)
von: Yano, Chihiro, et al.
Veröffentlicht: (2024)
Turning English-centric LLMs Into Polyglots: How Much Multilinguality Is Needed?
von: Kew, Tannon, et al.
Veröffentlicht: (2023)
von: Kew, Tannon, et al.
Veröffentlicht: (2023)
How Much Can RAG Help the Reasoning of LLM?
von: Liu, Jingyu, et al.
Veröffentlicht: (2024)
von: Liu, Jingyu, et al.
Veröffentlicht: (2024)
NusaBERT: Teaching IndoBERT to be Multilingual and Multicultural
von: Wongso, Wilson, et al.
Veröffentlicht: (2024)
von: Wongso, Wilson, et al.
Veröffentlicht: (2024)
How do Large Language Models Handle Multilingualism?
von: Zhao, Yiran, et al.
Veröffentlicht: (2024)
von: Zhao, Yiran, et al.
Veröffentlicht: (2024)
UoL-UPF at TSAR 2025 Shared Task A Generate-and-Select Approach for Readability-Controlled Text Simplification.
von: Hayakawa, Akio, et al.
Veröffentlicht: (2025)
von: Hayakawa, Akio, et al.
Veröffentlicht: (2025)
How does a Multilingual LM Handle Multiple Languages?
von: Kakarla, Santhosh, et al.
Veröffentlicht: (2025)
von: Kakarla, Santhosh, et al.
Veröffentlicht: (2025)
How Effectively Can BERT Models Interpret Context and Detect Bengali Communal Violent Text?
von: Khondoker, Abdullah, et al.
Veröffentlicht: (2025)
von: Khondoker, Abdullah, et al.
Veröffentlicht: (2025)
Datasets for Multilingual Answer Sentence Selection
von: Gabburo, Matteo, et al.
Veröffentlicht: (2024)
von: Gabburo, Matteo, et al.
Veröffentlicht: (2024)
Detecting Redundant Health Survey Questions Using Language-agnostic BERT Sentence Embedding (LaBSE)
von: Kang, Sunghoon, et al.
Veröffentlicht: (2024)
von: Kang, Sunghoon, et al.
Veröffentlicht: (2024)
Fine-tuning the SwissBERT Encoder Model for Embedding Sentences and Documents
von: Grosjean, Juri, et al.
Veröffentlicht: (2024)
von: Grosjean, Juri, et al.
Veröffentlicht: (2024)
How Much Can We Forget about Data Contamination?
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
Sentiment Informed Sentence BERT-Ensemble Algorithm for Depression Detection
von: Ogunleye, Bayode, et al.
Veröffentlicht: (2024)
von: Ogunleye, Bayode, et al.
Veröffentlicht: (2024)
Dependency Annotation of Ottoman Turkish with Multilingual BERT
von: Özateş, Şaziye Betül, et al.
Veröffentlicht: (2024)
von: Özateş, Şaziye Betül, et al.
Veröffentlicht: (2024)
SwissBERT: The Multilingual Language Model for Switzerland
von: Vamvas, Jannis, et al.
Veröffentlicht: (2023)
von: Vamvas, Jannis, et al.
Veröffentlicht: (2023)
How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM Hallucination
von: Islam, Saad Obaid ul, et al.
Veröffentlicht: (2025)
von: Islam, Saad Obaid ul, et al.
Veröffentlicht: (2025)
Topic mining based on fine-tuning Sentence-BERT and LDA
von: Li, Jianheng, et al.
Veröffentlicht: (2025)
von: Li, Jianheng, et al.
Veröffentlicht: (2025)
How Much of Your Data Can Suck? Thresholds for Domain Performance and Emergent Misalignment in LLMs
von: Ouyang, Jian, et al.
Veröffentlicht: (2025)
von: Ouyang, Jian, et al.
Veröffentlicht: (2025)
SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in HuBERT
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023)
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023)
Benchmarking LLM Guardrails in Handling Multilingual Toxicity
von: Yang, Yahan, et al.
Veröffentlicht: (2024)
von: Yang, Yahan, et al.
Veröffentlicht: (2024)
Adaptative Bilingual Aligning Using Multilingual Sentence Embedding
von: Kraif, Olivier
Veröffentlicht: (2024)
von: Kraif, Olivier
Veröffentlicht: (2024)
Comparing Human and Language Models Sentence Processing Difficulties on Complex Structures
von: Amouyal, Samuel Joseph, et al.
Veröffentlicht: (2025)
von: Amouyal, Samuel Joseph, et al.
Veröffentlicht: (2025)
How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM?
von: Pletenev, Sergey, et al.
Veröffentlicht: (2025)
von: Pletenev, Sergey, et al.
Veröffentlicht: (2025)
CoT-BERT: Enhancing Unsupervised Sentence Representation through Chain-of-Thought
von: Zhang, Bowen, et al.
Veröffentlicht: (2023)
von: Zhang, Bowen, et al.
Veröffentlicht: (2023)
Improving Sampling Methods for Fine-tuning SentenceBERT in Text Streams
von: Garcia, Cristiano Mesquita, et al.
Veröffentlicht: (2024)
von: Garcia, Cristiano Mesquita, et al.
Veröffentlicht: (2024)
Towards Building Efficient Sentence BERT Models using Layer Pruning
von: Shelke, Anushka, et al.
Veröffentlicht: (2024)
von: Shelke, Anushka, et al.
Veröffentlicht: (2024)
Predicting Antibiotic Resistance Patterns Using Sentence-BERT: A Machine Learning Approach
von: Alwakeel, Mahmoud, et al.
Veröffentlicht: (2025)
von: Alwakeel, Mahmoud, et al.
Veröffentlicht: (2025)
ColBERT: Using BERT Sentence Embedding in Parallel Neural Networks for Computational Humor
von: Annamoradnejad, Issa, et al.
Veröffentlicht: (2020)
von: Annamoradnejad, Issa, et al.
Veröffentlicht: (2020)
EMS: Efficient and Effective Massively Multilingual Sentence Embedding Learning
von: Mao, Zhuoyuan, et al.
Veröffentlicht: (2022)
von: Mao, Zhuoyuan, et al.
Veröffentlicht: (2022)
Multilingual JobBERT for Cross-Lingual Job Title Matching
von: Decorte, Jens-Joris, et al.
Veröffentlicht: (2025)
von: Decorte, Jens-Joris, et al.
Veröffentlicht: (2025)
A Comprehensive Survey of Sentence Representations: From the BERT Epoch to the ChatGPT Era and Beyond
von: Kashyap, Abhinav Ramesh, et al.
Veröffentlicht: (2023)
von: Kashyap, Abhinav Ramesh, et al.
Veröffentlicht: (2023)
Benchmarking BERT-based Models for Sentence-level Topic Classification in Nepali Language
von: Karki, Nischal, et al.
Veröffentlicht: (2026)
von: Karki, Nischal, et al.
Veröffentlicht: (2026)
Bot Meets Shortcut: How Can LLMs Aid in Handling Unknown Invariance OOD Scenarios?
von: Zheng, Shiyan, et al.
Veröffentlicht: (2025)
von: Zheng, Shiyan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification
von: Hilasaca, Kenji, et al.
Veröffentlicht: (2026) -
To Predict or Not to Predict? Towards reliable uncertainty estimation in the presence of noise
von: Khallaf, Nouran, et al.
Veröffentlicht: (2026) -
Reading Between the Lines: A dataset and a study on why some texts are tougher than others
von: Khallaf, Nouran, et al.
Veröffentlicht: (2025) -
Can LLM Reasoning Be Trusted? A Comparative Study: Using Human Benchmarking on Statistical Tasks
von: Nagarkar, Crish, et al.
Veröffentlicht: (2026) -
A Multilingual Human Annotated Corpus of Original and Easy-to-Read Texts to Support Access to Democratic Participatory Processes
von: Bott, Stefan, et al.
Veröffentlicht: (2026)