Reading Between the Lines: A dataset and a study on why some texts are tougher than others
Fuente:
arXiv
Salvato in:
| Autori principali: | Khallaf, Nouran, Eugeni, Carlo, Sharoff, Serge |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
To Predict or Not to Predict? Towards reliable uncertainty estimation in the presence of noise
di: Khallaf, Nouran, et al.
Pubblicazione: (2026)
di: Khallaf, Nouran, et al.
Pubblicazione: (2026)
How Much Noise Can BERT Handle? Insights from Multilingual Sentence Difficulty Detection
di: Khallaf, Nouran, et al.
Pubblicazione: (2026)
di: Khallaf, Nouran, et al.
Pubblicazione: (2026)
Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification
di: Hilasaca, Kenji, et al.
Pubblicazione: (2026)
di: Hilasaca, Kenji, et al.
Pubblicazione: (2026)
A Multilingual Human Annotated Corpus of Original and Easy-to-Read Texts to Support Access to Democratic Participatory Processes
di: Bott, Stefan, et al.
Pubblicazione: (2026)
di: Bott, Stefan, et al.
Pubblicazione: (2026)
Can LLM Reasoning Be Trusted? A Comparative Study: Using Human Benchmarking on Statistical Tasks
di: Nagarkar, Crish, et al.
Pubblicazione: (2026)
di: Nagarkar, Crish, et al.
Pubblicazione: (2026)
Controlling Out-of-Domain Gaps in LLMs for Genre Classification and Generated Text Detection
di: Roussinov, Dmitri, et al.
Pubblicazione: (2024)
di: Roussinov, Dmitri, et al.
Pubblicazione: (2024)
Almost Clinical: Linguistic properties of synthetic electronic health records
di: Sharoff, Serge, et al.
Pubblicazione: (2026)
di: Sharoff, Serge, et al.
Pubblicazione: (2026)
UoL-UPF at TSAR 2025 Shared Task A Generate-and-Select Approach for Readability-Controlled Text Simplification.
di: Hayakawa, Akio, et al.
Pubblicazione: (2025)
di: Hayakawa, Akio, et al.
Pubblicazione: (2025)
Are some books better than others?
di: Rosenbusch, Hannes, et al.
Pubblicazione: (2025)
di: Rosenbusch, Hannes, et al.
Pubblicazione: (2025)
Reading Between the Lines: Classifying Resume Seniority with Large Language Models
di: Cohen, Matan, et al.
Pubblicazione: (2025)
di: Cohen, Matan, et al.
Pubblicazione: (2025)
Read Between the Lines: A Benchmark for Uncovering Political Bias in Bangla News Articles
di: Lia, Nusrat Jahan, et al.
Pubblicazione: (2025)
di: Lia, Nusrat Jahan, et al.
Pubblicazione: (2025)
Reading Between the Lines: The One-Sided Conversation Problem
di: Ebert, Victoria, et al.
Pubblicazione: (2025)
di: Ebert, Victoria, et al.
Pubblicazione: (2025)
Reading Between the Lines: How Electronic Nonverbal Cues shape Emotion Decoding
di: Kumar, Taara, et al.
Pubblicazione: (2026)
di: Kumar, Taara, et al.
Pubblicazione: (2026)
"I Wrote, I Paused, I Rewrote" Teaching LLMs to Read Between the Lines of Student Writing
di: Zafar, Samra, et al.
Pubblicazione: (2025)
di: Zafar, Samra, et al.
Pubblicazione: (2025)
Reading the unreadable: Creating a dataset of 19th century English newspapers using image-to-text language models
di: Bourne, Jonathan
Pubblicazione: (2025)
di: Bourne, Jonathan
Pubblicazione: (2025)
A multi-level multi-label text classification dataset of 19th century Ottoman and Russian literary and critical texts
di: Gokceoglu, Gokcen, et al.
Pubblicazione: (2024)
di: Gokceoglu, Gokcen, et al.
Pubblicazione: (2024)
ProText: A benchmark dataset for measuring (mis)gendering in long-form texts
di: Kotek, Hadas, et al.
Pubblicazione: (2026)
di: Kotek, Hadas, et al.
Pubblicazione: (2026)
Reading Between the Lines: Combining Pause Dynamics and Semantic Coherence for Automated Assessment of Thought Disorder
di: Chen, Feng, et al.
Pubblicazione: (2025)
di: Chen, Feng, et al.
Pubblicazione: (2025)
A large-scale image-text dataset benchmark for farmland segmentation
di: Tao, Chao, et al.
Pubblicazione: (2025)
di: Tao, Chao, et al.
Pubblicazione: (2025)
Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?
di: Nunez, Jeanmely Rojas, et al.
Pubblicazione: (2026)
di: Nunez, Jeanmely Rojas, et al.
Pubblicazione: (2026)
BeanCounter: A low-toxicity, large-scale, and open dataset of business-oriented text
di: Wang, Siyan, et al.
Pubblicazione: (2024)
di: Wang, Siyan, et al.
Pubblicazione: (2024)
GPT or BERT: why not both?
di: Charpentier, Lucas Georges Gabriel, et al.
Pubblicazione: (2024)
di: Charpentier, Lucas Georges Gabriel, et al.
Pubblicazione: (2024)
Tgea: An error-annotated dataset and benchmark tasks for text generation from pretrained language models
di: He, Jie, et al.
Pubblicazione: (2025)
di: He, Jie, et al.
Pubblicazione: (2025)
Can AI Read Between The Lines? Benchmarking LLMs On Financial Nuance
di: Kubica, Dominick, et al.
Pubblicazione: (2025)
di: Kubica, Dominick, et al.
Pubblicazione: (2025)
ReadCtrl: Personalizing text generation with readability-controlled instruction learning
di: Tran, Hieu, et al.
Pubblicazione: (2024)
di: Tran, Hieu, et al.
Pubblicazione: (2024)
LLMs can hide text in other text of the same length
di: Norelli, Antonio, et al.
Pubblicazione: (2025)
di: Norelli, Antonio, et al.
Pubblicazione: (2025)
Entailed Between the Lines: Incorporating Implication into NLI
di: Havaldar, Shreya, et al.
Pubblicazione: (2025)
di: Havaldar, Shreya, et al.
Pubblicazione: (2025)
The Thin Line Between Comprehension and Persuasion in LLMs
di: de Wynter, Adrian, et al.
Pubblicazione: (2025)
di: de Wynter, Adrian, et al.
Pubblicazione: (2025)
Less than one percent of words would be affected by gender-inclusive language in German press texts
di: Müller-Spitzer, Carolin, et al.
Pubblicazione: (2024)
di: Müller-Spitzer, Carolin, et al.
Pubblicazione: (2024)
TAGLAS: An atlas of text-attributed graph datasets in the era of large graph and language models
di: Feng, Jiarui, et al.
Pubblicazione: (2024)
di: Feng, Jiarui, et al.
Pubblicazione: (2024)
Reading Between the Lines: Towards Reliable Black-box LLM Fingerprinting via Zeroth-order Gradient Estimation
di: Shao, Shuo, et al.
Pubblicazione: (2025)
di: Shao, Shuo, et al.
Pubblicazione: (2025)
MedReadCtrl: Personalizing medical text generation with readability-controlled instruction learning
di: Tran, Hieu, et al.
Pubblicazione: (2025)
di: Tran, Hieu, et al.
Pubblicazione: (2025)
MaterioMiner -- An ontology-based text mining dataset for extraction of process-structure-property entities
di: Durmaz, Ali Riza, et al.
Pubblicazione: (2024)
di: Durmaz, Ali Riza, et al.
Pubblicazione: (2024)
AlleNoise: large-scale text classification benchmark dataset with real-world label noise
di: Rączkowska, Alicja, et al.
Pubblicazione: (2024)
di: Rączkowska, Alicja, et al.
Pubblicazione: (2024)
Synthetically generated text for supervised text analysis
di: Halterman, Andrew
Pubblicazione: (2023)
di: Halterman, Andrew
Pubblicazione: (2023)
The why, what, and how of AI-based coding in scientific research
di: Zhuang, Tonghe, et al.
Pubblicazione: (2024)
di: Zhuang, Tonghe, et al.
Pubblicazione: (2024)
Taec: a Manually annotated text dataset for trait and phenotype extraction and entity linking in wheat breeding literature
di: Nédellec, Claire, et al.
Pubblicazione: (2024)
di: Nédellec, Claire, et al.
Pubblicazione: (2024)
Designing large language model prompts to extract scores from messy text: A shared dataset and challenge
di: Thelwall, Mike
Pubblicazione: (2026)
di: Thelwall, Mike
Pubblicazione: (2026)
Spider4SSC & S2CLite: A text-to-multi-query-language dataset using lightweight ontology-agnostic SPARQL to Cypher parser
di: Vejvar, Martin, et al.
Pubblicazione: (2025)
di: Vejvar, Martin, et al.
Pubblicazione: (2025)
Reading Between the Prompts: How Stereotypes Shape LLM's Implicit Personalization
di: Neplenbroek, Vera, et al.
Pubblicazione: (2025)
di: Neplenbroek, Vera, et al.
Pubblicazione: (2025)
Documenti analoghi
-
To Predict or Not to Predict? Towards reliable uncertainty estimation in the presence of noise
di: Khallaf, Nouran, et al.
Pubblicazione: (2026) -
How Much Noise Can BERT Handle? Insights from Multilingual Sentence Difficulty Detection
di: Khallaf, Nouran, et al.
Pubblicazione: (2026) -
Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification
di: Hilasaca, Kenji, et al.
Pubblicazione: (2026) -
A Multilingual Human Annotated Corpus of Original and Easy-to-Read Texts to Support Access to Democratic Participatory Processes
di: Bott, Stefan, et al.
Pubblicazione: (2026) -
Can LLM Reasoning Be Trusted? A Comparative Study: Using Human Benchmarking on Statistical Tasks
di: Nagarkar, Crish, et al.
Pubblicazione: (2026)