Beyond Flesch-Kincaid: Prompt-based Metrics Improve Difficulty Classification of Educational Texts
Fuente:
arXiv
Guardado en:
| Autores principales: | Rooein, Donya, Rottger, Paul, Shaitarova, Anastassia, Hovy, Dirk |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Conversations as a Source for Teaching Scientific Concepts at Different Education Levels
por: Rooein, Donya, et al.
Publicado: (2024)
por: Rooein, Donya, et al.
Publicado: (2024)
Biased Tales: Cultural and Topic Bias in Generating Children's Stories
por: Rooein, Donya, et al.
Publicado: (2025)
por: Rooein, Donya, et al.
Publicado: (2025)
Exploring Subjective Tasks in Farsi: A Survey Analysis and Evaluation of Language Models
por: Rooein, Donya, et al.
Publicado: (2025)
por: Rooein, Donya, et al.
Publicado: (2025)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
por: Röttger, Paul, et al.
Publicado: (2024)
por: Röttger, Paul, et al.
Publicado: (2024)
Compromesso! Italian Many-Shot Jailbreaks Undermine the Safety of Large Language Models
por: Pernisi, Fabio, et al.
Publicado: (2024)
por: Pernisi, Fabio, et al.
Publicado: (2024)
Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions
por: Orlikowski, Matthias, et al.
Publicado: (2025)
por: Orlikowski, Matthias, et al.
Publicado: (2025)
Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance
por: de Araujo, Pedro Henrique Luz, et al.
Publicado: (2025)
por: de Araujo, Pedro Henrique Luz, et al.
Publicado: (2025)
The Ecological Fallacy in Annotation: Modelling Human Label Variation goes beyond Sociodemographics
por: Orlikowski, Matthias, et al.
Publicado: (2023)
por: Orlikowski, Matthias, et al.
Publicado: (2023)
PATS: Personality-Aware Teaching Strategies with Large Language Model Tutors
por: Rooein, Donya, et al.
Publicado: (2026)
por: Rooein, Donya, et al.
Publicado: (2026)
Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification
por: Xiong, Chenfei, et al.
Publicado: (2025)
por: Xiong, Chenfei, et al.
Publicado: (2025)
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
por: Russo, Giuseppe, et al.
Publicado: (2025)
por: Russo, Giuseppe, et al.
Publicado: (2025)
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
por: Röttger, Paul, et al.
Publicado: (2025)
por: Röttger, Paul, et al.
Publicado: (2025)
MSTS: A Multimodal Safety Test Suite for Vision-Language Models
por: Röttger, Paul, et al.
Publicado: (2025)
por: Röttger, Paul, et al.
Publicado: (2025)
Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
por: Ni, Jingwei, et al.
Publicado: (2025)
por: Ni, Jingwei, et al.
Publicado: (2025)
No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
por: Plaza-del-Arco, Flor Miriam, et al.
Publicado: (2025)
por: Plaza-del-Arco, Flor Miriam, et al.
Publicado: (2025)
Multilingual Performance Biases of Large Language Models in Education
por: Gupta, Vansh, et al.
Publicado: (2025)
por: Gupta, Vansh, et al.
Publicado: (2025)
Do Prompts Reshape Representations? An Empirical Study of Prompting Effects on Embeddings
por: Gonzalez-Gutierrez, Cesar, et al.
Publicado: (2025)
por: Gonzalez-Gutierrez, Cesar, et al.
Publicado: (2025)
"My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
por: Wang, Xinpeng, et al.
Publicado: (2024)
por: Wang, Xinpeng, et al.
Publicado: (2024)
Diffusion Language Models Are Natively Length-Aware
por: Rossi, Vittorio, et al.
Publicado: (2026)
por: Rossi, Vittorio, et al.
Publicado: (2026)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
por: Hu, Tiancheng, et al.
Publicado: (2025)
por: Hu, Tiancheng, et al.
Publicado: (2025)
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
por: Röttger, Paul, et al.
Publicado: (2023)
por: Röttger, Paul, et al.
Publicado: (2023)
Educators' Perceptions of Large Language Models as Tutors: Comparing Human and AI Tutors in a Blind Text-only Setting
por: Chowdhury, Sankalan Pal, et al.
Publicado: (2025)
por: Chowdhury, Sankalan Pal, et al.
Publicado: (2025)
Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
por: Baumann, Joachim, et al.
Publicado: (2025)
por: Baumann, Joachim, et al.
Publicado: (2025)
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
por: Röttger, Paul, et al.
Publicado: (2024)
por: Röttger, Paul, et al.
Publicado: (2024)
DADIT: A Dataset for Demographic Classification of Italian Twitter Users and a Comparison of Prediction Methods
por: Lupo, Lorenzo, et al.
Publicado: (2024)
por: Lupo, Lorenzo, et al.
Publicado: (2024)
Narratives at Conflict: Computational Analysis of News Framing in Multilingual Disinformation Campaigns
por: Sinelnik, Antonina, et al.
Publicado: (2024)
por: Sinelnik, Antonina, et al.
Publicado: (2024)
SpiritRAG: A Q&A System for Religion and Spirituality in the United Nations Archive
por: Gao, Yingqiang, et al.
Publicado: (2025)
por: Gao, Yingqiang, et al.
Publicado: (2025)
Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation
por: Zengaffinen, Yanick, et al.
Publicado: (2026)
por: Zengaffinen, Yanick, et al.
Publicado: (2026)
Improving Adversarial Data Collection by Supporting Annotators: Lessons from GAHD, a German Hate Speech Dataset
por: Goldzycher, Janis, et al.
Publicado: (2024)
por: Goldzycher, Janis, et al.
Publicado: (2024)
Teacher Demonstrations in a BabyLM's Zone of Proximal Development for Contingent Multi-Turn Interaction
por: Salhan, Suchir, et al.
Publicado: (2025)
por: Salhan, Suchir, et al.
Publicado: (2025)
Information-Theoretic Complementary Prompts for Improved Continual Text Classification
por: Zhang, Duzhen, et al.
Publicado: (2025)
por: Zhang, Duzhen, et al.
Publicado: (2025)
The AI Gap: How Socioeconomic Status Affects Language Technology Interactions
por: Bassignana, Elisa, et al.
Publicado: (2025)
por: Bassignana, Elisa, et al.
Publicado: (2025)
Navigating the Prompt Space: Improving LLM Classification of Social Science Texts Through Prompt Engineering
por: Gunes, Erkan, et al.
Publicado: (2026)
por: Gunes, Erkan, et al.
Publicado: (2026)
Towards Human-Level Text Coding with LLMs: The Case of Fatherhood Roles in Public Policy Documents
por: Lupo, Lorenzo, et al.
Publicado: (2023)
por: Lupo, Lorenzo, et al.
Publicado: (2023)
The Call for Socially Aware Language Technologies
por: Yang, Diyi, et al.
Publicado: (2024)
por: Yang, Diyi, et al.
Publicado: (2024)
Twists, Humps, and Pebbles: Multilingual Speech Recognition Models Exhibit Gender Performance Gaps
por: Attanasio, Giuseppe, et al.
Publicado: (2024)
por: Attanasio, Giuseppe, et al.
Publicado: (2024)
Impoverished Language Technology: The Lack of (Social) Class in NLP
por: Curry, Amanda Cercas, et al.
Publicado: (2024)
por: Curry, Amanda Cercas, et al.
Publicado: (2024)
Around the World in 24 Hours: Probing LLM Knowledge of Time and Place
por: Holtermann, Carolin, et al.
Publicado: (2025)
por: Holtermann, Carolin, et al.
Publicado: (2025)
Classist Tools: Social Class Correlates with Performance in NLP
por: Curry, Amanda Cercas, et al.
Publicado: (2024)
por: Curry, Amanda Cercas, et al.
Publicado: (2024)
Wisdom of Instruction-Tuned Language Model Crowds. Exploring Model Label Variation
por: Plaza-del-Arco, Flor Miriam, et al.
Publicado: (2023)
por: Plaza-del-Arco, Flor Miriam, et al.
Publicado: (2023)
Ejemplares similares
-
Conversations as a Source for Teaching Scientific Concepts at Different Education Levels
por: Rooein, Donya, et al.
Publicado: (2024) -
Biased Tales: Cultural and Topic Bias in Generating Children's Stories
por: Rooein, Donya, et al.
Publicado: (2025) -
Exploring Subjective Tasks in Farsi: A Survey Analysis and Evaluation of Language Models
por: Rooein, Donya, et al.
Publicado: (2025) -
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
por: Röttger, Paul, et al.
Publicado: (2024) -
Compromesso! Italian Many-Shot Jailbreaks Undermine the Safety of Large Language Models
por: Pernisi, Fabio, et al.
Publicado: (2024)