Are LLM-generated plain language summaries truly understandable? A large-scale crowdsourced evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Yue, Sohn, Jae Ho, Leroy, Gondy, Cohen, Trevor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Retrieval augmentation of large language models for lay language generation
von: Guo, Yue, et al.
Veröffentlicht: (2022)
von: Guo, Yue, et al.
Veröffentlicht: (2022)
APPLS: Evaluating Evaluation Metrics for Plain Language Summarization
von: Guo, Yue, et al.
Veröffentlicht: (2023)
von: Guo, Yue, et al.
Veröffentlicht: (2023)
Role of Dependency Distance in Text Simplification: A Human vs ChatGPT Simplification Comparison
von: Lee, Sumi, et al.
Veröffentlicht: (2024)
von: Lee, Sumi, et al.
Veröffentlicht: (2024)
Into the crossfire: evaluating the use of a language model to crowdsource gun violence reports
von: Belisario, Adriano, et al.
Veröffentlicht: (2024)
von: Belisario, Adriano, et al.
Veröffentlicht: (2024)
DocLLM: A layout-aware generative language model for multimodal document understanding
von: Wang, Dongsheng, et al.
Veröffentlicht: (2023)
von: Wang, Dongsheng, et al.
Veröffentlicht: (2023)
Utilizing Large Language Models to Generate Synthetic Data to Increase the Performance of BERT-Based Neural Networks
von: Woolsey, Chancellor R., et al.
Veröffentlicht: (2024)
von: Woolsey, Chancellor R., et al.
Veröffentlicht: (2024)
AIDBench: A benchmark for evaluating the authorship identification capability of large language models
von: Wen, Zichen, et al.
Veröffentlicht: (2024)
von: Wen, Zichen, et al.
Veröffentlicht: (2024)
Automated Feedback Loops to Protect Text Simplification with Generative AI from Information Loss
von: Nandiraju, Abhay Kumara Sri Krishna, et al.
Veröffentlicht: (2025)
von: Nandiraju, Abhay Kumara Sri Krishna, et al.
Veröffentlicht: (2025)
GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models
von: Jassim, Serwan, et al.
Veröffentlicht: (2023)
von: Jassim, Serwan, et al.
Veröffentlicht: (2023)
From tools to thieves: Measuring and understanding public perceptions of AI through crowdsourced metaphors
von: Cheng, Myra, et al.
Veröffentlicht: (2025)
von: Cheng, Myra, et al.
Veröffentlicht: (2025)
Text and Audio Simplification: Human vs. ChatGPT
von: Leroy, Gondy, et al.
Veröffentlicht: (2024)
von: Leroy, Gondy, et al.
Veröffentlicht: (2024)
Effects of Added Emphasis and Pause in Audio Delivery of Health Information
von: Ahmed, Arif, et al.
Veröffentlicht: (2024)
von: Ahmed, Arif, et al.
Veröffentlicht: (2024)
A benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models
von: Li, Kunning, et al.
Veröffentlicht: (2025)
von: Li, Kunning, et al.
Veröffentlicht: (2025)
ECG-LLM -- training and evaluation of domain-specific large language models for electrocardiography
von: Ahrens, Lara, et al.
Veröffentlicht: (2025)
von: Ahrens, Lara, et al.
Veröffentlicht: (2025)
Zero-shot generation of synthetic neurosurgical data with large language models
von: Barr, Austin A., et al.
Veröffentlicht: (2025)
von: Barr, Austin A., et al.
Veröffentlicht: (2025)
Can large language models understand uncommon meanings of common words?
von: Wu, Jinyang, et al.
Veröffentlicht: (2024)
von: Wu, Jinyang, et al.
Veröffentlicht: (2024)
Re-evaluating Theory of Mind evaluation in large language models
von: Hu, Jennifer, et al.
Veröffentlicht: (2025)
von: Hu, Jennifer, et al.
Veröffentlicht: (2025)
The "LLM World of Words" English free association norms generated by large language models
von: Abramski, Katherine, et al.
Veröffentlicht: (2024)
von: Abramski, Katherine, et al.
Veröffentlicht: (2024)
Factual consistency evaluation of summarization in the Era of large language models
von: Luo, Zheheng, et al.
Veröffentlicht: (2024)
von: Luo, Zheheng, et al.
Veröffentlicht: (2024)
Multi-Agent LLM Judge: automatic personalized LLM judge design for evaluating natural language generation applications
von: Cao, Hongliu, et al.
Veröffentlicht: (2025)
von: Cao, Hongliu, et al.
Veröffentlicht: (2025)
Revisiting subword tokenization: A case study on affixal negation in large language models
von: Truong, Thinh Hung, et al.
Veröffentlicht: (2024)
von: Truong, Thinh Hung, et al.
Veröffentlicht: (2024)
AI-AI Bias: large language models favor communications generated by large language models
von: Laurito, Walter, et al.
Veröffentlicht: (2024)
von: Laurito, Walter, et al.
Veröffentlicht: (2024)
What does it mean to understand language?
von: Casto, Colton, et al.
Veröffentlicht: (2025)
von: Casto, Colton, et al.
Veröffentlicht: (2025)
Biomedical knowledge graph-optimized prompt generation for large language models
von: Soman, Karthik, et al.
Veröffentlicht: (2023)
von: Soman, Karthik, et al.
Veröffentlicht: (2023)
Integrating large language models and active inference to understand eye movements in reading and dyslexia
von: Donnarumma, Francesco, et al.
Veröffentlicht: (2023)
von: Donnarumma, Francesco, et al.
Veröffentlicht: (2023)
The creative psychometric item generator: a framework for item generation and validation using large language models
von: Laverghetta Jr., Antonio, et al.
Veröffentlicht: (2024)
von: Laverghetta Jr., Antonio, et al.
Veröffentlicht: (2024)
NLD-LLM: A systematic framework for evaluating small language transformer models on natural language description
von: Jelodar, Hamed, et al.
Veröffentlicht: (2025)
von: Jelodar, Hamed, et al.
Veröffentlicht: (2025)
CVE-LLM : Automatic vulnerability evaluation in medical device industry using large language models
von: Ghosh, Rikhiya, et al.
Veröffentlicht: (2024)
von: Ghosh, Rikhiya, et al.
Veröffentlicht: (2024)
Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation
von: Carandang, Kristine Ann M., et al.
Veröffentlicht: (2025)
von: Carandang, Kristine Ann M., et al.
Veröffentlicht: (2025)
rollama: An R package for using generative large language models through Ollama
von: Gruber, Johannes B., et al.
Veröffentlicht: (2024)
von: Gruber, Johannes B., et al.
Veröffentlicht: (2024)
Large-scale moral machine experiment on large language models
von: Ahmad, Muhammad Shahrul Zaim bin, et al.
Veröffentlicht: (2024)
von: Ahmad, Muhammad Shahrul Zaim bin, et al.
Veröffentlicht: (2024)
Automatic design optimization of preference-based subjective evaluation with online learning in crowdsourcing environment
von: Yasuda, Yusuke, et al.
Veröffentlicht: (2024)
von: Yasuda, Yusuke, et al.
Veröffentlicht: (2024)
Comparing large language models and human programmers for generating programming code
von: Hou, Wenpin, et al.
Veröffentlicht: (2024)
von: Hou, Wenpin, et al.
Veröffentlicht: (2024)
Generative adversarial networks vs large language models: a comparative study on synthetic tabular data generation
von: Barr, Austin A., et al.
Veröffentlicht: (2025)
von: Barr, Austin A., et al.
Veröffentlicht: (2025)
CMMLU: Measuring massive multitask language understanding in Chinese
von: Li, Haonan, et al.
Veröffentlicht: (2023)
von: Li, Haonan, et al.
Veröffentlicht: (2023)
Emergent effects of scaling on the functional hierarchies within large language models
von: Bogdan, Paul C.
Veröffentlicht: (2025)
von: Bogdan, Paul C.
Veröffentlicht: (2025)
\textsc{CantoNLU}: A benchmark for Cantonese natural language understanding
von: Min, Junghyun, et al.
Veröffentlicht: (2025)
von: Min, Junghyun, et al.
Veröffentlicht: (2025)
Can Generative AI Support Patients' & Caregivers' Informational Needs? Towards Task-Centric Evaluation Of AI Systems
von: Rajagopal, Shreya, et al.
Veröffentlicht: (2024)
von: Rajagopal, Shreya, et al.
Veröffentlicht: (2024)
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs?
von: Lei, Zhikai, et al.
Veröffentlicht: (2024)
von: Lei, Zhikai, et al.
Veröffentlicht: (2024)
Assessment and manipulation of latent constructs in pre-trained language models using psychometric scales
von: Reuben, Maor, et al.
Veröffentlicht: (2024)
von: Reuben, Maor, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Retrieval augmentation of large language models for lay language generation
von: Guo, Yue, et al.
Veröffentlicht: (2022) -
APPLS: Evaluating Evaluation Metrics for Plain Language Summarization
von: Guo, Yue, et al.
Veröffentlicht: (2023) -
Role of Dependency Distance in Text Simplification: A Human vs ChatGPT Simplification Comparison
von: Lee, Sumi, et al.
Veröffentlicht: (2024) -
Into the crossfire: evaluating the use of a language model to crowdsource gun violence reports
von: Belisario, Adriano, et al.
Veröffentlicht: (2024) -
DocLLM: A layout-aware generative language model for multimodal document understanding
von: Wang, Dongsheng, et al.
Veröffentlicht: (2023)