Not too long do read: Evaluating LLM-generated extreme scientific summaries
Fuente:
arXiv
Guardado en:
| Autores principales: | Lyu, Zhuoqi, Ke, Qing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
COMO: Closed-Loop Optical Molecule Recognition with Minimum Risk Training
por: Lyu, Zhuoqi, et al.
Publicado: (2026)
por: Lyu, Zhuoqi, et al.
Publicado: (2026)
Unused information in token probability distribution of generative LLM: improving LLM reading comprehension through calculation of expected values
por: Zawistowski, Krystian
Publicado: (2024)
por: Zawistowski, Krystian
Publicado: (2024)
Towards understanding evolution of science through language model series
por: Dong, Junjie, et al.
Publicado: (2024)
por: Dong, Junjie, et al.
Publicado: (2024)
Reveal and Release: Iterative LLM Unlearning with Self-generated Data
por: Xie, Linxi, et al.
Publicado: (2025)
por: Xie, Linxi, et al.
Publicado: (2025)
Towards an automatic method for generating topical vocabulary test forms for specific reading passages
por: Flor, Michael, et al.
Publicado: (2025)
por: Flor, Michael, et al.
Publicado: (2025)
Talking to Machines: do you read me?
por: Rojas-Barahona, Lina M.
Publicado: (2024)
por: Rojas-Barahona, Lina M.
Publicado: (2024)
LLM as Attention-Informed NTM and Topic Modeling as long-input Generation: Interpretability and long-Context Capability
por: Xu, Xuan, et al.
Publicado: (2025)
por: Xu, Xuan, et al.
Publicado: (2025)
Benchmark of stylistic variation in LLM-generated texts
por: Milička, Jiří, et al.
Publicado: (2025)
por: Milička, Jiří, et al.
Publicado: (2025)
Gender Bias in LLM-generated Interview Responses
por: Kong, Haein, et al.
Publicado: (2024)
por: Kong, Haein, et al.
Publicado: (2024)
Learning to Summarize from LLM-generated Feedback
por: Song, Hwanjun, et al.
Publicado: (2024)
por: Song, Hwanjun, et al.
Publicado: (2024)
Semantic uncertainty in advanced decoding methods for LLM generation
por: Foodeei, Darius, et al.
Publicado: (2025)
por: Foodeei, Darius, et al.
Publicado: (2025)
Temperature-scaling surprisal estimates improve fit to human reading times -- but does it do so for the "right reasons"?
por: Liu, Tong, et al.
Publicado: (2023)
por: Liu, Tong, et al.
Publicado: (2023)
AI-University: An LLM-based platform for instructional alignment to scientific classrooms
por: Shojaei, Mostafa Faghih, et al.
Publicado: (2025)
por: Shojaei, Mostafa Faghih, et al.
Publicado: (2025)
Modeling cognitive processes of natural reading with transformer-based Language Models
por: Bianchi, Bruno, et al.
Publicado: (2025)
por: Bianchi, Bruno, et al.
Publicado: (2025)
Multi-Agent LLM Judge: automatic personalized LLM judge design for evaluating natural language generation applications
por: Cao, Hongliu, et al.
Publicado: (2025)
por: Cao, Hongliu, et al.
Publicado: (2025)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
por: Wang, Chonghua, et al.
Publicado: (2024)
por: Wang, Chonghua, et al.
Publicado: (2024)
CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
por: Liu, Jiayu, et al.
Publicado: (2025)
por: Liu, Jiayu, et al.
Publicado: (2025)
Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM
por: Yuan, Dong, et al.
Publicado: (2024)
por: Yuan, Dong, et al.
Publicado: (2024)
CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation
por: Ke, Pei, et al.
Publicado: (2023)
por: Ke, Pei, et al.
Publicado: (2023)
CLAWS:Creativity detection for LLM-generated solutions using Attention Window of Sections
por: Kim, Keuntae, et al.
Publicado: (2025)
por: Kim, Keuntae, et al.
Publicado: (2025)
LLM-as-a-qualitative-judge: automating error analysis in natural language generation
por: Chirkova, Nadezhda, et al.
Publicado: (2025)
por: Chirkova, Nadezhda, et al.
Publicado: (2025)
Beyond speculation: Measuring the growing presence of LLM-generated texts in multilingual disinformation
por: Macko, Dominik, et al.
Publicado: (2025)
por: Macko, Dominik, et al.
Publicado: (2025)
Data Compressibility Quantifies LLM Memorization
por: Huang, Yizhan, et al.
Publicado: (2025)
por: Huang, Yizhan, et al.
Publicado: (2025)
Evaluating Metrics for Safety with LLM-as-Judges
por: Clegg, Kester, et al.
Publicado: (2025)
por: Clegg, Kester, et al.
Publicado: (2025)
CSCE: Boosting LLM Reasoning by Simultaneous Enhancing of Causal Significance and Consistency
por: Wang, Kangsheng, et al.
Publicado: (2024)
por: Wang, Kangsheng, et al.
Publicado: (2024)
LLM Prompt Evaluation for Educational Applications
por: Holmes, Langdon, et al.
Publicado: (2026)
por: Holmes, Langdon, et al.
Publicado: (2026)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
por: Kim, Eunsu, et al.
Publicado: (2024)
por: Kim, Eunsu, et al.
Publicado: (2024)
Defining and Evaluating Decision and Composite Risk in Language Models Applied to Natural Language Inference
por: Shen, Ke, et al.
Publicado: (2024)
por: Shen, Ke, et al.
Publicado: (2024)
Cross-cultural Inspiration Detection and Analysis in Real and LLM-generated Social Media Data
por: Ignat, Oana, et al.
Publicado: (2024)
por: Ignat, Oana, et al.
Publicado: (2024)
The "LLM World of Words" English free association norms generated by large language models
por: Abramski, Katherine, et al.
Publicado: (2024)
por: Abramski, Katherine, et al.
Publicado: (2024)
Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
por: Mushtaq, Abdullah, et al.
Publicado: (2025)
por: Mushtaq, Abdullah, et al.
Publicado: (2025)
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models
por: Lyu, Xinxi, et al.
Publicado: (2024)
por: Lyu, Xinxi, et al.
Publicado: (2024)
The Challenges of Evaluating LLM Applications: An Analysis of Automated, Human, and LLM-Based Approaches
por: Abeysinghe, Bhashithe, et al.
Publicado: (2024)
por: Abeysinghe, Bhashithe, et al.
Publicado: (2024)
Read Before You Think: Mitigating LLM Comprehension Failures with Step-by-Step Reading
por: Han, Feijiang, et al.
Publicado: (2025)
por: Han, Feijiang, et al.
Publicado: (2025)
The why, what, and how of AI-based coding in scientific research
por: Zhuang, Tonghe, et al.
Publicado: (2024)
por: Zhuang, Tonghe, et al.
Publicado: (2024)
Rethinking Human Preference Evaluation of LLM Rationales
por: Li, Ziang, et al.
Publicado: (2025)
por: Li, Ziang, et al.
Publicado: (2025)
Integrated Framework for LLM Evaluation with Answer Generation
por: Lee, Sujeong, et al.
Publicado: (2025)
por: Lee, Sujeong, et al.
Publicado: (2025)
Evaluating the Diversity and Quality of LLM Generated Content
por: Shypula, Alexander, et al.
Publicado: (2025)
por: Shypula, Alexander, et al.
Publicado: (2025)
Autorubric: Unifying Rubric-based LLM Evaluation
por: Rao, Delip, et al.
Publicado: (2026)
por: Rao, Delip, et al.
Publicado: (2026)
MERA: A Comprehensive LLM Evaluation in Russian
por: Fenogenova, Alena, et al.
Publicado: (2024)
por: Fenogenova, Alena, et al.
Publicado: (2024)
Ejemplares similares
-
COMO: Closed-Loop Optical Molecule Recognition with Minimum Risk Training
por: Lyu, Zhuoqi, et al.
Publicado: (2026) -
Unused information in token probability distribution of generative LLM: improving LLM reading comprehension through calculation of expected values
por: Zawistowski, Krystian
Publicado: (2024) -
Towards understanding evolution of science through language model series
por: Dong, Junjie, et al.
Publicado: (2024) -
Reveal and Release: Iterative LLM Unlearning with Self-generated Data
por: Xie, Linxi, et al.
Publicado: (2025) -
Towards an automatic method for generating topical vocabulary test forms for specific reading passages
por: Flor, Michael, et al.
Publicado: (2025)