Report Cards: Qualitative Evaluation of Language Models Using Natural Language Summaries
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Blair, Cui, Fuyang, Paster, Keiran, Ba, Jimmy, Vaezipoor, Pashootan, Pitis, Silviu, Zhang, Michael R. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving Context-Aware Preference Modeling for Language Models
von: Pitis, Silviu, et al.
Veröffentlicht: (2024)
von: Pitis, Silviu, et al.
Veröffentlicht: (2024)
LLMs and the Abstraction and Reasoning Corpus: Successes, Failures, and the Importance of Object-based Representations
von: Xu, Yudong, et al.
Veröffentlicht: (2023)
von: Xu, Yudong, et al.
Veröffentlicht: (2023)
STEVE-1: A Generative Model for Text-to-Behavior in Minecraft
von: Lifshitz, Shalev, et al.
Veröffentlicht: (2023)
von: Lifshitz, Shalev, et al.
Veröffentlicht: (2023)
Simulating Viva Voce Examinations to Evaluate Clinical Reasoning in Large Language Models
von: Chiu, Christopher, et al.
Veröffentlicht: (2025)
von: Chiu, Christopher, et al.
Veröffentlicht: (2025)
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
von: Ruan, Yangjun, et al.
Veröffentlicht: (2023)
von: Ruan, Yangjun, et al.
Veröffentlicht: (2023)
Llemma: An Open Language Model For Mathematics
von: Azerbayev, Zhangir, et al.
Veröffentlicht: (2023)
von: Azerbayev, Zhangir, et al.
Veröffentlicht: (2023)
Multilingual Natural Language Processing Model for Radiology Reports -- The Summary is all you need!
von: Lindo, Mariana, et al.
Veröffentlicht: (2023)
von: Lindo, Mariana, et al.
Veröffentlicht: (2023)
Adaptive Elicitation of Latent Information Using Natural Language
von: Wang, Jimmy, et al.
Veröffentlicht: (2025)
von: Wang, Jimmy, et al.
Veröffentlicht: (2025)
LLM-Supported Natural Language to Bash Translation
von: Westenfelder, Finnian, et al.
Veröffentlicht: (2025)
von: Westenfelder, Finnian, et al.
Veröffentlicht: (2025)
Evaluating Text Summaries Generated by Large Language Models Using OpenAI's GPT
von: Shakil, Hassan, et al.
Veröffentlicht: (2024)
von: Shakil, Hassan, et al.
Veröffentlicht: (2024)
The Phenomenology of Hallucinations
von: Ruscio, Valeria, et al.
Veröffentlicht: (2026)
von: Ruscio, Valeria, et al.
Veröffentlicht: (2026)
Reward Machines for Deep RL in Noisy and Uncertain Environments
von: Li, Andrew C., et al.
Veröffentlicht: (2024)
von: Li, Andrew C., et al.
Veröffentlicht: (2024)
LegalLens Shared Task 2024: Legal Violation Identification in Unstructured Text
von: Hagag, Ben, et al.
Veröffentlicht: (2024)
von: Hagag, Ben, et al.
Veröffentlicht: (2024)
Using Large Language Models for Hyperparameter Optimization
von: Zhang, Michael R., et al.
Veröffentlicht: (2023)
von: Zhang, Michael R., et al.
Veröffentlicht: (2023)
Cards Against LLMs: Benchmarking Humor Alignment in Large Language Models
von: Fettach, Yousra, et al.
Veröffentlicht: (2026)
von: Fettach, Yousra, et al.
Veröffentlicht: (2026)
Querying Structured Data Through Natural Language Using Language Models
von: Valentin-Micu, Hontan, et al.
Veröffentlicht: (2026)
von: Valentin-Micu, Hontan, et al.
Veröffentlicht: (2026)
Natural Language Counterfactual Explanations for Graphs Using Large Language Models
von: Giorgi, Flavio, et al.
Veröffentlicht: (2024)
von: Giorgi, Flavio, et al.
Veröffentlicht: (2024)
EvalCards: A Framework for Standardized Evaluation Reporting
von: Dhar, Ruchira, et al.
Veröffentlicht: (2025)
von: Dhar, Ruchira, et al.
Veröffentlicht: (2025)
Quantized Large Language Models in Biomedical Natural Language Processing: Evaluation and Recommendation
von: Zhan, Zaifu, et al.
Veröffentlicht: (2025)
von: Zhan, Zaifu, et al.
Veröffentlicht: (2025)
A Comprehensive Evaluation of Large Language Models on Benchmark Biomedical Text Processing Tasks
von: Jahan, Israt, et al.
Veröffentlicht: (2023)
von: Jahan, Israt, et al.
Veröffentlicht: (2023)
Defining and Evaluating Decision and Composite Risk in Language Models Applied to Natural Language Inference
von: Shen, Ke, et al.
Veröffentlicht: (2024)
von: Shen, Ke, et al.
Veröffentlicht: (2024)
Natural Language Satisfiability: Exploring the Problem Distribution and Evaluating Transformer-based Language Models
von: Madusanka, Tharindu, et al.
Veröffentlicht: (2025)
von: Madusanka, Tharindu, et al.
Veröffentlicht: (2025)
Evaluating Large Language Models on Financial Report Summarization: An Empirical Study
von: Yang, Xinqi, et al.
Veröffentlicht: (2024)
von: Yang, Xinqi, et al.
Veröffentlicht: (2024)
WRDScore: New Metric for Evaluation of Natural Language Generation Models
von: Mussabayev, Ravil
Veröffentlicht: (2024)
von: Mussabayev, Ravil
Veröffentlicht: (2024)
A Survey on Retrieval-Augmented Text Generation for Large Language Models
von: Huang, Yizheng, et al.
Veröffentlicht: (2024)
von: Huang, Yizheng, et al.
Veröffentlicht: (2024)
ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models
von: Baek, Jinheon, et al.
Veröffentlicht: (2024)
von: Baek, Jinheon, et al.
Veröffentlicht: (2024)
LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models
von: Elesedy, Hayder, et al.
Veröffentlicht: (2024)
von: Elesedy, Hayder, et al.
Veröffentlicht: (2024)
DataAgent: Evaluating Large Language Models' Ability to Answer Zero-Shot, Natural Language Queries
von: Mishra, Manit, et al.
Veröffentlicht: (2024)
von: Mishra, Manit, et al.
Veröffentlicht: (2024)
Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
von: Wu, Yulong, et al.
Veröffentlicht: (2025)
von: Wu, Yulong, et al.
Veröffentlicht: (2025)
Where Pretraining writes and Alignment reads: the asymmetry of Transformer weight space
von: Ruscio, Valeria, et al.
Veröffentlicht: (2026)
von: Ruscio, Valeria, et al.
Veröffentlicht: (2026)
Evaluating Neural Language Models as Cognitive Models of Language Acquisition
von: Martínez, Héctor Javier Vázquez, et al.
Veröffentlicht: (2023)
von: Martínez, Héctor Javier Vázquez, et al.
Veröffentlicht: (2023)
Consistency Evaluation of News Article Summaries Generated by Large (and Small) Language Models
von: Gilhuly, Colleen, et al.
Veröffentlicht: (2025)
von: Gilhuly, Colleen, et al.
Veröffentlicht: (2025)
BioInstruct: Instruction Tuning of Large Language Models for Biomedical Natural Language Processing
von: Tran, Hieu, et al.
Veröffentlicht: (2023)
von: Tran, Hieu, et al.
Veröffentlicht: (2023)
FEABench: Evaluating Language Models on Multiphysics Reasoning Ability
von: Mudur, Nayantara, et al.
Veröffentlicht: (2025)
von: Mudur, Nayantara, et al.
Veröffentlicht: (2025)
A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents
von: Mo, Lingbo, et al.
Veröffentlicht: (2024)
von: Mo, Lingbo, et al.
Veröffentlicht: (2024)
GeneSUM: Large Language Model-based Gene Summary Extraction
von: Chen, Zhijian, et al.
Veröffentlicht: (2024)
von: Chen, Zhijian, et al.
Veröffentlicht: (2024)
On the Importance and Evaluation of Narrativity in Natural Language AI Explanations
von: Cedro, Mateusz, et al.
Veröffentlicht: (2026)
von: Cedro, Mateusz, et al.
Veröffentlicht: (2026)
Language Bottleneck Models for Qualitative Knowledge State Modeling
von: Berthon, Antonin, et al.
Veröffentlicht: (2025)
von: Berthon, Antonin, et al.
Veröffentlicht: (2025)
Can Language Models Solve Graph Problems in Natural Language?
von: Wang, Heng, et al.
Veröffentlicht: (2023)
von: Wang, Heng, et al.
Veröffentlicht: (2023)
Knowledge-Augmented Large Language Models for Personalized Contextual Query Suggestion
von: Baek, Jinheon, et al.
Veröffentlicht: (2023)
von: Baek, Jinheon, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Improving Context-Aware Preference Modeling for Language Models
von: Pitis, Silviu, et al.
Veröffentlicht: (2024) -
LLMs and the Abstraction and Reasoning Corpus: Successes, Failures, and the Importance of Object-based Representations
von: Xu, Yudong, et al.
Veröffentlicht: (2023) -
STEVE-1: A Generative Model for Text-to-Behavior in Minecraft
von: Lifshitz, Shalev, et al.
Veröffentlicht: (2023) -
Simulating Viva Voce Examinations to Evaluate Clinical Reasoning in Large Language Models
von: Chiu, Christopher, et al.
Veröffentlicht: (2025) -
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
von: Ruan, Yangjun, et al.
Veröffentlicht: (2023)