AllSummedUp: un framework open-source pour comparer les metriques d'evaluation de resume
Fuente:
arXiv
Guardado en:
| Autores principales: | Herserant, Tanguy, Guigue, Vincent |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SEval-Ex: A Statement-Level Framework for Explainable Summarization Evaluation
por: Herserant, Tanguy, et al.
Publicado: (2025)
por: Herserant, Tanguy, et al.
Publicado: (2025)
Can open source large language models be used for tumor documentation in Germany? -- An evaluation on urological doctors' notes
por: Lenz, Stefan, et al.
Publicado: (2025)
por: Lenz, Stefan, et al.
Publicado: (2025)
Large Language Models, scientific knowledge and factuality: A framework to streamline human expert evaluation
por: Wysocka, Magdalena, et al.
Publicado: (2023)
por: Wysocka, Magdalena, et al.
Publicado: (2023)
AspirinSum: an Aspect-based utility-preserved de-identification Summarization framework
por: Li, Ya-Lun
Publicado: (2024)
por: Li, Ya-Lun
Publicado: (2024)
adaptNMT: an open-source, language-agnostic development environment for Neural Machine Translation
por: Lankford, Séamus, et al.
Publicado: (2024)
por: Lankford, Séamus, et al.
Publicado: (2024)
Berta: an open-source, modular tool for AI-enabled clinical documentation
por: Vaid, Samridhi, et al.
Publicado: (2026)
por: Vaid, Samridhi, et al.
Publicado: (2026)
Closing the gap between open-source and commercial large language models for medical evidence summarization
por: Zhang, Gongbo, et al.
Publicado: (2024)
por: Zhang, Gongbo, et al.
Publicado: (2024)
Towards Lighter and Robust Evaluation for Retrieval Augmented Generation
por: Ispas, Alex-Razvan, et al.
Publicado: (2025)
por: Ispas, Alex-Razvan, et al.
Publicado: (2025)
$π$-yalli: un nouveau corpus pour le nahuatl
por: Torres-Moreno, Juan-Manuel, et al.
Publicado: (2024)
por: Torres-Moreno, Juan-Manuel, et al.
Publicado: (2024)
A word association network methodology for evaluating implicit biases in LLMs compared to humans
por: Abramski, Katherine, et al.
Publicado: (2025)
por: Abramski, Katherine, et al.
Publicado: (2025)
Diverse And Private Synthetic Datasets Generation for RAG evaluation: A multi-agent framework
por: Driouich, Ilias, et al.
Publicado: (2025)
por: Driouich, Ilias, et al.
Publicado: (2025)
COGNET-MD, an evaluation framework and dataset for Large Language Model benchmarks in the medical domain
por: Panagoulias, Dimitrios P., et al.
Publicado: (2024)
por: Panagoulias, Dimitrios P., et al.
Publicado: (2024)
How well do LLMs cite relevant medical references? An evaluation framework and analyses
por: Wu, Kevin, et al.
Publicado: (2024)
por: Wu, Kevin, et al.
Publicado: (2024)
Align-then-Slide: A complete evaluation framework for Ultra-Long Document-Level Machine Translation
por: Guo, Jiaxin, et al.
Publicado: (2025)
por: Guo, Jiaxin, et al.
Publicado: (2025)
A unified foundational framework for knowledge injection and evaluation of Large Language Models in Combustion Science
por: Yang, Zonglin, et al.
Publicado: (2026)
por: Yang, Zonglin, et al.
Publicado: (2026)
Memorization Dynamics of Fill-in-the-Middle Pretraining
por: von Arx, Tobias, et al.
Publicado: (2026)
por: von Arx, Tobias, et al.
Publicado: (2026)
CBEval: A framework for evaluating and interpreting cognitive biases in LLMs
por: Shaikh, Ammar, et al.
Publicado: (2024)
por: Shaikh, Ammar, et al.
Publicado: (2024)
DiscoSum: Discourse-aware News Summarization
por: Spangher, Alexander, et al.
Publicado: (2025)
por: Spangher, Alexander, et al.
Publicado: (2025)
MovieSum: An Abstractive Summarization Dataset for Movie Screenplays
por: Saxena, Rohit, et al.
Publicado: (2024)
por: Saxena, Rohit, et al.
Publicado: (2024)
Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Literal Extraction, Logical Inference, and Hallucination Risks in Long-Context LLMs
por: Ebrahimzadeh, Amirali, et al.
Publicado: (2026)
por: Ebrahimzadeh, Amirali, et al.
Publicado: (2026)
Comparative analysis of privacy-preserving open-source LLMs regarding extraction of diagnostic information from clinical CMR imaging reports
por: Amirrajab, Sina, et al.
Publicado: (2025)
por: Amirrajab, Sina, et al.
Publicado: (2025)
Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
por: Sun, Shengyin, et al.
Publicado: (2025)
por: Sun, Shengyin, et al.
Publicado: (2025)
ARE: Scaling Up Agent Environments and Evaluations
por: Froger, Romain, et al.
Publicado: (2025)
por: Froger, Romain, et al.
Publicado: (2025)
On Speeding Up Language Model Evaluation
por: Zhou, Jin Peng, et al.
Publicado: (2024)
por: Zhou, Jin Peng, et al.
Publicado: (2024)
ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition
por: Khan, Haidar, et al.
Publicado: (2025)
por: Khan, Haidar, et al.
Publicado: (2025)
NexusSum: Hierarchical LLM Agents for Long-Form Narrative Summarization
por: Kim, Hyuntak, et al.
Publicado: (2025)
por: Kim, Hyuntak, et al.
Publicado: (2025)
HeSum: a Novel Dataset for Abstractive Text Summarization in Hebrew
por: Paz-Argaman, Tzuf, et al.
Publicado: (2024)
por: Paz-Argaman, Tzuf, et al.
Publicado: (2024)
AgenticSum: An Agentic Inference-Time Framework for Faithful Clinical Text Summarization
por: Piya, Fahmida Liza, et al.
Publicado: (2026)
por: Piya, Fahmida Liza, et al.
Publicado: (2026)
MetaSumPerceiver: Multimodal Multi-Document Evidence Summarization for Fact-Checking
por: Chen, Ting-Chih, et al.
Publicado: (2024)
por: Chen, Ting-Chih, et al.
Publicado: (2024)
ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition
por: Alyahya, Hisham A., et al.
Publicado: (2025)
por: Alyahya, Hisham A., et al.
Publicado: (2025)
Reasoning Up the Instruction Ladder for Controllable Language Models
por: Zheng, Zishuo, et al.
Publicado: (2025)
por: Zheng, Zishuo, et al.
Publicado: (2025)
Re-evaluating Theory of Mind evaluation in large language models
por: Hu, Jennifer, et al.
Publicado: (2025)
por: Hu, Jennifer, et al.
Publicado: (2025)
UniSumEval: Towards Unified, Fine-Grained, Multi-Dimensional Summarization Evaluation for LLMs
por: Lee, Yuho, et al.
Publicado: (2024)
por: Lee, Yuho, et al.
Publicado: (2024)
More Than Sum of Its Parts: Deciphering Intent Shifts in Multimodal Hate Speech Detection
por: Sun, Runze, et al.
Publicado: (2026)
por: Sun, Runze, et al.
Publicado: (2026)
When LLMs Team Up: The Emergence of Collaborative Affective Computing
por: Lai, Wenna, et al.
Publicado: (2025)
por: Lai, Wenna, et al.
Publicado: (2025)
Contrast Is All You Need
por: Kilic, Burak, et al.
Publicado: (2023)
por: Kilic, Burak, et al.
Publicado: (2023)
dVoting: Fast Voting for dLLMs
por: Feng, Sicheng, et al.
Publicado: (2026)
por: Feng, Sicheng, et al.
Publicado: (2026)
Advancing Complex Medical Communication in Arabic with Sporo AraSum: Surpassing Existing Large Language Models
por: Lee, Chanseo, et al.
Publicado: (2024)
por: Lee, Chanseo, et al.
Publicado: (2024)
On the effectiveness of LLMs for automatic grading of open-ended questions in Spanish
por: Capdehourat, Germán, et al.
Publicado: (2025)
por: Capdehourat, Germán, et al.
Publicado: (2025)
Look It Up: Analysing Internal Web Search Capabilities of Modern LLMs
por: Kale, Sahil
Publicado: (2025)
por: Kale, Sahil
Publicado: (2025)
Ejemplares similares
-
SEval-Ex: A Statement-Level Framework for Explainable Summarization Evaluation
por: Herserant, Tanguy, et al.
Publicado: (2025) -
Can open source large language models be used for tumor documentation in Germany? -- An evaluation on urological doctors' notes
por: Lenz, Stefan, et al.
Publicado: (2025) -
Large Language Models, scientific knowledge and factuality: A framework to streamline human expert evaluation
por: Wysocka, Magdalena, et al.
Publicado: (2023) -
AspirinSum: an Aspect-based utility-preserved de-identification Summarization framework
por: Li, Ya-Lun
Publicado: (2024) -
adaptNMT: an open-source, language-agnostic development environment for Neural Machine Translation
por: Lankford, Séamus, et al.
Publicado: (2024)