Dynamic benchmarking framework for LLM-based conversational data capture
Fuente:
arXiv
Salvato in:
| Autori principali: | Aluffi, Pietro Alessandro, Zietkiewicz, Patrick, Bazzi, Marya, Arderne, Matt, Murevics, Vladimirs |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Categorising SME Bank Transactions with Machine Learning and Synthetic Data Generation
di: Alessandro, Aluffi Pietro, et al.
Pubblicazione: (2025)
di: Alessandro, Aluffi Pietro, et al.
Pubblicazione: (2025)
Tag and correct: high precision post-editing approach to correction of speech recognition errors
di: Ziętkiewicz, Tomasz
Pubblicazione: (2024)
di: Ziętkiewicz, Tomasz
Pubblicazione: (2024)
Enhancing Debunking Effectiveness through LLM-based Personality Adaptation
di: Dell'Oglio, Pietro, et al.
Pubblicazione: (2026)
di: Dell'Oglio, Pietro, et al.
Pubblicazione: (2026)
Efficient RL for optimizing conversation level outcomes with an LLM-based tutor
di: Nam, Hyunji, et al.
Pubblicazione: (2025)
di: Nam, Hyunji, et al.
Pubblicazione: (2025)
COGNET-MD, an evaluation framework and dataset for Large Language Model benchmarks in the medical domain
di: Panagoulias, Dimitrios P., et al.
Pubblicazione: (2024)
di: Panagoulias, Dimitrios P., et al.
Pubblicazione: (2024)
LLMzSzŁ: a comprehensive LLM benchmark for Polish
di: Jassem, Krzysztof, et al.
Pubblicazione: (2025)
di: Jassem, Krzysztof, et al.
Pubblicazione: (2025)
Exploring the generalization of LLM truth directions on conversational formats
di: Ichmoukhamedov, Timour, et al.
Pubblicazione: (2025)
di: Ichmoukhamedov, Timour, et al.
Pubblicazione: (2025)
Developing an AI framework to automatically detect shared decision-making in patient-doctor conversations
di: Ponce-Ponte, Oscar J., et al.
Pubblicazione: (2025)
di: Ponce-Ponte, Oscar J., et al.
Pubblicazione: (2025)
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
di: Pacchiardi, Lorenzo, et al.
Pubblicazione: (2024)
di: Pacchiardi, Lorenzo, et al.
Pubblicazione: (2024)
TelcoLM: collecting data, adapting, and benchmarking language models for the telecommunication domain
di: Barboule, Camille, et al.
Pubblicazione: (2024)
di: Barboule, Camille, et al.
Pubblicazione: (2024)
MinorBench: A hand-built benchmark for content-based risks for children
di: Khoo, Shaun, et al.
Pubblicazione: (2025)
di: Khoo, Shaun, et al.
Pubblicazione: (2025)
Suvach -- Generated Hindi QA benchmark
di: Narayanan, Vaishak, et al.
Pubblicazione: (2024)
di: Narayanan, Vaishak, et al.
Pubblicazione: (2024)
FLUID-LLM: Learning Computational Fluid Dynamics with Spatiotemporal-aware Large Language Models
di: Zhu, Max, et al.
Pubblicazione: (2024)
di: Zhu, Max, et al.
Pubblicazione: (2024)
IoT-LLM: a framework for enhancing Large Language Model reasoning from real-world sensor data
di: An, Tuo, et al.
Pubblicazione: (2024)
di: An, Tuo, et al.
Pubblicazione: (2024)
Mapping the Course for Prompt-based Structured Prediction
di: Pauk, Matt, et al.
Pubblicazione: (2025)
di: Pauk, Matt, et al.
Pubblicazione: (2025)
Language Ranker: A Lightweight Ranking framework for LLM Decoding
di: Zhang, Chenheng, et al.
Pubblicazione: (2025)
di: Zhang, Chenheng, et al.
Pubblicazione: (2025)
An LLM-enabled semantic-centric framework to consume privacy policies
di: Zhao, Rui, et al.
Pubblicazione: (2025)
di: Zhao, Rui, et al.
Pubblicazione: (2025)
WHODUNIT: Evaluation benchmark for culprit detection in mystery stories
di: Gupta, Kshitij
Pubblicazione: (2025)
di: Gupta, Kshitij
Pubblicazione: (2025)
Collaborative Storytelling and LLM: A Linguistic Analysis of Automatically-Generated Role-Playing Game Sessions
di: Maisto, Alessandro
Pubblicazione: (2025)
di: Maisto, Alessandro
Pubblicazione: (2025)
Automatic benchmarking of large multimodal models via iterative experiment programming
di: Conti, Alessandro, et al.
Pubblicazione: (2024)
di: Conti, Alessandro, et al.
Pubblicazione: (2024)
DynamicNER: A Dynamic, Multilingual, and Fine-Grained Dataset for LLM-based Named Entity Recognition
di: Luo, Hanjun, et al.
Pubblicazione: (2024)
di: Luo, Hanjun, et al.
Pubblicazione: (2024)
Where is the Mind? Persona Vectors and LLM Individuation
di: Beckmann, Pierre, et al.
Pubblicazione: (2026)
di: Beckmann, Pierre, et al.
Pubblicazione: (2026)
PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice
di: Liu, Shuyu, et al.
Pubblicazione: (2025)
di: Liu, Shuyu, et al.
Pubblicazione: (2025)
The Impact of Persona-based Political Perspectives on Hateful Content Detection
di: Civelli, Stefano, et al.
Pubblicazione: (2025)
di: Civelli, Stefano, et al.
Pubblicazione: (2025)
Batayan: A Filipino NLP benchmark for evaluating Large Language Models
di: Montalan, Jann Railey, et al.
Pubblicazione: (2025)
di: Montalan, Jann Railey, et al.
Pubblicazione: (2025)
LongTail-Swap: benchmarking language models' abilities on rare words
di: Algayres, Robin, et al.
Pubblicazione: (2025)
di: Algayres, Robin, et al.
Pubblicazione: (2025)
Polish-English medical knowledge transfer: A new benchmark and results
di: Grzybowski, Łukasz, et al.
Pubblicazione: (2024)
di: Grzybowski, Łukasz, et al.
Pubblicazione: (2024)
Halluverse-M^3: A multitask multilingual benchmark for hallucination in LLMs
di: Abdaljalil, Samir, et al.
Pubblicazione: (2026)
di: Abdaljalil, Samir, et al.
Pubblicazione: (2026)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
di: Wang, Chonghua, et al.
Pubblicazione: (2024)
di: Wang, Chonghua, et al.
Pubblicazione: (2024)
TAPS: Task Aware Proposal Distributions for Speculative Sampling
di: Zbib, Mohamad, et al.
Pubblicazione: (2026)
di: Zbib, Mohamad, et al.
Pubblicazione: (2026)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
di: Kim, Eunsu, et al.
Pubblicazione: (2024)
di: Kim, Eunsu, et al.
Pubblicazione: (2024)
The Russian-focused embedders' exploration: ruMTEB benchmark and Russian embedding model design
di: Snegirev, Artem, et al.
Pubblicazione: (2024)
di: Snegirev, Artem, et al.
Pubblicazione: (2024)
A benchmark for joint dialogue satisfaction, emotion recognition, and emotion state transition prediction
di: Bian, Jing, et al.
Pubblicazione: (2026)
di: Bian, Jing, et al.
Pubblicazione: (2026)
A thorough benchmark of automatic text classification: From traditional approaches to large language models
di: Cunha, Washington, et al.
Pubblicazione: (2025)
di: Cunha, Washington, et al.
Pubblicazione: (2025)
QuanTemp: A real-world open-domain benchmark for fact-checking numerical claims
di: V, Venktesh, et al.
Pubblicazione: (2024)
di: V, Venktesh, et al.
Pubblicazione: (2024)
TransitGPT: A Generative AI-based framework for interacting with GTFS data using Large Language Models
di: Devunuri, Saipraneeth, et al.
Pubblicazione: (2024)
di: Devunuri, Saipraneeth, et al.
Pubblicazione: (2024)
AspirinSum: an Aspect-based utility-preserved de-identification Summarization framework
di: Li, Ya-Lun
Pubblicazione: (2024)
di: Li, Ya-Lun
Pubblicazione: (2024)
Learning LLM Preference over Intra-Dialogue Pairs: A Framework for Utterance-level Understandings
di: Liu, Xuanqing, et al.
Pubblicazione: (2025)
di: Liu, Xuanqing, et al.
Pubblicazione: (2025)
Do language models capture implied discourse meanings? An investigation with exhaustivity implicatures of Korean morphology
di: Shin, Hagyeong, et al.
Pubblicazione: (2024)
di: Shin, Hagyeong, et al.
Pubblicazione: (2024)
Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation
di: Rädsch, Tim, et al.
Pubblicazione: (2025)
di: Rädsch, Tim, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Categorising SME Bank Transactions with Machine Learning and Synthetic Data Generation
di: Alessandro, Aluffi Pietro, et al.
Pubblicazione: (2025) -
Tag and correct: high precision post-editing approach to correction of speech recognition errors
di: Ziętkiewicz, Tomasz
Pubblicazione: (2024) -
Enhancing Debunking Effectiveness through LLM-based Personality Adaptation
di: Dell'Oglio, Pietro, et al.
Pubblicazione: (2026) -
Efficient RL for optimizing conversation level outcomes with an LLM-based tutor
di: Nam, Hyunji, et al.
Pubblicazione: (2025) -
COGNET-MD, an evaluation framework and dataset for Large Language Model benchmarks in the medical domain
di: Panagoulias, Dimitrios P., et al.
Pubblicazione: (2024)