Advancing Academic Chatbots: Evaluation of Non Traditional Outputs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Favero, Nicole, Salute, Francesca, Hardt, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
KatzBot: Revolutionizing Academic Chatbot for Enhanced Communication
von: Kumar, Sahil, et al.
Veröffentlicht: (2024)
von: Kumar, Sahil, et al.
Veröffentlicht: (2024)
Evaluating Commercial AI Chatbots as News Intermediaries
von: Suzgun, Mirac, et al.
Veröffentlicht: (2026)
von: Suzgun, Mirac, et al.
Veröffentlicht: (2026)
ChatbotManip: A Dataset to Facilitate Evaluation and Oversight of Manipulative Chatbot Behaviour
von: Contro, Jack, et al.
Veröffentlicht: (2025)
von: Contro, Jack, et al.
Veröffentlicht: (2025)
Evaluating language models as risk scores
von: Cruz, André F., et al.
Veröffentlicht: (2024)
von: Cruz, André F., et al.
Veröffentlicht: (2024)
LMStyle Benchmark: Evaluating Text Style Transfer for Chatbots
von: Chen, Jianlin
Veröffentlicht: (2024)
von: Chen, Jianlin
Veröffentlicht: (2024)
Test-Time Training on Nearest Neighbors for Large Language Models
von: Hardt, Moritz, et al.
Veröffentlicht: (2023)
von: Hardt, Moritz, et al.
Veröffentlicht: (2023)
Advancing Risk and Quality Assurance: A RAG Chatbot for Improved Regulatory Compliance
von: Hillebrand, Lars, et al.
Veröffentlicht: (2025)
von: Hillebrand, Lars, et al.
Veröffentlicht: (2025)
End-to-End Chatbot Evaluation with Adaptive Reasoning and Uncertainty Filtering
von: Dang, Nhi, et al.
Veröffentlicht: (2026)
von: Dang, Nhi, et al.
Veröffentlicht: (2026)
Training on the Test Task Confounds Evaluation and Emergence
von: Dominguez-Olmedo, Ricardo, et al.
Veröffentlicht: (2024)
von: Dominguez-Olmedo, Ricardo, et al.
Veröffentlicht: (2024)
TherapyGym: Evaluating and Aligning Clinical Fidelity and Safety in Therapy Chatbots
von: Huang, Fangrui, et al.
Veröffentlicht: (2026)
von: Huang, Fangrui, et al.
Veröffentlicht: (2026)
LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs
von: Long, Do Xuan, et al.
Veröffentlicht: (2024)
von: Long, Do Xuan, et al.
Veröffentlicht: (2024)
SESGO: Spanish Evaluation of Stereotypical Generative Outputs
von: Robles, Melissa, et al.
Veröffentlicht: (2025)
von: Robles, Melissa, et al.
Veröffentlicht: (2025)
Advancing Fairness in Natural Language Processing: From Traditional Methods to Explainability
von: Jourdan, Fanny
Veröffentlicht: (2024)
von: Jourdan, Fanny
Veröffentlicht: (2024)
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
von: Chiang, Wei-Lin, et al.
Veröffentlicht: (2024)
von: Chiang, Wei-Lin, et al.
Veröffentlicht: (2024)
Questioning the Survey Responses of Large Language Models
von: Dominguez-Olmedo, Ricardo, et al.
Veröffentlicht: (2023)
von: Dominguez-Olmedo, Ricardo, et al.
Veröffentlicht: (2023)
An Improved Traditional Chinese Evaluation Suite for Foundation Model
von: Tam, Zhi-Rui, et al.
Veröffentlicht: (2024)
von: Tam, Zhi-Rui, et al.
Veröffentlicht: (2024)
ARAGOG: Advanced RAG Output Grading
von: Eibich, Matouš, et al.
Veröffentlicht: (2024)
von: Eibich, Matouš, et al.
Veröffentlicht: (2024)
LLM as a Scorer: The Impact of Output Order on Dialogue Evaluation
von: Chen, Yi-Pei, et al.
Veröffentlicht: (2024)
von: Chen, Yi-Pei, et al.
Veröffentlicht: (2024)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
von: Chen, Junjie, et al.
Veröffentlicht: (2026)
von: Chen, Junjie, et al.
Veröffentlicht: (2026)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
von: Hamna, Hamna, et al.
Veröffentlicht: (2025)
von: Hamna, Hamna, et al.
Veröffentlicht: (2025)
Limits to Predicting Online Speech Using Large Language Models
von: Remeli, Mina, et al.
Veröffentlicht: (2024)
von: Remeli, Mina, et al.
Veröffentlicht: (2024)
Arabic Chatbot Technologies in Education: An Overview
von: Bourhil, Hicham, et al.
Veröffentlicht: (2025)
von: Bourhil, Hicham, et al.
Veröffentlicht: (2025)
On the Implications of Verbose LLM Outputs: A Case Study in Translation Evaluation
von: Briakou, Eleftheria, et al.
Veröffentlicht: (2024)
von: Briakou, Eleftheria, et al.
Veröffentlicht: (2024)
A Course Shared Task on Evaluating LLM Output for Clinical Questions
von: Hou, Yufang, et al.
Veröffentlicht: (2024)
von: Hou, Yufang, et al.
Veröffentlicht: (2024)
How Well Can LLMs Echo Us? Evaluating AI Chatbots' Role-Play Ability with ECHO
von: Ng, Man Tik, et al.
Veröffentlicht: (2024)
von: Ng, Man Tik, et al.
Veröffentlicht: (2024)
Domain-Specific Improvement on Psychotherapy Chatbot Using Assistant
von: Kang, Cheng, et al.
Veröffentlicht: (2024)
von: Kang, Cheng, et al.
Veröffentlicht: (2024)
Through the Prism of Culture: Evaluating LLMs' Understanding of Indian Subcultures and Traditions
von: Chhikara, Garima, et al.
Veröffentlicht: (2025)
von: Chhikara, Garima, et al.
Veröffentlicht: (2025)
Mixed Chain-of-Psychotherapies for Emotional Support Chatbot
von: Chen, Siyuan, et al.
Veröffentlicht: (2024)
von: Chen, Siyuan, et al.
Veröffentlicht: (2024)
LLM Roleplay: Simulating Human-Chatbot Interaction
von: Tamoyan, Hovhannes, et al.
Veröffentlicht: (2024)
von: Tamoyan, Hovhannes, et al.
Veröffentlicht: (2024)
Sólo Escúchame: Spanish Emotional Accompaniment Chatbot
von: Ramírez, Bruno Gil, et al.
Veröffentlicht: (2024)
von: Ramírez, Bruno Gil, et al.
Veröffentlicht: (2024)
Empirical Study of Symmetrical Reasoning in Conversational Chatbots
von: Rim, Daniela N., et al.
Veröffentlicht: (2024)
von: Rim, Daniela N., et al.
Veröffentlicht: (2024)
Evaluation of LLM Chatbots for OSINT-based Cyber Threat Awareness
von: Shafee, Samaneh, et al.
Veröffentlicht: (2024)
von: Shafee, Samaneh, et al.
Veröffentlicht: (2024)
The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models
von: Singh, Abhinav Kumar, et al.
Veröffentlicht: (2026)
von: Singh, Abhinav Kumar, et al.
Veröffentlicht: (2026)
Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework
von: Jain, Shomik, et al.
Veröffentlicht: (2025)
von: Jain, Shomik, et al.
Veröffentlicht: (2025)
First-Person Fairness in Chatbots
von: Eloundou, Tyna, et al.
Veröffentlicht: (2024)
von: Eloundou, Tyna, et al.
Veröffentlicht: (2024)
LLMs vs. Traditional Sentiment Tools in Psychology: An Evaluation on Belgian-Dutch Narratives
von: Kandala, Ratna, et al.
Veröffentlicht: (2025)
von: Kandala, Ratna, et al.
Veröffentlicht: (2025)
A Comparison of LLM Finetuning Methods & Evaluation Metrics with Travel Chatbot Use Case
von: Meyer, Sonia, et al.
Veröffentlicht: (2024)
von: Meyer, Sonia, et al.
Veröffentlicht: (2024)
Distinguishing Chatbot from Human
von: Godghase, Gauri Anil, et al.
Veröffentlicht: (2024)
von: Godghase, Gauri Anil, et al.
Veröffentlicht: (2024)
Knowledge Graph-Driven Retrieval-Augmented Generation: Integrating Deepseek-R1 with Weaviate for Advanced Chatbot Applications
von: Lecu, Alexandru, et al.
Veröffentlicht: (2025)
von: Lecu, Alexandru, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
KatzBot: Revolutionizing Academic Chatbot for Enhanced Communication
von: Kumar, Sahil, et al.
Veröffentlicht: (2024) -
Evaluating Commercial AI Chatbots as News Intermediaries
von: Suzgun, Mirac, et al.
Veröffentlicht: (2026) -
ChatbotManip: A Dataset to Facilitate Evaluation and Oversight of Manipulative Chatbot Behaviour
von: Contro, Jack, et al.
Veröffentlicht: (2025) -
Evaluating language models as risk scores
von: Cruz, André F., et al.
Veröffentlicht: (2024) -
LMStyle Benchmark: Evaluating Text Style Transfer for Chatbots
von: Chen, Jianlin
Veröffentlicht: (2024)