A Comparison of LLM Finetuning Methods & Evaluation Metrics with Travel Chatbot Use Case
Fuente:
arXiv
Saved in:
| Main Authors: | Meyer, Sonia, Singh, Shreya, Tam, Bertha, Ton, Christopher, Ren, Angel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Dynamics of LLM Finetuning
by: Ren, Yi, et al.
Published: (2024)
by: Ren, Yi, et al.
Published: (2024)
Empowering Air Travelers: A Chatbot for Canadian Air Passenger Rights
by: Taranukhin, Maksym, et al.
Published: (2024)
by: Taranukhin, Maksym, et al.
Published: (2024)
Building Trust in Mental Health Chatbots: Safety Metrics and LLM-Based Evaluation Tools
by: Park, Jung In, et al.
Published: (2024)
by: Park, Jung In, et al.
Published: (2024)
Evaluating Metrics for Safety with LLM-as-Judges
by: Clegg, Kester, et al.
Published: (2025)
by: Clegg, Kester, et al.
Published: (2025)
Evaluation of Finetuned LLMs in AMR Parsing
by: Ho, Shu Han
Published: (2025)
by: Ho, Shu Han
Published: (2025)
TripTide: A Benchmark for Adaptive Travel Planning under Disruptions
by: Karmakar, Priyanshu, et al.
Published: (2025)
by: Karmakar, Priyanshu, et al.
Published: (2025)
A Complete Survey on LLM-based AI Chatbots
by: Dam, Sumit Kumar, et al.
Published: (2024)
by: Dam, Sumit Kumar, et al.
Published: (2024)
Citation-Enhanced Generation for LLM-based Chatbots
by: Li, Weitao, et al.
Published: (2024)
by: Li, Weitao, et al.
Published: (2024)
The Knowledge-Behaviour Disconnect in LLM-based Chatbots
by: Broersen, Jan
Published: (2025)
by: Broersen, Jan
Published: (2025)
Leveraging LLM-Respondents for Item Evaluation: a Psychometric Analysis
by: Liu, Yunting, et al.
Published: (2024)
by: Liu, Yunting, et al.
Published: (2024)
TripCraft: A Benchmark for Spatio-Temporally Fine Grained Travel Planning
by: Chaudhuri, Soumyabrata, et al.
Published: (2025)
by: Chaudhuri, Soumyabrata, et al.
Published: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
by: Yang, Langqi, et al.
Published: (2025)
by: Yang, Langqi, et al.
Published: (2025)
I Need Help! Evaluating LLM's Ability to Ask for Users' Support: A Case Study on Text-to-SQL Generation
by: Wu, Cheng-Kuang, et al.
Published: (2024)
by: Wu, Cheng-Kuang, et al.
Published: (2024)
DialogueForge: LLM Simulation of Human-Chatbot Dialogue
by: Zhu, Ruizhe, et al.
Published: (2025)
by: Zhu, Ruizhe, et al.
Published: (2025)
Reasoning before Comparison: LLM-Enhanced Semantic Similarity Metrics for Domain Specialized Text Analysis
by: Xu, Shaochen, et al.
Published: (2024)
by: Xu, Shaochen, et al.
Published: (2024)
ACC-Collab: An Actor-Critic Approach to Multi-Agent LLM Collaboration
by: Estornell, Andrew, et al.
Published: (2024)
by: Estornell, Andrew, et al.
Published: (2024)
Differentially Private Zeroth-Order Methods for Scalable Large Language Model Finetuning
by: Liu, Z, et al.
Published: (2024)
by: Liu, Z, et al.
Published: (2024)
ChaI-TeA: A Benchmark for Evaluating Autocompletion of Interactions with LLM-based Chatbots
by: Goren, Shani, et al.
Published: (2024)
by: Goren, Shani, et al.
Published: (2024)
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
by: Chiang, Wei-Lin, et al.
Published: (2024)
by: Chiang, Wei-Lin, et al.
Published: (2024)
Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics
by: Cho, Yousang, et al.
Published: (2025)
by: Cho, Yousang, et al.
Published: (2025)
Beyond Next Word Prediction: Developing Comprehensive Evaluation Frameworks for measuring LLM performance on real world applications
by: Agrawal, Vishakha, et al.
Published: (2025)
by: Agrawal, Vishakha, et al.
Published: (2025)
ReFT: Representation Finetuning for Language Models
by: Wu, Zhengxuan, et al.
Published: (2024)
by: Wu, Zhengxuan, et al.
Published: (2024)
TravelBench : Exploring LLM Performance in Low-Resource Domains
by: Billa, Srinivas, et al.
Published: (2025)
by: Billa, Srinivas, et al.
Published: (2025)
Case-Based Calibration of Adaptive Reasoning and Execution for LLM Tool Use
by: Pang, Renning, et al.
Published: (2026)
by: Pang, Renning, et al.
Published: (2026)
Hybrid-NL2SVA: Integrating RAG and Finetuning for LLM-based NL2SVA
by: Xiao, Weihua, et al.
Published: (2025)
by: Xiao, Weihua, et al.
Published: (2025)
FineScope : SAE-guided Data Selection Enables Domain Specific LLM Pruning and Finetuning
by: Bhattacharyya, Chaitali, et al.
Published: (2025)
by: Bhattacharyya, Chaitali, et al.
Published: (2025)
Arithmetic Reasoning with LLM: Prolog Generation & Permutation
by: Yang, Xiaocheng, et al.
Published: (2024)
by: Yang, Xiaocheng, et al.
Published: (2024)
Whisper Finetuning on Nepali Language
by: Rijal, Sanjay, et al.
Published: (2024)
by: Rijal, Sanjay, et al.
Published: (2024)
Blending Human and LLM Expertise to Detect Hallucinations and Omissions in Mental Health Chatbot Responses
by: Hussain, Khizar, et al.
Published: (2026)
by: Hussain, Khizar, et al.
Published: (2026)
AGENTiGraph: A Multi-Agent Knowledge Graph Framework for Interactive, Domain-Specific LLM Chatbots
by: Zhao, Xinjie, et al.
Published: (2025)
by: Zhao, Xinjie, et al.
Published: (2025)
SpiroLLM: Finetuning Pretrained LLMs to Understand Spirogram Time Series with Clinical Validation in COPD Reporting
by: Mei, Shuhao, et al.
Published: (2025)
by: Mei, Shuhao, et al.
Published: (2025)
When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation
by: Sun, Bian, et al.
Published: (2026)
by: Sun, Bian, et al.
Published: (2026)
Simulating Meaning, Nevermore! Introducing ICR: A Semiotic-Hermeneutic Metric for Evaluating Meaning in LLM Text Summaries
by: Perez, Natalie, et al.
Published: (2026)
by: Perez, Natalie, et al.
Published: (2026)
Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic
by: Ingimundarson, Finnur Ágúst, et al.
Published: (2026)
by: Ingimundarson, Finnur Ágúst, et al.
Published: (2026)
From Biased Chatbots to Biased Agents: Examining Role Assignment Effects on LLM Agent Robustness
by: Cao, Linbo, et al.
Published: (2026)
by: Cao, Linbo, et al.
Published: (2026)
Understanding Learner-LLM Chatbot Interactions and the Impact of Prompting Guidelines
by: Koyuturk, Cansu, et al.
Published: (2025)
by: Koyuturk, Cansu, et al.
Published: (2025)
LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces
by: Kirgis, Peter, et al.
Published: (2026)
by: Kirgis, Peter, et al.
Published: (2026)
TherapyGym: Evaluating and Aligning Clinical Fidelity and Safety in Therapy Chatbots
by: Huang, Fangrui, et al.
Published: (2026)
by: Huang, Fangrui, et al.
Published: (2026)
A Chatbot for Asylum-Seeking Migrants in Europe
by: Fazzinga, Bettina, et al.
Published: (2024)
by: Fazzinga, Bettina, et al.
Published: (2024)
TravelAgent: An AI Assistant for Personalized Travel Planning
by: Chen, Aili, et al.
Published: (2024)
by: Chen, Aili, et al.
Published: (2024)
Similar Items
-
Learning Dynamics of LLM Finetuning
by: Ren, Yi, et al.
Published: (2024) -
Empowering Air Travelers: A Chatbot for Canadian Air Passenger Rights
by: Taranukhin, Maksym, et al.
Published: (2024) -
Building Trust in Mental Health Chatbots: Safety Metrics and LLM-Based Evaluation Tools
by: Park, Jung In, et al.
Published: (2024) -
Evaluating Metrics for Safety with LLM-as-Judges
by: Clegg, Kester, et al.
Published: (2025) -
Evaluation of Finetuned LLMs in AMR Parsing
by: Ho, Shu Han
Published: (2025)