Multi-Dimensional Evaluation of Sustainable City Trips with LLM-as-a-Judge and Human-in-the-Loop
Fuente:
arXiv
Guardado en:
| Autores principales: | Banerjee, Ashmi, Satish, Adithi, Wörndl, Wolfgang, Deldjoo, Yashar |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TRACE: A Conversational Framework for Sustainable Tourism Recommendation with Agentic Counterfactual Explanations
por: Banerjee, Ashmi, et al.
Publicado: (2026)
por: Banerjee, Ashmi, et al.
Publicado: (2026)
Collab-REC: An LLM-based Agentic Framework for Balancing Recommendations in Tourism
por: Banerjee, Ashmi, et al.
Publicado: (2025)
por: Banerjee, Ashmi, et al.
Publicado: (2025)
SynthTRIPs: A Knowledge-Grounded Framework for Benchmark Query Generation for Personalized Tourism Recommenders
por: Banerjee, Ashmi, et al.
Publicado: (2025)
por: Banerjee, Ashmi, et al.
Publicado: (2025)
Enhancing Tourism Recommender Systems for Sustainable City Trips Using Retrieval-Augmented Generation
por: Banerjee, Ashmi, et al.
Publicado: (2024)
por: Banerjee, Ashmi, et al.
Publicado: (2024)
A User Interface Study on Sustainable City Trip Recommendations
por: Banerjee, Ashmi, et al.
Publicado: (2024)
por: Banerjee, Ashmi, et al.
Publicado: (2024)
SmartSustain Recommender System: Navigating Sustainability Trade-offs in Personalized City Trip Planning
por: Banerjee, Ashmi, et al.
Publicado: (2025)
por: Banerjee, Ashmi, et al.
Publicado: (2025)
Modeling Sustainable City Trips: Integrating CO2e Emissions, Popularity, and Seasonality into Tourism Recommender Systems
por: Banerjee, Ashmi, et al.
Publicado: (2024)
por: Banerjee, Ashmi, et al.
Publicado: (2024)
A Normative Framework for Benchmarking Consumer Fairness in Large Language Model Recommender System
por: Deldjoo, Yashar, et al.
Publicado: (2024)
por: Deldjoo, Yashar, et al.
Publicado: (2024)
A Personalized Framework for Consumer and Producer Group Fairness Optimization in Recommender Systems
por: Rahmani, Hossein A., et al.
Publicado: (2024)
por: Rahmani, Hossein A., et al.
Publicado: (2024)
XAI4LLM. Let Machine Learning Models and LLMs Collaborate for Enhanced In-Context Learning in Healthcare
por: Nazary, Fatemeh, et al.
Publicado: (2024)
por: Nazary, Fatemeh, et al.
Publicado: (2024)
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
por: Do, Hyo Jin, et al.
Publicado: (2025)
por: Do, Hyo Jin, et al.
Publicado: (2025)
Toward Holistic Evaluation of Recommender Systems Powered by Generative Models
por: Deldjoo, Yashar, et al.
Publicado: (2025)
por: Deldjoo, Yashar, et al.
Publicado: (2025)
Judging the Judges: Human Validation of Multi-LLM Evaluation for High-Quality K--12 Science Instructional Materials
por: He, Peng, et al.
Publicado: (2026)
por: He, Peng, et al.
Publicado: (2026)
Towards Recommender Systems LLMs Playground (RecSysLLMsP): Exploring Polarization and Engagement in Simulated Social Networks
por: Bojic, Ljubisa, et al.
Publicado: (2025)
por: Bojic, Ljubisa, et al.
Publicado: (2025)
Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering
por: Zhao, Zixiao, et al.
Publicado: (2026)
por: Zhao, Zixiao, et al.
Publicado: (2026)
Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines
por: Soumik, Sadman Kabir
Publicado: (2026)
por: Soumik, Sadman Kabir
Publicado: (2026)
Understanding Biases in ChatGPT-based Recommender Systems: Provider Fairness, Temporal Stability, and Recency
por: Deldjoo, Yashar
Publicado: (2024)
por: Deldjoo, Yashar
Publicado: (2024)
LoopBench: Discovering Emergent Symmetry Breaking Strategies with LLM Swarms
por: Parsaee, Ali, et al.
Publicado: (2025)
por: Parsaee, Ali, et al.
Publicado: (2025)
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments
por: Li, Yuran, et al.
Publicado: (2025)
por: Li, Yuran, et al.
Publicado: (2025)
Zero-Shot LLMs in Human-in-the-Loop RL: Replacing Human Feedback for Reward Shaping
por: Nazir, Mohammad Saif, et al.
Publicado: (2025)
por: Nazir, Mohammad Saif, et al.
Publicado: (2025)
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
por: Li, Xiaochuan, et al.
Publicado: (2025)
por: Li, Xiaochuan, et al.
Publicado: (2025)
CPS-LLM: Large Language Model based Safe Usage Plan Generator for Human-in-the-Loop Human-in-the-Plant Cyber-Physical System
por: Banerjee, Ayan, et al.
Publicado: (2024)
por: Banerjee, Ayan, et al.
Publicado: (2024)
Evaluating Metrics for Safety with LLM-as-Judges
por: Clegg, Kester, et al.
Publicado: (2025)
por: Clegg, Kester, et al.
Publicado: (2025)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
por: Wang, Ruiqi, et al.
Publicado: (2025)
por: Wang, Ruiqi, et al.
Publicado: (2025)
MADP: A Multi-Agent Pipeline for Sustainable Document Processing with Human-in-the-Loop
por: Gosmar, Diego, et al.
Publicado: (2026)
por: Gosmar, Diego, et al.
Publicado: (2026)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
por: Zhou, Xin, et al.
Publicado: (2025)
por: Zhou, Xin, et al.
Publicado: (2025)
Machine-learned Adversarial Attacks against Fault Prediction Systems in Smart Electrical Grids
por: Ardito, Carmelo, et al.
Publicado: (2023)
por: Ardito, Carmelo, et al.
Publicado: (2023)
MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation
por: Wang, Yutong, et al.
Publicado: (2025)
por: Wang, Yutong, et al.
Publicado: (2025)
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
por: Han, Steve, et al.
Publicado: (2025)
por: Han, Steve, et al.
Publicado: (2025)
A Multi-Agent Human-LLM Collaborative Framework for Closed-Loop Scientific Literature Summarization
por: Jacobson, Maxwell J., et al.
Publicado: (2026)
por: Jacobson, Maxwell J., et al.
Publicado: (2026)
Beyond correlation: The Impact of Human Uncertainty in Measuring the Effectiveness of Automatic Evaluation and LLM-as-a-Judge
por: Elangovan, Aparna, et al.
Publicado: (2024)
por: Elangovan, Aparna, et al.
Publicado: (2024)
Adobe Summit Concierge Evaluation with Human in the Loop
por: Chen, Yiru, et al.
Publicado: (2025)
por: Chen, Yiru, et al.
Publicado: (2025)
Beyond Consensus: Mitigating the Agreeableness Bias in LLM Judge Evaluations
por: Jain, Suryaansh, et al.
Publicado: (2025)
por: Jain, Suryaansh, et al.
Publicado: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
por: Tan, Sijun, et al.
Publicado: (2024)
por: Tan, Sijun, et al.
Publicado: (2024)
Multi-Dimensional Behavioral Evaluation of Agentic Stock Prediction Systems Using Large Language Model Judges with Closed-Loop Reinforcement Learning Feedback
por: Ridhawi, Mohammad Al, et al.
Publicado: (2026)
por: Ridhawi, Mohammad Al, et al.
Publicado: (2026)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
por: Liu, Yixin, et al.
Publicado: (2025)
por: Liu, Yixin, et al.
Publicado: (2025)
Multi-Agent Debate for LLM Judges with Adaptive Stability Detection
por: Hu, Tianyu, et al.
Publicado: (2025)
por: Hu, Tianyu, et al.
Publicado: (2025)
Agent-in-the-Loop: A Data Flywheel for Continuous Improvement in LLM-based Customer Support
por: Zhao, Cen Mia, et al.
Publicado: (2025)
por: Zhao, Cen Mia, et al.
Publicado: (2025)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
por: Saha, Swarnadeep, et al.
Publicado: (2025)
por: Saha, Swarnadeep, et al.
Publicado: (2025)
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
por: Huang, Tzu-Heng, et al.
Publicado: (2025)
por: Huang, Tzu-Heng, et al.
Publicado: (2025)
Ejemplares similares
-
TRACE: A Conversational Framework for Sustainable Tourism Recommendation with Agentic Counterfactual Explanations
por: Banerjee, Ashmi, et al.
Publicado: (2026) -
Collab-REC: An LLM-based Agentic Framework for Balancing Recommendations in Tourism
por: Banerjee, Ashmi, et al.
Publicado: (2025) -
SynthTRIPs: A Knowledge-Grounded Framework for Benchmark Query Generation for Personalized Tourism Recommenders
por: Banerjee, Ashmi, et al.
Publicado: (2025) -
Enhancing Tourism Recommender Systems for Sustainable City Trips Using Retrieval-Augmented Generation
por: Banerjee, Ashmi, et al.
Publicado: (2024) -
A User Interface Study on Sustainable City Trip Recommendations
por: Banerjee, Ashmi, et al.
Publicado: (2024)