Estimating LLM Consistency: A User Baseline vs Surrogate Metrics
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Xiaoyuan, Lin, Weiran, Akgul, Omer, Bauer, Lujo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses
di: Lin, Weiran, et al.
Pubblicazione: (2024)
di: Lin, Weiran, et al.
Pubblicazione: (2024)
User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios
di: Wu, Xiaoyuan, et al.
Pubblicazione: (2025)
di: Wu, Xiaoyuan, et al.
Pubblicazione: (2025)
Building Trust in Mental Health Chatbots: Safety Metrics and LLM-Based Evaluation Tools
di: Park, Jung In, et al.
Pubblicazione: (2024)
di: Park, Jung In, et al.
Pubblicazione: (2024)
UI-JEPA: Towards Active Perception of User Intent through Onscreen User Activity
di: Fu, Yicheng, et al.
Pubblicazione: (2024)
di: Fu, Yicheng, et al.
Pubblicazione: (2024)
Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models
di: Qi, Jirui, et al.
Pubblicazione: (2023)
di: Qi, Jirui, et al.
Pubblicazione: (2023)
Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood Consistency
di: Xu, Haoming, et al.
Pubblicazione: (2026)
di: Xu, Haoming, et al.
Pubblicazione: (2026)
LLM Attributor: Interactive Visual Attribution for LLM Generation
di: Lee, Seongmin, et al.
Pubblicazione: (2024)
di: Lee, Seongmin, et al.
Pubblicazione: (2024)
Agent Laboratory: Using LLM Agents as Research Assistants
di: Schmidgall, Samuel, et al.
Pubblicazione: (2025)
di: Schmidgall, Samuel, et al.
Pubblicazione: (2025)
"Excuse me, may I say something..." CoLabScience, A Proactive AI Assistant for Biomedical Discovery and LLM-Expert Collaborations
di: Wu, Yang, et al.
Pubblicazione: (2026)
di: Wu, Yang, et al.
Pubblicazione: (2026)
FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users
di: Singh, Anikait, et al.
Pubblicazione: (2025)
di: Singh, Anikait, et al.
Pubblicazione: (2025)
Survey of User Interface Design and Interaction Techniques in Generative AI Applications
di: Luera, Reuben, et al.
Pubblicazione: (2024)
di: Luera, Reuben, et al.
Pubblicazione: (2024)
Augmenting Automation: Intent-Based User Instruction Classification with Machine Learning
di: Basyal, Lochan, et al.
Pubblicazione: (2024)
di: Basyal, Lochan, et al.
Pubblicazione: (2024)
Evaluating LLM Adaptation to Sociodemographic Factors: User Profile vs. Dialogue History
di: Zhong, Qishuai, et al.
Pubblicazione: (2025)
di: Zhong, Qishuai, et al.
Pubblicazione: (2025)
Talking with Oompa Loompas: A novel framework for evaluating linguistic acquisition of LLM agents
di: Swain, Sankalp Tattwadarshi, et al.
Pubblicazione: (2025)
di: Swain, Sankalp Tattwadarshi, et al.
Pubblicazione: (2025)
DiscoverLLM: From Executing Intents to Discovering Them
di: Kim, Tae Soo, et al.
Pubblicazione: (2026)
di: Kim, Tae Soo, et al.
Pubblicazione: (2026)
Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors
di: Lam, Michelle S., et al.
Pubblicazione: (2024)
di: Lam, Michelle S., et al.
Pubblicazione: (2024)
Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
di: Thakkar, Nitya, et al.
Pubblicazione: (2025)
di: Thakkar, Nitya, et al.
Pubblicazione: (2025)
Multimodal Behavioral Patterns Analysis with Eye-Tracking and LLM-Based Reasoning
di: Guo, Dongyang, et al.
Pubblicazione: (2025)
di: Guo, Dongyang, et al.
Pubblicazione: (2025)
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
di: Cook, Jonathan, et al.
Pubblicazione: (2024)
di: Cook, Jonathan, et al.
Pubblicazione: (2024)
Creative Beam Search: LLM-as-a-Judge For Improving Response Generation
di: Franceschelli, Giorgio, et al.
Pubblicazione: (2024)
di: Franceschelli, Giorgio, et al.
Pubblicazione: (2024)
LLM Comparator: Visual Analytics for Side-by-Side Evaluation of Large Language Models
di: Kahng, Minsuk, et al.
Pubblicazione: (2024)
di: Kahng, Minsuk, et al.
Pubblicazione: (2024)
Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
di: Zhang, Lechen, et al.
Pubblicazione: (2025)
di: Zhang, Lechen, et al.
Pubblicazione: (2025)
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs
di: Baidya, Avinash, et al.
Pubblicazione: (2025)
di: Baidya, Avinash, et al.
Pubblicazione: (2025)
Never Start from Scratch: Expediting On-Device LLM Personalization via Explainable Model Selection
di: Wang, Haoming, et al.
Pubblicazione: (2025)
di: Wang, Haoming, et al.
Pubblicazione: (2025)
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
di: Kim, Jane Paik
Pubblicazione: (2026)
di: Kim, Jane Paik
Pubblicazione: (2026)
ParlAI Vote: A Web Platform for Analyzing Gender and Political Bias in Large Language Models
di: Lin, Wenjie, et al.
Pubblicazione: (2025)
di: Lin, Wenjie, et al.
Pubblicazione: (2025)
Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users
di: Karamolegkou, Antonia, et al.
Pubblicazione: (2025)
di: Karamolegkou, Antonia, et al.
Pubblicazione: (2025)
Are You Being Tracked? Discover the Power of Zero-Shot Trajectory Tracing with LLMs!
di: Yang, Huanqi, et al.
Pubblicazione: (2024)
di: Yang, Huanqi, et al.
Pubblicazione: (2024)
Large Language Models Can Infer Personality from Free-Form User Interactions
di: Peters, Heinrich, et al.
Pubblicazione: (2024)
di: Peters, Heinrich, et al.
Pubblicazione: (2024)
Llms, Virtual Users, and Bias: Predicting Any Survey Question Without Human Data
di: Sinacola, Enzo, et al.
Pubblicazione: (2025)
di: Sinacola, Enzo, et al.
Pubblicazione: (2025)
Who is a Better Matchmaker? Human vs. Algorithmic Judge Assignment in a High-Stakes Startup Competition
di: Xi, Sarina, et al.
Pubblicazione: (2025)
di: Xi, Sarina, et al.
Pubblicazione: (2025)
Properties and Challenges of LLM-Generated Explanations
di: Kunz, Jenny, et al.
Pubblicazione: (2024)
di: Kunz, Jenny, et al.
Pubblicazione: (2024)
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
di: Fu, Xiao, et al.
Pubblicazione: (2025)
di: Fu, Xiao, et al.
Pubblicazione: (2025)
Teaching According to Students' Aptitude: Personalized Mathematics Tutoring via Persona-, Memory-, and Forgetting-Aware LLMs
di: Wu, Yang, et al.
Pubblicazione: (2025)
di: Wu, Yang, et al.
Pubblicazione: (2025)
Do It For Me vs. Do It With Me: Investigating User Perceptions of Different Paradigms of Automation in Copilots for Feature-Rich Software
di: Khurana, Anjali, et al.
Pubblicazione: (2025)
di: Khurana, Anjali, et al.
Pubblicazione: (2025)
Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations
di: Seshadri, Preethi, et al.
Pubblicazione: (2026)
di: Seshadri, Preethi, et al.
Pubblicazione: (2026)
ABLEIST: Intersectional Disability Bias in LLM-Generated Hiring Scenarios
di: Phutane, Mahika, et al.
Pubblicazione: (2025)
di: Phutane, Mahika, et al.
Pubblicazione: (2025)
A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic
di: Brodeur, Peter, et al.
Pubblicazione: (2026)
di: Brodeur, Peter, et al.
Pubblicazione: (2026)
LLM4PM: A case study on using Large Language Models for Process Modeling in Enterprise Organizations
di: Ziche, Clara, et al.
Pubblicazione: (2024)
di: Ziche, Clara, et al.
Pubblicazione: (2024)
"They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated Conversations
di: Dammu, Preetam Prabhu Srikar, et al.
Pubblicazione: (2024)
di: Dammu, Preetam Prabhu Srikar, et al.
Pubblicazione: (2024)
Documenti analoghi
-
LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses
di: Lin, Weiran, et al.
Pubblicazione: (2024) -
User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios
di: Wu, Xiaoyuan, et al.
Pubblicazione: (2025) -
Building Trust in Mental Health Chatbots: Safety Metrics and LLM-Based Evaluation Tools
di: Park, Jung In, et al.
Pubblicazione: (2024) -
UI-JEPA: Towards Active Perception of User Intent through Onscreen User Activity
di: Fu, Yicheng, et al.
Pubblicazione: (2024) -
Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models
di: Qi, Jirui, et al.
Pubblicazione: (2023)