Personalized Benchmarking: Evaluating LLMs by Individual Preferences
Fuente:
arXiv
Saved in:
| Main Authors: | Garbacea, Cristina, Wang, Heran, Tan, Chenhao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ink and Individuality: Crafting a Personalised Narrative in the Age of LLMs
by: Wasi, Azmine Toushik, et al.
Published: (2024)
by: Wasi, Azmine Toushik, et al.
Published: (2024)
Bridging the Skills Gap: Evaluating an AI-Assisted Provider Platform to Support Care Providers with Empathetic Delivery of Protocolized Therapy
by: Kearns, William R., et al.
Published: (2024)
by: Kearns, William R., et al.
Published: (2024)
RuleAlign: Making Large Language Models Better Physicians with Diagnostic Rule Alignment
by: Wang, Xiaohan, et al.
Published: (2024)
by: Wang, Xiaohan, et al.
Published: (2024)
MeMemo: On-device Retrieval Augmentation for Private and Personalized Text Generation
by: Wang, Zijie J., et al.
Published: (2024)
by: Wang, Zijie J., et al.
Published: (2024)
SCOPE: A Lightweight-training LLM Framework for Air Traffic Control Readback Monitoring
by: Deng, Qihan, et al.
Published: (2026)
by: Deng, Qihan, et al.
Published: (2026)
RoTRAG: Rule of Thumb Reasoning for Conversation Harm Detection with Retrieval-Augmented Generation
by: Lee, Juhyeon, et al.
Published: (2026)
by: Lee, Juhyeon, et al.
Published: (2026)
From Bytes to Biases: Investigating the Cultural Self-Perception of Large Language Models
by: Messner, Wolfgang, et al.
Published: (2023)
by: Messner, Wolfgang, et al.
Published: (2023)
EasyInstruct: An Easy-to-use Instruction Processing Framework for Large Language Models
by: Ou, Yixin, et al.
Published: (2024)
by: Ou, Yixin, et al.
Published: (2024)
OmniThink: Expanding Knowledge Boundaries in Machine Writing through Thinking
by: Xi, Zekun, et al.
Published: (2025)
by: Xi, Zekun, et al.
Published: (2025)
User Perception of Attention Visualizations: Effects on Interpretability Across Evidence-Based Medical Documents
by: Carvallo, Andrés, et al.
Published: (2025)
by: Carvallo, Andrés, et al.
Published: (2025)
Cascading Adaptors to Leverage English Data to Improve Performance of Question Answering for Low-Resource Languages
by: Pandya, Hariom A., et al.
Published: (2021)
by: Pandya, Hariom A., et al.
Published: (2021)
WildVis: Open Source Visualizer for Million-Scale Chat Logs in the Wild
by: Deng, Yuntian, et al.
Published: (2024)
by: Deng, Yuntian, et al.
Published: (2024)
Generative AI in the Construction Industry: A State-of-the-art Analysis
by: Taiwo, Ridwan, et al.
Published: (2024)
by: Taiwo, Ridwan, et al.
Published: (2024)
Making Language Models Better Tool Learners with Execution Feedback
by: Qiao, Shuofei, et al.
Published: (2023)
by: Qiao, Shuofei, et al.
Published: (2023)
Towards End-to-End Open Conversational Machine Reading
by: Zhou, Sizhe, et al.
Published: (2022)
by: Zhou, Sizhe, et al.
Published: (2022)
Observations on LLMs for Telecom Domain: Capabilities and Limitations
by: Soman, Sumit, et al.
Published: (2023)
by: Soman, Sumit, et al.
Published: (2023)
To LLM, or Not to LLM: How Designers and Developers Navigate LLMs as Tools or Teammates
by: Vishwarupe, Varad, et al.
Published: (2026)
by: Vishwarupe, Varad, et al.
Published: (2026)
User Preference Modeling for Conversational LLM Agents: Weak Rewards from Retrieval-Augmented Interaction
by: Hao, Yuren, et al.
Published: (2026)
by: Hao, Yuren, et al.
Published: (2026)
The Agentic Leash: Extracting Causal Feedback Fuzzy Cognitive Maps with LLMs
by: Panda, Akash Kumar, et al.
Published: (2025)
by: Panda, Akash Kumar, et al.
Published: (2025)
User Simulation in the Era of Generative AI: User Modeling, Synthetic Data Generation, and System Evaluation
by: Balog, Krisztian, et al.
Published: (2025)
by: Balog, Krisztian, et al.
Published: (2025)
Do LLMs Provide Consistent Answers to Health-Related Questions across Languages?
by: Schlicht, Ipek Baris, et al.
Published: (2025)
by: Schlicht, Ipek Baris, et al.
Published: (2025)
Anna Karenina Strikes Again: Pre-Trained LLM Embeddings May Favor High-Performing Learners
by: Schleifer, Abigail Gurin, et al.
Published: (2024)
by: Schleifer, Abigail Gurin, et al.
Published: (2024)
Semantic Interaction for Narrative Map Sensemaking: An Insight-based Evaluation
by: Keith-Norambuena, Brian Felipe, et al.
Published: (2026)
by: Keith-Norambuena, Brian Felipe, et al.
Published: (2026)
Emotion-Driven Personalized Recommendation for AI-Generated Content Using Multi-Modal Sentiment and Intent Analysis
by: Hu, Zheqi, et al.
Published: (2025)
by: Hu, Zheqi, et al.
Published: (2025)
SteerEval: A Framework for Evaluating Steerability with Natural Language Profiles for Recommendation
by: Zhou, Joyce, et al.
Published: (2026)
by: Zhou, Joyce, et al.
Published: (2026)
USE: Dynamic User Modeling with Stateful Sequence Models
by: Zhou, Zhihan, et al.
Published: (2024)
by: Zhou, Zhihan, et al.
Published: (2024)
Rewriting Conversational Utterances with Instructed Large Language Models
by: Galimzhanova, Elnara, et al.
Published: (2024)
by: Galimzhanova, Elnara, et al.
Published: (2024)
Retrieve, Annotate, Evaluate, Repeat: Leveraging Multimodal LLMs for Large-Scale Product Retrieval Evaluation
by: Hosseini, Kasra, et al.
Published: (2024)
by: Hosseini, Kasra, et al.
Published: (2024)
What should I wear to a party in a Greek taverna? Evaluation for Conversational Agents in the Fashion Domain
by: Maronikolakis, Antonis, et al.
Published: (2024)
by: Maronikolakis, Antonis, et al.
Published: (2024)
EHR-MCP: Real-world Evaluation of Clinical Information Retrieval by Large Language Models via Model Context Protocol
by: Masayoshi, Kanato, et al.
Published: (2025)
by: Masayoshi, Kanato, et al.
Published: (2025)
Retentive Relevance: Capturing Long-Term User Value in Recommendation Systems
by: Bakhshi, Saeideh, et al.
Published: (2025)
by: Bakhshi, Saeideh, et al.
Published: (2025)
GraphSeek: Next-Generation Graph Analytics with LLMs
by: Besta, Maciej, et al.
Published: (2026)
by: Besta, Maciej, et al.
Published: (2026)
Clinical Reasoning AI for Oncology Treatment Planning: A Multi-Specialty Case-Based Evaluation
by: Spiess, Philippe E., et al.
Published: (2026)
by: Spiess, Philippe E., et al.
Published: (2026)
Are Generative AI Agents Effective Personalized Financial Advisors?
by: Takayanagi, Takehiro, et al.
Published: (2025)
by: Takayanagi, Takehiro, et al.
Published: (2025)
The Collaboration Gap in Human-AI Work
by: Vishwarupe, Varad, et al.
Published: (2026)
by: Vishwarupe, Varad, et al.
Published: (2026)
From Rights to Rites: Expectations Management in Smart-Home AI
by: Vishwarupe, Varad, et al.
Published: (2026)
by: Vishwarupe, Varad, et al.
Published: (2026)
FairEval: Evaluating Fairness in LLM-Based Recommendations with Personality Awareness
by: Sah, Chandan Kumar, et al.
Published: (2025)
by: Sah, Chandan Kumar, et al.
Published: (2025)
Almanac Copilot: Towards Autonomous Electronic Health Record Navigation
by: Zakka, Cyril, et al.
Published: (2024)
by: Zakka, Cyril, et al.
Published: (2024)
TreeHop: Generate and Filter Next Query Embeddings Efficiently for Multi-hop Question Answering
by: Li, Zhonghao, et al.
Published: (2025)
by: Li, Zhonghao, et al.
Published: (2025)
Knowledge Sharing in Manufacturing using Large Language Models: User Evaluation and Model Benchmarking
by: Freire, Samuel Kernan, et al.
Published: (2024)
by: Freire, Samuel Kernan, et al.
Published: (2024)
Similar Items
-
Ink and Individuality: Crafting a Personalised Narrative in the Age of LLMs
by: Wasi, Azmine Toushik, et al.
Published: (2024) -
Bridging the Skills Gap: Evaluating an AI-Assisted Provider Platform to Support Care Providers with Empathetic Delivery of Protocolized Therapy
by: Kearns, William R., et al.
Published: (2024) -
RuleAlign: Making Large Language Models Better Physicians with Diagnostic Rule Alignment
by: Wang, Xiaohan, et al.
Published: (2024) -
MeMemo: On-device Retrieval Augmentation for Private and Personalized Text Generation
by: Wang, Zijie J., et al.
Published: (2024) -
SCOPE: A Lightweight-training LLM Framework for Air Traffic Control Readback Monitoring
by: Deng, Qihan, et al.
Published: (2026)