Cost-Effective Online Multi-LLM Selection with Versatile Reward Models
Fuente:
arXiv
Saved in:
| Main Authors: | Dai, Xiangxiang, Li, Jin, Liu, Xutong, Yu, Anqi, Lui, John C. S. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Multi-Agent Conversational Bandit Approach to Online Evaluation and Selection of User-Aligned LLM Responses
by: Dai, Xiangxiang, et al.
Published: (2025)
by: Dai, Xiangxiang, et al.
Published: (2025)
Steering Frozen LLMs: Adaptive Social Alignment via Online Prompt Routing
by: Zhang, Zeyu, et al.
Published: (2026)
by: Zhang, Zeyu, et al.
Published: (2026)
A User Study on Contrastive Explanations for Multi-Effector Temporal Planning with Non-Stationary Costs
by: Liu, Xiaowei, et al.
Published: (2024)
by: Liu, Xiaowei, et al.
Published: (2024)
Explanation through Reward Model Reconciliation using POMDP Tree Search
by: Kraske, Benjamin D., et al.
Published: (2023)
by: Kraske, Benjamin D., et al.
Published: (2023)
Hindsight PRIORs for Reward Learning from Human Preferences
by: Verma, Mudit, et al.
Published: (2024)
by: Verma, Mudit, et al.
Published: (2024)
Why and When LLM-Based Assistants Can Go Wrong: Investigating the Effectiveness of Prompt-Based Interactions for Software Help-Seeking
by: Khurana, Anjali, et al.
Published: (2024)
by: Khurana, Anjali, et al.
Published: (2024)
Integrating Policy Summaries with Reward Decomposition for Explaining Reinforcement Learning Agents
by: Septon, Yael, et al.
Published: (2022)
by: Septon, Yael, et al.
Published: (2022)
Analyzing Character Representation in Media Content using Multimodal Foundation Model: Effectiveness and Trust
by: Taka, Evdoxia, et al.
Published: (2025)
by: Taka, Evdoxia, et al.
Published: (2025)
Never Start from Scratch: Expediting On-Device LLM Personalization via Explainable Model Selection
by: Wang, Haoming, et al.
Published: (2025)
by: Wang, Haoming, et al.
Published: (2025)
NeuroRVQ: Multi-Scale Biosignal Tokenization for Generative Foundation Models
by: Barmpas, Konstantinos, et al.
Published: (2025)
by: Barmpas, Konstantinos, et al.
Published: (2025)
Classroom Simulacra: Building Contextual Student Generative Agents in Online Education for Learning Behavioral Simulation
by: Xu, Songlin, et al.
Published: (2025)
by: Xu, Songlin, et al.
Published: (2025)
AdaShadow: Responsive Test-time Model Adaptation in Non-stationary Mobile Environments
by: Fang, Cheng, et al.
Published: (2024)
by: Fang, Cheng, et al.
Published: (2024)
Robots That Know What to Ask: Recovering Misaligned Rewards through Targeted Explanations
by: Merker, Helena, et al.
Published: (2026)
by: Merker, Helena, et al.
Published: (2026)
Improving User Experience in Preference-Based Optimization of Reward Functions for Assistive Robots
by: Dennler, Nathaniel, et al.
Published: (2024)
by: Dennler, Nathaniel, et al.
Published: (2024)
From Commands to Prompts: LLM-based Semantic File System for AIOS
by: Shi, Zeru, et al.
Published: (2024)
by: Shi, Zeru, et al.
Published: (2024)
Off-Policy Selection for Initiating Human-Centric Experimental Design
by: Gao, Ge, et al.
Published: (2024)
by: Gao, Ge, et al.
Published: (2024)
OnGoal: Tracking and Visualizing Conversational Goals in Multi-Turn Dialogue with Large Language Models
by: Coscia, Adam, et al.
Published: (2025)
by: Coscia, Adam, et al.
Published: (2025)
AI Agents for Inventory Control: Human-LLM-OR Complementarity
by: Baek, Jackie, et al.
Published: (2026)
by: Baek, Jackie, et al.
Published: (2026)
A Multi-Component AI Framework for Computational Psychology: From Robust Predictive Modeling to Deployed Generative Dialogue
by: Pareek, Anant
Published: (2025)
by: Pareek, Anant
Published: (2025)
Helping Customers in Distress: An LLM-powered Agent that Converses, Probes, and Routes
by: Atreya, Alankar, et al.
Published: (2026)
by: Atreya, Alankar, et al.
Published: (2026)
Personas Evolved: Designing Ethical LLM-Based Conversational Agent Personalities
by: Desai, Smit, et al.
Published: (2025)
by: Desai, Smit, et al.
Published: (2025)
DiscoverLLM: From Executing Intents to Discovering Them
by: Kim, Tae Soo, et al.
Published: (2026)
by: Kim, Tae Soo, et al.
Published: (2026)
Learning Social Cost Functions for Human-Aware Path Planning
by: Eirale, Andrea, et al.
Published: (2024)
by: Eirale, Andrea, et al.
Published: (2024)
RuleEdit: Failure-Guided Human-AI Model Editing with Prospective Impact Preview
by: Lee, Min Hun, et al.
Published: (2026)
by: Lee, Min Hun, et al.
Published: (2026)
Deception at Scale: Deceptive Designs in 1K LLM-Generated Ecommerce Components
by: Chen, Ziwei, et al.
Published: (2025)
by: Chen, Ziwei, et al.
Published: (2025)
Enhancing Adaptive Behavioral Interventions with LLM Inference from Participant-Described States
by: Karine, Karine, et al.
Published: (2025)
by: Karine, Karine, et al.
Published: (2025)
LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
by: Park, Joon Sung, et al.
Published: (2024)
by: Park, Joon Sung, et al.
Published: (2024)
Roamify: Designing and Evaluating an LLM Based Google Chrome Extension for Personalised Itinerary Planning
by: Udandarao, Vikranth, et al.
Published: (2025)
by: Udandarao, Vikranth, et al.
Published: (2025)
U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning
by: Lee, Christine P, et al.
Published: (2026)
by: Lee, Christine P, et al.
Published: (2026)
Interaction Dynamics as a Reward Signal for LLMs
by: Gooding, Sian, et al.
Published: (2025)
by: Gooding, Sian, et al.
Published: (2025)
Learning to Assist Humans without Inferring Rewards
by: Myers, Vivek, et al.
Published: (2024)
by: Myers, Vivek, et al.
Published: (2024)
Multi-Turn Human-LLM Interaction Through the Lens of a Two-Way Intelligibility Protocol
by: Mestha, Harshvardhan, et al.
Published: (2024)
by: Mestha, Harshvardhan, et al.
Published: (2024)
LLMs May Not Be Human-Level Players, But They Can Be Testers: Measuring Game Difficulty with LLM Agents
by: Xiao, Chang, et al.
Published: (2024)
by: Xiao, Chang, et al.
Published: (2024)
Evaluating Interactive 2D Visualization as a Sample Selection Strategy for Biomedical Time-Series Data Annotation
by: Vaaras, Einari, et al.
Published: (2026)
by: Vaaras, Einari, et al.
Published: (2026)
Vi(E)va LLM! A Conceptual Stack for Evaluating and Interpreting Generative AI-based Visualizations
by: Podo, Luca, et al.
Published: (2024)
by: Podo, Luca, et al.
Published: (2024)
Scaling Wearable Foundation Models
by: Narayanswamy, Girish, et al.
Published: (2024)
by: Narayanswamy, Girish, et al.
Published: (2024)
Agent Laboratory: Using LLM Agents as Research Assistants
by: Schmidgall, Samuel, et al.
Published: (2025)
by: Schmidgall, Samuel, et al.
Published: (2025)
When the Loop Closes: Architectural Limits of In-Context Isolation, Metacognitive Co-option, and the Two-Target Design Problem in Human-LLM Systems
by: Cheng, Z., et al.
Published: (2026)
by: Cheng, Z., et al.
Published: (2026)
Exploring Emotions in Multi-componential Space using Interactive VR Games
by: Somarathna, Rukshani, et al.
Published: (2024)
by: Somarathna, Rukshani, et al.
Published: (2024)
CooT: Learning to Coordinate In-Context with Coordination Transformers
by: Wang, Huai-Chih, et al.
Published: (2025)
by: Wang, Huai-Chih, et al.
Published: (2025)
Similar Items
-
A Multi-Agent Conversational Bandit Approach to Online Evaluation and Selection of User-Aligned LLM Responses
by: Dai, Xiangxiang, et al.
Published: (2025) -
Steering Frozen LLMs: Adaptive Social Alignment via Online Prompt Routing
by: Zhang, Zeyu, et al.
Published: (2026) -
A User Study on Contrastive Explanations for Multi-Effector Temporal Planning with Non-Stationary Costs
by: Liu, Xiaowei, et al.
Published: (2024) -
Explanation through Reward Model Reconciliation using POMDP Tree Search
by: Kraske, Benjamin D., et al.
Published: (2023) -
Hindsight PRIORs for Reward Learning from Human Preferences
by: Verma, Mudit, et al.
Published: (2024)