LifeSim: Long-Horizon User Life Simulator for Personalized Assistant Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Duan, Feiyu, Huang, Xuanjing, Wei, Zhongyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CURP: Codebook-based Continuous User Representation for Personalized Generation with LLMs
von: Wang, Liang, et al.
Veröffentlicht: (2026)
von: Wang, Liang, et al.
Veröffentlicht: (2026)
ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulation
von: Kim, Jiho, et al.
Veröffentlicht: (2025)
von: Kim, Jiho, et al.
Veröffentlicht: (2025)
Unveiling the Truth and Facilitating Change: Towards Agent-based Large-scale Social Movement Simulation
von: Mou, Xinyi, et al.
Veröffentlicht: (2024)
von: Mou, Xinyi, et al.
Veröffentlicht: (2024)
StreamProfileBench: A Benchmark for Fine-Grained User Profile Inference in Real-World Streaming Scenarios
von: Wang, Sizhe, et al.
Veröffentlicht: (2026)
von: Wang, Sizhe, et al.
Veröffentlicht: (2026)
EcoLANG: Efficient and Effective Agent Communication Language Induction for Social Simulation
von: Mou, Xinyi, et al.
Veröffentlicht: (2025)
von: Mou, Xinyi, et al.
Veröffentlicht: (2025)
MemSim: A Bayesian Simulator for Evaluating Memory of LLM-based Personal Assistants
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?
von: Dou, Yao, et al.
Veröffentlicht: (2025)
von: Dou, Yao, et al.
Veröffentlicht: (2025)
ElectionSim: Massive Population Election Simulation Powered by Large Language Model Driven Agents
von: Zhang, Xinnong, et al.
Veröffentlicht: (2024)
von: Zhang, Xinnong, et al.
Veröffentlicht: (2024)
Synergistic Multi-Agent Framework with Trajectory Learning for Knowledge-Intensive Tasks
von: Yue, Shengbin, et al.
Veröffentlicht: (2024)
von: Yue, Shengbin, et al.
Veröffentlicht: (2024)
PIORS: Personalized Intelligent Outpatient Reception based on Large Language Model with Multi-Agents Medical Scenario Simulation
von: Bao, Zhijie, et al.
Veröffentlicht: (2024)
von: Bao, Zhijie, et al.
Veröffentlicht: (2024)
Mem-PAL: Towards Memory-based Personalized Dialogue Assistants for Long-term User-Agent Interaction
von: Huang, Zhaopei, et al.
Veröffentlicht: (2025)
von: Huang, Zhaopei, et al.
Veröffentlicht: (2025)
Beyond Isolated Behaviors: Hierarchical User Modeling for LLM Personalization
von: Wang, Liang, et al.
Veröffentlicht: (2026)
von: Wang, Liang, et al.
Veröffentlicht: (2026)
When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation
von: Zou, Henry Peng, et al.
Veröffentlicht: (2026)
von: Zou, Henry Peng, et al.
Veröffentlicht: (2026)
Multi-Agent Simulator Drives Language Models for Legal Intensive Interaction
von: Yue, Shengbin, et al.
Veröffentlicht: (2025)
von: Yue, Shengbin, et al.
Veröffentlicht: (2025)
HorizonBench: Long-Horizon Personalization with Evolving Preferences
von: Li, Shuyue Stella, et al.
Veröffentlicht: (2026)
von: Li, Shuyue Stella, et al.
Veröffentlicht: (2026)
SalesSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators
von: Pruksachatkun, Yada, et al.
Veröffentlicht: (2026)
von: Pruksachatkun, Yada, et al.
Veröffentlicht: (2026)
Inside Out: Evolving User-Centric Core Memory Trees for Long-Term Personalized Dialogue Systems
von: Zhao, Jihao, et al.
Veröffentlicht: (2026)
von: Zhao, Jihao, et al.
Veröffentlicht: (2026)
AI-Press: A Multi-Agent News Generating and Feedback Simulation System Powered by Large Language Models
von: Liu, Xiawei, et al.
Veröffentlicht: (2024)
von: Liu, Xiawei, et al.
Veröffentlicht: (2024)
ALaRM: Align Language Models via Hierarchical Rewards Modeling
von: Lai, Yuhang, et al.
Veröffentlicht: (2024)
von: Lai, Yuhang, et al.
Veröffentlicht: (2024)
EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models
von: Du, Mengfei, et al.
Veröffentlicht: (2024)
von: Du, Mengfei, et al.
Veröffentlicht: (2024)
Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants
von: Suh, Joseph, et al.
Veröffentlicht: (2026)
von: Suh, Joseph, et al.
Veröffentlicht: (2026)
AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversations
von: Jiayang, Cheng, et al.
Veröffentlicht: (2026)
von: Jiayang, Cheng, et al.
Veröffentlicht: (2026)
Social Life Simulation for Non-Cognitive Skills Learning
von: Yan, Zihan, et al.
Veröffentlicht: (2024)
von: Yan, Zihan, et al.
Veröffentlicht: (2024)
Debatrix: Multi-dimensional Debate Judge with Iterative Chronological Analysis Based on LLM
von: Liang, Jingcong, et al.
Veröffentlicht: (2024)
von: Liang, Jingcong, et al.
Veröffentlicht: (2024)
Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis
von: Mok, Jisoo, et al.
Veröffentlicht: (2025)
von: Mok, Jisoo, et al.
Veröffentlicht: (2025)
VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
von: Li, Zejun, et al.
Veröffentlicht: (2024)
von: Li, Zejun, et al.
Veröffentlicht: (2024)
HAF-RM: A Hybrid Alignment Framework for Reward Model Training
von: Liu, Shujun, et al.
Veröffentlicht: (2024)
von: Liu, Shujun, et al.
Veröffentlicht: (2024)
Eval4Sim: An Evaluation Framework for Persona Simulation
von: Bao, Eliseo, et al.
Veröffentlicht: (2026)
von: Bao, Eliseo, et al.
Veröffentlicht: (2026)
LLM-based Medical Assistant Personalization with Short- and Long-Term Memory Coordination
von: Zhang, Kai, et al.
Veröffentlicht: (2023)
von: Zhang, Kai, et al.
Veröffentlicht: (2023)
From Individual to Society: A Survey on Social Simulation Driven by Large Language Model-based Agents
von: Mou, Xinyi, et al.
Veröffentlicht: (2024)
von: Mou, Xinyi, et al.
Veröffentlicht: (2024)
MemoryCD: Benchmarking Long-Context User Memory of LLM Agents for Lifelong Cross-Domain Personalization
von: Zhang, Weizhi, et al.
Veröffentlicht: (2026)
von: Zhang, Weizhi, et al.
Veröffentlicht: (2026)
How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation
von: Long, Zhuohang, et al.
Veröffentlicht: (2025)
von: Long, Zhuohang, et al.
Veröffentlicht: (2025)
EmoHarbor: Evaluating Personalized Emotional Support by Simulating the User's Internal World
von: Ye, Jing, et al.
Veröffentlicht: (2026)
von: Ye, Jing, et al.
Veröffentlicht: (2026)
InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
CauESC: A Causal Aware Model for Emotional Support Conversation
von: Chen, Wei, et al.
Veröffentlicht: (2024)
von: Chen, Wei, et al.
Veröffentlicht: (2024)
CL-bench Life: Can Language Models Learn from Real-Life Context?
von: Dou, Shihan, et al.
Veröffentlicht: (2026)
von: Dou, Shihan, et al.
Veröffentlicht: (2026)
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
von: Ding, Shuangrui, et al.
Veröffentlicht: (2026)
von: Ding, Shuangrui, et al.
Veröffentlicht: (2026)
OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions
von: Xu, Fangzhi, et al.
Veröffentlicht: (2026)
von: Xu, Fangzhi, et al.
Veröffentlicht: (2026)
AdaMem: Adaptive User-Centric Memory for Long-Horizon Dialogue Agents
von: Yan, Shannan, et al.
Veröffentlicht: (2026)
von: Yan, Shannan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CURP: Codebook-based Continuous User Representation for Personalized Generation with LLMs
von: Wang, Liang, et al.
Veröffentlicht: (2026) -
ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulation
von: Kim, Jiho, et al.
Veröffentlicht: (2025) -
Unveiling the Truth and Facilitating Change: Towards Agent-based Large-scale Social Movement Simulation
von: Mou, Xinyi, et al.
Veröffentlicht: (2024) -
StreamProfileBench: A Benchmark for Fine-Grained User Profile Inference in Real-World Streaming Scenarios
von: Wang, Sizhe, et al.
Veröffentlicht: (2026) -
EcoLANG: Efficient and Effective Agent Communication Language Induction for Social Simulation
von: Mou, Xinyi, et al.
Veröffentlicht: (2025)