HorizonBench: Long-Horizon Personalization with Evolving Preferences
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Shuyue Stella, Paranjape, Bhargavi, Oktar, Kerem, Ma, Zhongyao, Zhou, Gelin, Guan, Lin, Zhang, Na, Park, Sem, Chen, Lin, Yang, Diyi, Tsvetkov, Yulia, Celikyilmaz, Asli |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PrefPalette: Personalized Preference Modeling with Latent Attributes
von: Li, Shuyue Stella, et al.
Veröffentlicht: (2025)
von: Li, Shuyue Stella, et al.
Veröffentlicht: (2025)
Cold-Start Personalization via Training-Free Priors from Structured World Models
von: Bose, Avinandan, et al.
Veröffentlicht: (2026)
von: Bose, Avinandan, et al.
Veröffentlicht: (2026)
Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning
von: Sclar, Melanie, et al.
Veröffentlicht: (2024)
von: Sclar, Melanie, et al.
Veröffentlicht: (2024)
ComPO: Community Preferences for Language Model Personalization
von: Kumar, Sachin, et al.
Veröffentlicht: (2024)
von: Kumar, Sachin, et al.
Veröffentlicht: (2024)
Structured Preference Optimization for Vision-Language Long-Horizon Task Planning
von: Liang, Xiwen, et al.
Veröffentlicht: (2025)
von: Liang, Xiwen, et al.
Veröffentlicht: (2025)
PrefDisco: Benchmarking Proactive Personalized Reasoning
von: Li, Shuyue Stella, et al.
Veröffentlicht: (2025)
von: Li, Shuyue Stella, et al.
Veröffentlicht: (2025)
$π$-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
von: Zhang, Haoran, et al.
Veröffentlicht: (2026)
von: Zhang, Haoran, et al.
Veröffentlicht: (2026)
Cognitive Foundations for Reasoning and Their Manifestation in LLMs
von: Kargupta, Priyanka, et al.
Veröffentlicht: (2025)
von: Kargupta, Priyanka, et al.
Veröffentlicht: (2025)
O-Mem: Omni Memory System for Personalized, Long Horizon, Self-Evolving Agents
von: Wang, Piaohong, et al.
Veröffentlicht: (2025)
von: Wang, Piaohong, et al.
Veröffentlicht: (2025)
Towards Realistic Personalization: Evaluating Long-Horizon Preference Following in Personalized User-LLM Interactions
von: Guo, Qianyun, et al.
Veröffentlicht: (2026)
von: Guo, Qianyun, et al.
Veröffentlicht: (2026)
EvoLM: Self-Evolving Language Models through Co-Evolved Discriminative Rubrics
von: Li, Shuyue Stella, et al.
Veröffentlicht: (2026)
von: Li, Shuyue Stella, et al.
Veröffentlicht: (2026)
CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2024)
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2024)
TSUBASA: Improving Long-Horizon Personalization via Evolving Memory and Self-Learning with Context Distillation
von: Zhang, Xinliang Frederick, et al.
Veröffentlicht: (2026)
von: Zhang, Xinliang Frederick, et al.
Veröffentlicht: (2026)
Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search
von: Yen, Howard, et al.
Veröffentlicht: (2025)
von: Yen, Howard, et al.
Veröffentlicht: (2025)
Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
von: Wu, Addison J., et al.
Veröffentlicht: (2026)
von: Wu, Addison J., et al.
Veröffentlicht: (2026)
InfoGatherer: Principled Information Seeking via Evidence Retrieval and Strategic Questioning
von: Taranukhin, Maksym, et al.
Veröffentlicht: (2026)
von: Taranukhin, Maksym, et al.
Veröffentlicht: (2026)
BLAB: Brutally Long Audio Bench
von: Ahia, Orevaoghene, et al.
Veröffentlicht: (2025)
von: Ahia, Orevaoghene, et al.
Veröffentlicht: (2025)
ValueScope: Unveiling Implicit Norms and Values via Return Potential Model of Social Interactions
von: Park, Chan Young, et al.
Veröffentlicht: (2024)
von: Park, Chan Young, et al.
Veröffentlicht: (2024)
COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context
von: Wan, Guangya, et al.
Veröffentlicht: (2025)
von: Wan, Guangya, et al.
Veröffentlicht: (2025)
Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM Agents
von: Liu, Zhihan, et al.
Veröffentlicht: (2026)
von: Liu, Zhihan, et al.
Veröffentlicht: (2026)
Temporal Preferences in Language Models for Long-Horizon Assistance
von: Mazyaki, Ali, et al.
Veröffentlicht: (2025)
von: Mazyaki, Ali, et al.
Veröffentlicht: (2025)
AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications
von: Zhao, Yujie, et al.
Veröffentlicht: (2026)
von: Zhao, Yujie, et al.
Veröffentlicht: (2026)
EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning
von: Yu, Chengjun, et al.
Veröffentlicht: (2026)
von: Yu, Chengjun, et al.
Veröffentlicht: (2026)
reWordBench: Benchmarking and Improving the Robustness of Reward Models with Transformed Inputs
von: Wu, Zhaofeng, et al.
Veröffentlicht: (2025)
von: Wu, Zhaofeng, et al.
Veröffentlicht: (2025)
WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents
von: Liu, Junteng, et al.
Veröffentlicht: (2025)
von: Liu, Junteng, et al.
Veröffentlicht: (2025)
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
von: Ding, Shuangrui, et al.
Veröffentlicht: (2026)
von: Ding, Shuangrui, et al.
Veröffentlicht: (2026)
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
von: Zhao, Bingchen, et al.
Veröffentlicht: (2026)
von: Zhao, Bingchen, et al.
Veröffentlicht: (2026)
LifeBench: A Benchmark for Long-Horizon Multi-Source Memory
von: Cheng, Zihao, et al.
Veröffentlicht: (2026)
von: Cheng, Zihao, et al.
Veröffentlicht: (2026)
KellyBench: A Benchmark for Long-Horizon Sequential Decision Making
von: Grady, Thomas, et al.
Veröffentlicht: (2026)
von: Grady, Thomas, et al.
Veröffentlicht: (2026)
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
von: Song, Yuanyi, et al.
Veröffentlicht: (2025)
von: Song, Yuanyi, et al.
Veröffentlicht: (2025)
Adaptive Decoding via Latent Preference Optimization
von: Dhuliawala, Shehzaad, et al.
Veröffentlicht: (2024)
von: Dhuliawala, Shehzaad, et al.
Veröffentlicht: (2024)
Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks
von: Yang, Cheng, et al.
Veröffentlicht: (2025)
von: Yang, Cheng, et al.
Veröffentlicht: (2025)
Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks
von: Wu, Xiyang, et al.
Veröffentlicht: (2026)
von: Wu, Xiyang, et al.
Veröffentlicht: (2026)
LongBench: Evaluating Robotic Manipulation Policies on Real-World Long-Horizon Tasks
von: Chen, Xueyao, et al.
Veröffentlicht: (2026)
von: Chen, Xueyao, et al.
Veröffentlicht: (2026)
Precise Information Control in Long-Form Text Generation
von: He, Jacqueline, et al.
Veröffentlicht: (2025)
von: He, Jacqueline, et al.
Veröffentlicht: (2025)
Evolving Horizons: The Changing Generation of Legacy System
von: Rawal, Nitin
Veröffentlicht: (2023)
von: Rawal, Nitin
Veröffentlicht: (2023)
Evolving Horizons: The Changing Generation of Legacy System
von: Rawal, Nitin
Veröffentlicht: (2023)
von: Rawal, Nitin
Veröffentlicht: (2023)
MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning
von: Li, Shuyue Stella, et al.
Veröffentlicht: (2024)
von: Li, Shuyue Stella, et al.
Veröffentlicht: (2024)
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
SokoBench: Evaluating Long-Horizon Planning and Reasoning in Large Language Models
von: Monti, Sebastiano, et al.
Veröffentlicht: (2026)
von: Monti, Sebastiano, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
PrefPalette: Personalized Preference Modeling with Latent Attributes
von: Li, Shuyue Stella, et al.
Veröffentlicht: (2025) -
Cold-Start Personalization via Training-Free Priors from Structured World Models
von: Bose, Avinandan, et al.
Veröffentlicht: (2026) -
Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning
von: Sclar, Melanie, et al.
Veröffentlicht: (2024) -
ComPO: Community Preferences for Language Model Personalization
von: Kumar, Sachin, et al.
Veröffentlicht: (2024) -
Structured Preference Optimization for Vision-Language Long-Horizon Task Planning
von: Liang, Xiwen, et al.
Veröffentlicht: (2025)