VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yuhao, Xu, Yi, Ding, Xinyun, Fang, Xiang, Liu, Shuochen, Lin, Luxi, Zhang, Qingyu, Li, Ya, Liu, Quan, Xu, Tong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918407595622400
author Chen, Yuhao
Xu, Yi
Ding, Xinyun
Fang, Xiang
Liu, Shuochen
Lin, Luxi
Zhang, Qingyu
Li, Ya
Liu, Quan
Xu, Tong
author_facet Chen, Yuhao
Xu, Yi
Ding, Xinyun
Fang, Xiang
Liu, Shuochen
Lin, Luxi
Zhang, Qingyu
Li, Ya
Liu, Quan
Xu, Tong
contents With the growing demand for intelligent in-vehicle experiences, vehicle-based agents are evolving from simple assistants to long-term companions. This evolution requires agents to continuously model multi-user preferences and make reliable decisions in the face of inter-user preference conflicts and changing habits over time. However, existing benchmarks are largely limited to single-user, static question-answer settings, failing to capture the temporal evolution of preferences and the multi-user, tool-interactive nature of real vehicle environments. To address this gap, we introduce VehicleMemBench, a multi-user long-context memory benchmark built on an executable in-vehicle simulation environment. The benchmark evaluates tool use and memory by comparing the post-action environment state with a predefined target state, enabling objective and reproducible evaluation without LLM-based or human scoring. VehicleMemBench includes 23 tool modules, and each sample contains over 80 historical memory events. Experiments show that powerful models perform well on direct instruction tasks but struggle in scenarios involving memory evolution, particularly when user preferences change dynamically. Even advanced memory systems struggle to handle domain-specific memory requirements in this environment. These findings highlight the need for more robust and specialized memory management mechanisms to support long-term adaptive decision-making in real-world in-vehicle systems. To facilitate future research, we release the data and code.
format Preprint
id arxiv_https___arxiv_org_abs_2603_23840
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents
Chen, Yuhao
Xu, Yi
Ding, Xinyun
Fang, Xiang
Liu, Shuochen
Lin, Luxi
Zhang, Qingyu
Li, Ya
Liu, Quan
Xu, Tong
Artificial Intelligence
Computation and Language
With the growing demand for intelligent in-vehicle experiences, vehicle-based agents are evolving from simple assistants to long-term companions. This evolution requires agents to continuously model multi-user preferences and make reliable decisions in the face of inter-user preference conflicts and changing habits over time. However, existing benchmarks are largely limited to single-user, static question-answer settings, failing to capture the temporal evolution of preferences and the multi-user, tool-interactive nature of real vehicle environments. To address this gap, we introduce VehicleMemBench, a multi-user long-context memory benchmark built on an executable in-vehicle simulation environment. The benchmark evaluates tool use and memory by comparing the post-action environment state with a predefined target state, enabling objective and reproducible evaluation without LLM-based or human scoring. VehicleMemBench includes 23 tool modules, and each sample contains over 80 historical memory events. Experiments show that powerful models perform well on direct instruction tasks but struggle in scenarios involving memory evolution, particularly when user preferences change dynamically. Even advanced memory systems struggle to handle domain-specific memory requirements in this environment. These findings highlight the need for more robust and specialized memory management mechanisms to support long-term adaptive decision-making in real-world in-vehicle systems. To facilitate future research, we release the data and code.
title VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2603.23840