To See or To Read: User Behavior Reasoning in Multimodal LLMs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Dong, Tianning, Ma, Luyi, Vasudevan, Varun, Cho, Jason, Kumar, Sushant, Achan, Kannan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912689894195200
author Dong, Tianning
Ma, Luyi
Vasudevan, Varun
Cho, Jason
Kumar, Sushant
Achan, Kannan
author_facet Dong, Tianning
Ma, Luyi
Vasudevan, Varun
Cho, Jason
Kumar, Sushant
Achan, Kannan
contents Multimodal Large Language Models (MLLMs) are reshaping how modern agentic systems reason over sequential user-behavior data. However, whether textual or image representations of user behavior data are more effective for maximizing MLLM performance remains underexplored. We present \texttt{BehaviorLens}, a systematic benchmarking framework for assessing modality trade-offs in user-behavior reasoning across six MLLMs by representing transaction data as (1) a text paragraph, (2) a scatter plot, and (3) a flowchart. Using a real-world purchase-sequence dataset, we find that when data is represented as images, MLLMs next-purchase prediction accuracy is improved by 87.5% compared with an equivalent textual representation without any additional computational cost.
format Preprint
id arxiv_https___arxiv_org_abs_2511_03845
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle To See or To Read: User Behavior Reasoning in Multimodal LLMs
Dong, Tianning
Ma, Luyi
Vasudevan, Varun
Cho, Jason
Kumar, Sushant
Achan, Kannan
Artificial Intelligence
Machine Learning
Multimodal Large Language Models (MLLMs) are reshaping how modern agentic systems reason over sequential user-behavior data. However, whether textual or image representations of user behavior data are more effective for maximizing MLLM performance remains underexplored. We present \texttt{BehaviorLens}, a systematic benchmarking framework for assessing modality trade-offs in user-behavior reasoning across six MLLMs by representing transaction data as (1) a text paragraph, (2) a scatter plot, and (3) a flowchart. Using a real-world purchase-sequence dataset, we find that when data is represented as images, MLLMs next-purchase prediction accuracy is improved by 87.5% compared with an equivalent textual representation without any additional computational cost.
title To See or To Read: User Behavior Reasoning in Multimodal LLMs
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2511.03845