To See or To Read: User Behavior Reasoning in Multimodal LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866912689894195200 |
|---|---|
| author | Dong, Tianning Ma, Luyi Vasudevan, Varun Cho, Jason Kumar, Sushant Achan, Kannan |
| author_facet | Dong, Tianning Ma, Luyi Vasudevan, Varun Cho, Jason Kumar, Sushant Achan, Kannan |
| contents | Multimodal Large Language Models (MLLMs) are reshaping how modern agentic systems reason over sequential user-behavior data. However, whether textual or image representations of user behavior data are more effective for maximizing MLLM performance remains underexplored. We present \texttt{BehaviorLens}, a systematic benchmarking framework for assessing modality trade-offs in user-behavior reasoning across six MLLMs by representing transaction data as (1) a text paragraph, (2) a scatter plot, and (3) a flowchart. Using a real-world purchase-sequence dataset, we find that when data is represented as images, MLLMs next-purchase prediction accuracy is improved by 87.5% compared with an equivalent textual representation without any additional computational cost. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_03845 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | To See or To Read: User Behavior Reasoning in Multimodal LLMs Dong, Tianning Ma, Luyi Vasudevan, Varun Cho, Jason Kumar, Sushant Achan, Kannan Artificial Intelligence Machine Learning Multimodal Large Language Models (MLLMs) are reshaping how modern agentic systems reason over sequential user-behavior data. However, whether textual or image representations of user behavior data are more effective for maximizing MLLM performance remains underexplored. We present \texttt{BehaviorLens}, a systematic benchmarking framework for assessing modality trade-offs in user-behavior reasoning across six MLLMs by representing transaction data as (1) a text paragraph, (2) a scatter plot, and (3) a flowchart. Using a real-world purchase-sequence dataset, we find that when data is represented as images, MLLMs next-purchase prediction accuracy is improved by 87.5% compared with an equivalent textual representation without any additional computational cost. |
| title | To See or To Read: User Behavior Reasoning in Multimodal LLMs |
| topic | Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2511.03845 |