CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hwang, Juyeong, Hong, Seong-Eun, Kim, Jinhyun, Seon, JaeYoung, Nam, Giljoo, Jang, Hanyoung, Kang, HyeongYeop
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914451389677568
author Hwang, Juyeong
Hong, Seong-Eun
Kim, Jinhyun
Seon, JaeYoung
Nam, Giljoo
Jang, Hanyoung
Kang, HyeongYeop
author_facet Hwang, Juyeong
Hong, Seong-Eun
Kim, Jinhyun
Seon, JaeYoung
Nam, Giljoo
Jang, Hanyoung
Kang, HyeongYeop
contents Crowds do not merely move; they decide. Human navigation is inherently contextual: people interpret the meaning of space, social norms, and potential consequences before acting. Sidewalks invite walking, crosswalks invite crossing, and deviations are weighed against urgency and safety. Yet most crowd simulation methods reduce navigation to geometry and collision avoidance, producing motion that is plausible but rarely intentional. We introduce CrowdVLA, a new formulation of crowd simulation that models each pedestrian as a Vision-Language-Action (VLA) agent. Instead of replaying recorded trajectories, CrowdVLA enables agents to interpret scene semantics and social norms from visual observations and language instructions, and to select actions through consequence-aware reasoning. CrowdVLA addresses three key challenges-limited agent-centric supervision in crowd datasets, unstable per-frame control, and success-biased datasets-through: (i) agent-centric visual supervision via semantically reconstructed environments and Low-Rank Adaptation (LoRA) fine-tuning of a pretrained vision-language model, (ii) a motion skill action space that bridges symbolic decision making and continuous locomotion, and (iii) exploration-based question answering that exposes agents to counterfactual actions and their outcomes through simulation rollouts. Our results shift crowd simulation from motion-centric synthesis toward perception-driven, consequence-aware decision making, enabling crowds that move not just realistically, but meaningfully.
format Preprint
id arxiv_https___arxiv_org_abs_2604_05525
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation
Hwang, Juyeong
Hong, Seong-Eun
Kim, Jinhyun
Seon, JaeYoung
Nam, Giljoo
Jang, Hanyoung
Kang, HyeongYeop
Graphics
Crowds do not merely move; they decide. Human navigation is inherently contextual: people interpret the meaning of space, social norms, and potential consequences before acting. Sidewalks invite walking, crosswalks invite crossing, and deviations are weighed against urgency and safety. Yet most crowd simulation methods reduce navigation to geometry and collision avoidance, producing motion that is plausible but rarely intentional. We introduce CrowdVLA, a new formulation of crowd simulation that models each pedestrian as a Vision-Language-Action (VLA) agent. Instead of replaying recorded trajectories, CrowdVLA enables agents to interpret scene semantics and social norms from visual observations and language instructions, and to select actions through consequence-aware reasoning. CrowdVLA addresses three key challenges-limited agent-centric supervision in crowd datasets, unstable per-frame control, and success-biased datasets-through: (i) agent-centric visual supervision via semantically reconstructed environments and Low-Rank Adaptation (LoRA) fine-tuning of a pretrained vision-language model, (ii) a motion skill action space that bridges symbolic decision making and continuous locomotion, and (iii) exploration-based question answering that exposes agents to counterfactual actions and their outcomes through simulation rollouts. Our results shift crowd simulation from motion-centric synthesis toward perception-driven, consequence-aware decision making, enabling crowds that move not just realistically, but meaningfully.
title CrowdVLA: Embodied Vision-Language-Action Agents for Context-Aware Crowd Simulation
topic Graphics
url https://arxiv.org/abs/2604.05525