SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Jianping, Xiao, Weiye, Lin, Zhengyu, Zhang, Huaizhong, Ren, Tianxiang, Gao, Yang, Lin, Zhiqian, Cai, Zhongang, Yang, Lei, Liu, Ziwei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniTalker: Scaling up Audio-Driven 3D Facial Animation through A Unified Model
by: Fan, Xiangyu, et al.
Published: (2024)
by: Fan, Xiangyu, et al.
Published: (2024)
Playing for 3D Human Recovery
by: Cai, Zhongang, et al.
Published: (2021)
by: Cai, Zhongang, et al.
Published: (2021)
OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction
by: Zhang, Haonan, et al.
Published: (2025)
by: Zhang, Haonan, et al.
Published: (2025)
Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection
by: Yang, Le, et al.
Published: (2024)
by: Yang, Le, et al.
Published: (2024)
WHAC: World-grounded Humans and Cameras
by: Yin, Wanqi, et al.
Published: (2024)
by: Yin, Wanqi, et al.
Published: (2024)
Bridging Semantic and Kinematic Conditions with Diffusion-based Discrete Motion Tokenizer
by: Gu, Chenyang, et al.
Published: (2026)
by: Gu, Chenyang, et al.
Published: (2026)
Phased DMD: Few-step Distribution Matching Distillation via Score Matching within Subintervals
by: Fan, Xiangyu, et al.
Published: (2025)
by: Fan, Xiangyu, et al.
Published: (2025)
Disco4D: Disentangled 4D Human Generation and Animation from a Single Image
by: Pang, Hui En, et al.
Published: (2024)
by: Pang, Hui En, et al.
Published: (2024)
TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models
by: Zhang, Zongzheng, et al.
Published: (2025)
by: Zhang, Zongzheng, et al.
Published: (2025)
GenLARP: Enabling Immersive Live Action Role-Play through LLM-Generated Worlds and Characters
by: Yu, Yichen, et al.
Published: (2025)
by: Yu, Yichen, et al.
Published: (2025)
Declaración SOLAMI. "Promoviendo el Orgullo de ser Internista"
by:
Published: (2019)
by:
Published: (2019)
Deblur-Avatar: Animatable Avatars from Motion-Blurred Monocular Videos
by: Luo, Xianrui, et al.
Published: (2025)
by: Luo, Xianrui, et al.
Published: (2025)
AttriHuman-3D: Editable 3D Human Avatar Generation with Attribute Decomposition and Indexing
by: Yang, Fan, et al.
Published: (2023)
by: Yang, Fan, et al.
Published: (2023)
A Survey on Vision-Language-Action Models for Autonomous Driving
by: Jiang, Sicong, et al.
Published: (2025)
by: Jiang, Sicong, et al.
Published: (2025)
Exploring Immersive Social-Physical Interaction with Virtual Characters through Coordinated Robotic Encountered-Type Contact
by: Godden, Eric, et al.
Published: (2025)
by: Godden, Eric, et al.
Published: (2025)
StyleVLA: Driving Style-Aware Vision Language Action Model for Autonomous Driving
by: Gao, Yuan, et al.
Published: (2026)
by: Gao, Yuan, et al.
Published: (2026)
Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations
by: Chen, Boxu, et al.
Published: (2025)
by: Chen, Boxu, et al.
Published: (2025)
Bridging Vision, Language, and Mathematics: Pictographic Character Reconstruction with Bézier Curves
by: Wan, Zihao, et al.
Published: (2025)
by: Wan, Zihao, et al.
Published: (2025)
CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
by: Du, Fan, et al.
Published: (2026)
by: Du, Fan, et al.
Published: (2026)
Learning Vision-Language-Action World Models for Autonomous Driving
by: Wang, Guoqing, et al.
Published: (2026)
by: Wang, Guoqing, et al.
Published: (2026)
Hierarchical Vision-Language Interaction for Facial Action Unit Detection
by: Li, Yong, et al.
Published: (2026)
by: Li, Yong, et al.
Published: (2026)
TINA: Think, Interaction, and Action Framework for Zero-Shot Vision Language Navigation
by: Li, Dingbang, et al.
Published: (2024)
by: Li, Dingbang, et al.
Published: (2024)
Imagine360: Immersive 360 Video Generation from Perspective Anchor
by: Tan, Jing, et al.
Published: (2024)
by: Tan, Jing, et al.
Published: (2024)
ConsistCompose: Unified Multimodal Layout Control for Image Composition
by: Shi, Xuanke, et al.
Published: (2025)
by: Shi, Xuanke, et al.
Published: (2025)
Judge, Then Drive: A Critic-Centric Vision Language Action Framework for Autonomous Driving
by: Yang, Lijin, et al.
Published: (2026)
by: Yang, Lijin, et al.
Published: (2026)
Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories
by: Yang, Rushuai, et al.
Published: (2025)
by: Yang, Rushuai, et al.
Published: (2025)
VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving
by: Zhao, Rui, et al.
Published: (2026)
by: Zhao, Rui, et al.
Published: (2026)
OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning
by: Lin, Fanqi, et al.
Published: (2025)
by: Lin, Fanqi, et al.
Published: (2025)
Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization
by: Huang, Jialei, et al.
Published: (2025)
by: Huang, Jialei, et al.
Published: (2025)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
by: Lin, Kevin Qinghong, et al.
Published: (2024)
by: Lin, Kevin Qinghong, et al.
Published: (2024)
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models
by: Ling, Yiran, et al.
Published: (2026)
by: Ling, Yiran, et al.
Published: (2026)
EndoVLA: Dual-Phase Vision-Language-Action Model for Autonomous Tracking in Endoscopy
by: Ng, Chi Kit, et al.
Published: (2025)
by: Ng, Chi Kit, et al.
Published: (2025)
Autonomous Character-Scene Interaction Synthesis from Text Instruction
by: Jiang, Nan, et al.
Published: (2024)
by: Jiang, Nan, et al.
Published: (2024)
APPLV: Adaptive Planner Parameter Learning from Vision-Language-Action Model
by: Lu, Yuanjie, et al.
Published: (2026)
by: Lu, Yuanjie, et al.
Published: (2026)
Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
by: Hu, Tianshuai, et al.
Published: (2025)
by: Hu, Tianshuai, et al.
Published: (2025)
ATA: Bridging Implicit Reasoning with Attention-Guided and Action-Guided Inference for Vision-Language Action Models
by: Yang, Cheng, et al.
Published: (2026)
by: Yang, Cheng, et al.
Published: (2026)
Beyond World-Frame Action Heads: Motion-Centric Action Frames for Vision-Language-Action Models
by: Yang, Huoren, et al.
Published: (2026)
by: Yang, Huoren, et al.
Published: (2026)
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting
by: Lin, Juyi, et al.
Published: (2025)
by: Lin, Juyi, et al.
Published: (2025)
DPoser-X: Diffusion Model as Robust 3D Whole-body Human Pose Prior
by: Lu, Junzhe, et al.
Published: (2025)
by: Lu, Junzhe, et al.
Published: (2025)
Pay Less Attention to Function Words for Free Robustness of Vision-Language Models
by: Tian, Qiwei, et al.
Published: (2025)
by: Tian, Qiwei, et al.
Published: (2025)
Similar Items
-
UniTalker: Scaling up Audio-Driven 3D Facial Animation through A Unified Model
by: Fan, Xiangyu, et al.
Published: (2024) -
Playing for 3D Human Recovery
by: Cai, Zhongang, et al.
Published: (2021) -
OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction
by: Zhang, Haonan, et al.
Published: (2025) -
Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection
by: Yang, Le, et al.
Published: (2024) -
WHAC: World-grounded Humans and Cameras
by: Yin, Wanqi, et al.
Published: (2024)