Vision-language models lag human performance on physical dynamics and intent reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Gu, Tianjun, Gong, Jingyu, Zhang, Zhizhong, Xie, Yuan, Ma, Lizhuang, Tan, Xin, V, Athanasios |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GSCompleter: A Distillation-Free Plugin for Metric-Aware 3D Gaussian Splatting Completion in Seconds
by: Gao, Ao, et al.
Published: (2026)
by: Gao, Ao, et al.
Published: (2026)
S2GS: Streaming Semantic Gaussian Splatting for Online Scene Understanding and Reconstruction
by: Zhang, Renhe, et al.
Published: (2026)
by: Zhang, Renhe, et al.
Published: (2026)
COTR: Compact Occupancy TRansformer for Vision-based 3D Occupancy Prediction
by: Ma, Qihang, et al.
Published: (2023)
by: Ma, Qihang, et al.
Published: (2023)
UniForward: Unified 3D Scene and Semantic Field Reconstruction via Feed-Forward Gaussian Splatting from Only Sparse-View Images
by: Tian, Qijian, et al.
Published: (2025)
by: Tian, Qijian, et al.
Published: (2025)
From Enhancement to Understanding: Build a Generalized Bridge for Low-light Vision via Semantically Consistent Unsupervised Fine-tuning
by: Wang, Sen, et al.
Published: (2025)
by: Wang, Sen, et al.
Published: (2025)
DEMOS: Dynamic Environment Motion Synthesis in 3D Scenes via Local Spherical-BEV Perception
by: Gong, Jingyu, et al.
Published: (2024)
by: Gong, Jingyu, et al.
Published: (2024)
PFDepth: Heterogeneous Pinhole-Fisheye Joint Depth Estimation via Distortion-aware Gaussian-Splatted Volumetric Fusion
by: Zhang, Zhiwei, et al.
Published: (2025)
by: Zhang, Zhiwei, et al.
Published: (2025)
PromptAD: Learning Prompts with only Normal Samples for Few-Shot Anomaly Detection
by: Li, Xiaofan, et al.
Published: (2024)
by: Li, Xiaofan, et al.
Published: (2024)
DrivingForward: Feed-forward 3D Gaussian Splatting for Driving Scene Reconstruction from Flexible Surround-view Input
by: Tian, Qijian, et al.
Published: (2024)
by: Tian, Qijian, et al.
Published: (2024)
YouTube-Occ: Learning Indoor 3D Semantic Occupancy Prediction from YouTube Videos
by: Chen, Haoming, et al.
Published: (2025)
by: Chen, Haoming, et al.
Published: (2025)
GEOcc: Geometrically Enhanced 3D Occupancy Network with Implicit-Explicit Depth Fusion and Contextual Self-Supervision
by: Tan, Xin, et al.
Published: (2024)
by: Tan, Xin, et al.
Published: (2024)
Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy
by: Jingyu, Gong, et al.
Published: (2025)
by: Jingyu, Gong, et al.
Published: (2025)
Diffusion Implicit Policy for Unpaired Scene-aware Motion Synthesis
by: Gong, Jingyu, et al.
Published: (2024)
by: Gong, Jingyu, et al.
Published: (2024)
PIG: Prompt Images Guidance for Night-Time Scene Parsing
by: Xie, Zhifeng, et al.
Published: (2024)
by: Xie, Zhifeng, et al.
Published: (2024)
FLEG: Feed-Forward Language Embedded Gaussian Splatting from Any Views via Compact Semantic Representation
by: Tian, Qijian, et al.
Published: (2025)
by: Tian, Qijian, et al.
Published: (2025)
DORAEMON: Decentralized Ontology-aware Reliable Agent with Enhanced Memory Oriented Navigation
by: Gu, Tianjun, et al.
Published: (2025)
by: Gu, Tianjun, et al.
Published: (2025)
Textual Decomposition Then Sub-motion-space Scattering for Open-Vocabulary Motion Generation
by: Fan, Ke, et al.
Published: (2024)
by: Fan, Ke, et al.
Published: (2024)
One-for-More: Continual Diffusion Model for Anomaly Detection
by: Li, Xiaofan, et al.
Published: (2025)
by: Li, Xiaofan, et al.
Published: (2025)
Beyond the Label Itself: Latent Labels Enhance Semi-supervised Point Cloud Panoptic Segmentation
by: Chen, Yujun, et al.
Published: (2023)
by: Chen, Yujun, et al.
Published: (2023)
Exploring the Untouched Sweeps for Conflict-Aware 3D Segmentation Pretraining
by: Sun, Tianfang, et al.
Published: (2024)
by: Sun, Tianfang, et al.
Published: (2024)
World2Minecraft: Occupancy-Driven Simulated Scenes Construction
by: Zhang, Lechao, et al.
Published: (2026)
by: Zhang, Lechao, et al.
Published: (2026)
Omni-Supervised Motion Editing: Balancing Change and Invariance through Positive-Negative Learning
by: Shi, Zhenwu, et al.
Published: (2026)
by: Shi, Zhenwu, et al.
Published: (2026)
Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration
by: Wang, Sen, et al.
Published: (2026)
by: Wang, Sen, et al.
Published: (2026)
S2D: Sparse to Dense Lifting for 3D Reconstruction with Minimal Inputs
by: Ji, Yuzhou, et al.
Published: (2026)
by: Ji, Yuzhou, et al.
Published: (2026)
Building a Strong Pre-Training Baseline for Universal 3D Large-Scale Perception
by: Chen, Haoming, et al.
Published: (2024)
by: Chen, Haoming, et al.
Published: (2024)
Mutual Information Guided Optimal Transport for Unsupervised Visible-Infrared Person Re-identification
by: Zhang, Zhizhong, et al.
Published: (2024)
by: Zhang, Zhizhong, et al.
Published: (2024)
LidarPainter: One-Step Away From Any Lidar View To Novel Guidance
by: Ji, Yuzhou, et al.
Published: (2025)
by: Ji, Yuzhou, et al.
Published: (2025)
Emphasizing Semantic Consistency of Salient Posture for Speech-Driven Gesture Generation
by: Liu, Fengqi, et al.
Published: (2024)
by: Liu, Fengqi, et al.
Published: (2024)
Teaching large language models to reason like expert diagnosticians
by: Buckley, Thomas A., et al.
Published: (2025)
by: Buckley, Thomas A., et al.
Published: (2025)
Look, Remember and Reason: Grounded reasoning in videos with language models
by: Bhattacharyya, Apratim, et al.
Published: (2023)
by: Bhattacharyya, Apratim, et al.
Published: (2023)
FastLGS: Speeding up Language Embedded Gaussians with Feature Grid Mapping
by: Ji, Yuzhou, et al.
Published: (2024)
by: Ji, Yuzhou, et al.
Published: (2024)
DailyArt: Discovering Articulation from Single Static Images via Latent Dynamics
by: Zhang, Hang, et al.
Published: (2026)
by: Zhang, Hang, et al.
Published: (2026)
Continuous Piecewise-Affine Based Motion Model for Image Animation
by: Wang, Hexiang, et al.
Published: (2024)
by: Wang, Hexiang, et al.
Published: (2024)
CurvNet: Latent Contour Representation and Iterative Data Engine for Curvature Angle Estimation
by: Shao, Zhiwen, et al.
Published: (2024)
by: Shao, Zhiwen, et al.
Published: (2024)
MOS: Modeling Object-Scene Associations in Generalized Category Discovery
by: Peng, Zhengyuan, et al.
Published: (2025)
by: Peng, Zhengyuan, et al.
Published: (2025)
AUFormer: Vision Transformers are Parameter-Efficient Facial Action Unit Detectors
by: Yuan, Kaishen, et al.
Published: (2024)
by: Yuan, Kaishen, et al.
Published: (2024)
Efficient Multimodal Large Language Models: A Survey
by: Jin, Yizhang, et al.
Published: (2024)
by: Jin, Yizhang, et al.
Published: (2024)
Reconstructing In-the-Wild Open-Vocabulary Human-Object Interactions
by: Wen, Boran, et al.
Published: (2025)
by: Wen, Boran, et al.
Published: (2025)
MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
by: Xing, Yang, et al.
Published: (2026)
by: Xing, Yang, et al.
Published: (2026)
StyleRWKV: High-Quality and High-Efficiency Style Transfer with RWKV-like Architecture
by: Dai, Miaomiao, et al.
Published: (2024)
by: Dai, Miaomiao, et al.
Published: (2024)
Similar Items
-
GSCompleter: A Distillation-Free Plugin for Metric-Aware 3D Gaussian Splatting Completion in Seconds
by: Gao, Ao, et al.
Published: (2026) -
S2GS: Streaming Semantic Gaussian Splatting for Online Scene Understanding and Reconstruction
by: Zhang, Renhe, et al.
Published: (2026) -
COTR: Compact Occupancy TRansformer for Vision-based 3D Occupancy Prediction
by: Ma, Qihang, et al.
Published: (2023) -
UniForward: Unified 3D Scene and Semantic Field Reconstruction via Feed-Forward Gaussian Splatting from Only Sparse-View Images
by: Tian, Qijian, et al.
Published: (2025) -
From Enhancement to Understanding: Build a Generalized Bridge for Low-light Vision via Semantically Consistent Unsupervised Fine-tuning
by: Wang, Sen, et al.
Published: (2025)