Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Jinzhou, zhang, Jusheng, Liu, Sidi, Xiu, Waikit, Lv, Qinhan, Li, Xiying |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LASAR: Towards Spatio-temporal Reasoning with Latent Cognitive Map
by: Tang, Jinzhou, et al.
Published: (2026)
by: Tang, Jinzhou, et al.
Published: (2026)
HiVA: Self-organized Hierarchical Variable Agent via Goal-driven Semantic-Topological Evolution
by: Tang, Jinzhou, et al.
Published: (2025)
by: Tang, Jinzhou, et al.
Published: (2025)
Contrastive Learning-Driven Traffic Sign Perception: Multi-Modal Fusion of Text and Vision
by: Lu, Qiang, et al.
Published: (2025)
by: Lu, Qiang, et al.
Published: (2025)
Traffic-MLLM: Curiosity-Regularized Supervised Learning for Traffic Scenario Case-Based Reasoning
by: Xiu, Waikit, et al.
Published: (2025)
by: Xiu, Waikit, et al.
Published: (2025)
From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
DepthSSC: Monocular 3D Semantic Scene Completion via Depth-Spatial Alignment and Voxel Adaptation
by: Yao, Jiawei, et al.
Published: (2023)
by: Yao, Jiawei, et al.
Published: (2023)
ResAgent: Entropy-based Prior Point Discovery and Visual Reasoning for Referring Expression Segmentation
by: Wang, Yihao, et al.
Published: (2026)
by: Wang, Yihao, et al.
Published: (2026)
Beyond Pixels: Leveraging the Language of Soccer to Improve Spatio-Temporal Action Detection in Broadcast Videos
by: Ochin, Jeremie, et al.
Published: (2025)
by: Ochin, Jeremie, et al.
Published: (2025)
Task-Adapter++: Task-specific Adaptation with Order-aware Alignment for Few-shot Action Recognition
by: Cao, Congqi, et al.
Published: (2025)
by: Cao, Congqi, et al.
Published: (2025)
HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
Multimodal Spatio-temporal Graph Learning for Alignment-free RGBT Video Object Detection
by: Wang, Qishun, et al.
Published: (2025)
by: Wang, Qishun, et al.
Published: (2025)
PastNet: Introducing Physical Inductive Biases for Spatio-temporal Video Prediction
by: Wu, Hao, et al.
Published: (2023)
by: Wu, Hao, et al.
Published: (2023)
GeoVideo: Introducing Geometric Regularization into Video Generation Model
by: Bai, Yunpeng, et al.
Published: (2025)
by: Bai, Yunpeng, et al.
Published: (2025)
MoWM: Mixture-of-World-Models for Embodied Planning via Latent-to-Pixel Feature Modulation
by: Yu, Yangcheng, et al.
Published: (2025)
by: Yu, Yangcheng, et al.
Published: (2025)
Video-Language Alignment via Spatio-Temporal Graph Transformer
by: Zhang, Shi-Xue, et al.
Published: (2024)
by: Zhang, Shi-Xue, et al.
Published: (2024)
KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration
by: Zhang, Ruicheng, et al.
Published: (2026)
by: Zhang, Ruicheng, et al.
Published: (2026)
Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control
by: Han, Minghao, et al.
Published: (2025)
by: Han, Minghao, et al.
Published: (2025)
SGR-OCC: Evolving Monocular Priors for Embodied 3D Occupancy Prediction via Soft-Gating Lifting and Semantic-Adaptive Geometric Refinement
by: Guo, Yiran, et al.
Published: (2026)
by: Guo, Yiran, et al.
Published: (2026)
CoAgent: Collaborative Planning and Consistency Agent for Coherent Video Generation
by: Zeng, Qinglin, et al.
Published: (2025)
by: Zeng, Qinglin, et al.
Published: (2025)
Fast Omni-Directional Image Super-Resolution: Adapting the Implicit Image Function with Pixel and Semantic-Wise Spherical Geometric Priors
by: Shen, Xuelin, et al.
Published: (2025)
by: Shen, Xuelin, et al.
Published: (2025)
MM-CoT:A Benchmark for Probing Visual Chain-of-Thought Reasoning in Multimodal Models
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
SNOW: Spatio-Temporal Scene Understanding with World Knowledge for Open-World Embodied Reasoning
by: Sohn, Tin Stribor, et al.
Published: (2025)
by: Sohn, Tin Stribor, et al.
Published: (2025)
SeeClear: Semantic Distillation Enhances Pixel Condensation for Video Super-Resolution
by: Tang, Qi, et al.
Published: (2024)
by: Tang, Qi, et al.
Published: (2024)
Stable Single-Pixel Contrastive Learning for Semantic and Geometric Tasks
by: Pogorelyuk, Leonid, et al.
Published: (2025)
by: Pogorelyuk, Leonid, et al.
Published: (2025)
Mitigating Prior Shape Bias in Point Clouds via Differentiable Center Learning
by: Li, Zhe, et al.
Published: (2024)
by: Li, Zhe, et al.
Published: (2024)
Beyond Semantic Features: Pixel-level Mapping for Generalized AI-Generated Image Detection
by: Zhou, Chenming, et al.
Published: (2025)
by: Zhou, Chenming, et al.
Published: (2025)
Semantic Lens: Instance-Centric Semantic Alignment for Video Super-Resolution
by: Tang, Qi, et al.
Published: (2023)
by: Tang, Qi, et al.
Published: (2023)
Top-Down Semantic Refinement for Image Captioning
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
Wanderland: Geometrically Grounded Simulation for Open-World Embodied AI
by: Liu, Xinhao, et al.
Published: (2025)
by: Liu, Xinhao, et al.
Published: (2025)
SPARS3R: Semantic Prior Alignment and Regularization for Sparse 3D Reconstruction
by: Tang, Yutao, et al.
Published: (2024)
by: Tang, Yutao, et al.
Published: (2024)
STAF: 3D Human Mesh Recovery from Video with Spatio-Temporal Alignment Fusion
by: Yao, Wei, et al.
Published: (2024)
by: Yao, Wei, et al.
Published: (2024)
Blur-aware Spatio-temporal Sparse Transformer for Video Deblurring
by: Zhang, Huicong, et al.
Published: (2024)
by: Zhang, Huicong, et al.
Published: (2024)
STDiff: Spatio-temporal Diffusion for Continuous Stochastic Video Prediction
by: Ye, Xi, et al.
Published: (2023)
by: Ye, Xi, et al.
Published: (2023)
DreamSAC: Learning Hamiltonian World Models via Symmetry Exploration
by: Tang, Jinzhou, et al.
Published: (2026)
by: Tang, Jinzhou, et al.
Published: (2026)
DVFace: Spatio-Temporal Dual-Prior Diffusion for Video Face Restoration
by: Chen, Zheng, et al.
Published: (2026)
by: Chen, Zheng, et al.
Published: (2026)
Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction
by: Karypidis, Efstathios, et al.
Published: (2026)
by: Karypidis, Efstathios, et al.
Published: (2026)
Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models
by: Shen, Ruiqi, et al.
Published: (2025)
by: Shen, Ruiqi, et al.
Published: (2025)
Beyond Pixels: Semantic-aware Typographic Attack for Geo-Privacy Protection
by: Zhu, Jiayi, et al.
Published: (2025)
by: Zhu, Jiayi, et al.
Published: (2025)
Rethinking Video Generation Model for the Embodied World
by: Deng, Yufan, et al.
Published: (2026)
by: Deng, Yufan, et al.
Published: (2026)
Unleashing Semantic and Geometric Priors for 3D Scene Completion
by: Chen, Shiyuan, et al.
Published: (2025)
by: Chen, Shiyuan, et al.
Published: (2025)
Similar Items
-
LASAR: Towards Spatio-temporal Reasoning with Latent Cognitive Map
by: Tang, Jinzhou, et al.
Published: (2026) -
HiVA: Self-organized Hierarchical Variable Agent via Goal-driven Semantic-Topological Evolution
by: Tang, Jinzhou, et al.
Published: (2025) -
Contrastive Learning-Driven Traffic Sign Perception: Multi-Modal Fusion of Text and Vision
by: Lu, Qiang, et al.
Published: (2025) -
Traffic-MLLM: Curiosity-Regularized Supervised Learning for Traffic Scenario Case-Based Reasoning
by: Xiu, Waikit, et al.
Published: (2025) -
From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control
by: Zhang, Jusheng, et al.
Published: (2025)