Infrastructure-Centric World Models: Bridging Temporal Depth and Spatial Breadth for Roadside Perception
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Meng, Siyuan, Ai, Chengbo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
Joint Learning of Depth, Pose, and Local Radiance Field for Large Scale Monocular 3D Reconstruction
von: Syed, Shahram Najam, et al.
Veröffentlicht: (2025)
von: Syed, Shahram Najam, et al.
Veröffentlicht: (2025)
SimWorld: A Unified Benchmark for Simulator-Conditioned Scene Generation via World Model
von: Li, Xinqing, et al.
Veröffentlicht: (2025)
von: Li, Xinqing, et al.
Veröffentlicht: (2025)
Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
von: Mehta, Vinit, et al.
Veröffentlicht: (2025)
von: Mehta, Vinit, et al.
Veröffentlicht: (2025)
Vi-SAFE: A Spatial-Temporal Framework for Efficient Violence Detection in Public Surveillance
von: Chang, Ligang, et al.
Veröffentlicht: (2025)
von: Chang, Ligang, et al.
Veröffentlicht: (2025)
Temporally Consistent Object 6D Pose Estimation for Robot Control
von: Zorina, Kateryna, et al.
Veröffentlicht: (2026)
von: Zorina, Kateryna, et al.
Veröffentlicht: (2026)
Light Future: Multimodal Action Frame Prediction via InstructPix2Pix
von: Zhong, Zesen, et al.
Veröffentlicht: (2025)
von: Zhong, Zesen, et al.
Veröffentlicht: (2025)
SCA-Net: Spatial-Contextual Aggregation Network for Enhanced Small Building and Road Change Detection
von: Gholibeigi, Emad, et al.
Veröffentlicht: (2026)
von: Gholibeigi, Emad, et al.
Veröffentlicht: (2026)
Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens
von: Shen, Meng, et al.
Veröffentlicht: (2026)
von: Shen, Meng, et al.
Veröffentlicht: (2026)
EgoTraj: Real-World Egocentric Human Trajectory Dataset for Multimodal Prediction
von: Yehia, Ahmad, et al.
Veröffentlicht: (2026)
von: Yehia, Ahmad, et al.
Veröffentlicht: (2026)
Neuromorphic Monocular Depth Estimation with Uncertainty Modeling
von: Bergkvist, Viktor, et al.
Veröffentlicht: (2026)
von: Bergkvist, Viktor, et al.
Veröffentlicht: (2026)
Unveiling the Potential of iMarkers: Invisible Fiducial Markers for Advanced Robotics
von: Tourani, Ali, et al.
Veröffentlicht: (2025)
von: Tourani, Ali, et al.
Veröffentlicht: (2025)
SuperPoint-SLAM3: Augmenting ORB-SLAM3 with Deep Features, Adaptive NMS, and Learning-Based Loop Closure
von: Syed, Shahram Najam, et al.
Veröffentlicht: (2025)
von: Syed, Shahram Najam, et al.
Veröffentlicht: (2025)
vS-Graphs: Tightly Coupling Visual SLAM and 3D Scene Graphs Exploiting Hierarchical Scene Understanding
von: Tourani, Ali, et al.
Veröffentlicht: (2025)
von: Tourani, Ali, et al.
Veröffentlicht: (2025)
A Vision-Language Model for Focal Liver Lesion Classification
von: Jian, Song, et al.
Veröffentlicht: (2025)
von: Jian, Song, et al.
Veröffentlicht: (2025)
DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception
von: Deng, Pei, et al.
Veröffentlicht: (2025)
von: Deng, Pei, et al.
Veröffentlicht: (2025)
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
von: Lee, Kyuho, et al.
Veröffentlicht: (2025)
von: Lee, Kyuho, et al.
Veröffentlicht: (2025)
Selection, Not Fusion: Radar-Modulated State Space Models for Radar-Camera Depth Estimation
von: Hou, Zhangcheng, et al.
Veröffentlicht: (2026)
von: Hou, Zhangcheng, et al.
Veröffentlicht: (2026)
OCC-MLLM-CoT-Alpha: Towards Multi-stage Occlusion Recognition Based on Large Language Models via 3D-Aware Supervision and Chain-of-Thoughts Guidance
von: Wang, Chaoyi, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyi, et al.
Veröffentlicht: (2025)
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation
von: Xiao, Jiasong, et al.
Veröffentlicht: (2026)
von: Xiao, Jiasong, et al.
Veröffentlicht: (2026)
Single-Shot Metric Depth from Focused Plenoptic Cameras
von: Lasheras-Hernandez, Blanca, et al.
Veröffentlicht: (2024)
von: Lasheras-Hernandez, Blanca, et al.
Veröffentlicht: (2024)
Context in object detection: a systematic literature review
von: Jamali, Mahtab, et al.
Veröffentlicht: (2025)
von: Jamali, Mahtab, et al.
Veröffentlicht: (2025)
Mask-Conditioned Voxel Diffusion for Joint Geometry and Color Inpainting
von: Sumuk, Aarya
Veröffentlicht: (2026)
von: Sumuk, Aarya
Veröffentlicht: (2026)
Pedestrian Detection in Low-Light Conditions: A Comprehensive Survey
von: Ghari, Bahareh, et al.
Veröffentlicht: (2024)
von: Ghari, Bahareh, et al.
Veröffentlicht: (2024)
PhysVideoGenerator: Towards Physically Aware Video Generation via Latent Physics Guidance
von: Satish, Siddarth Nilol Kundur, et al.
Veröffentlicht: (2026)
von: Satish, Siddarth Nilol Kundur, et al.
Veröffentlicht: (2026)
FlowIBR: Leveraging Pre-Training for Efficient Neural Image-Based Rendering of Dynamic Scenes
von: Büsching, Marcel, et al.
Veröffentlicht: (2023)
von: Büsching, Marcel, et al.
Veröffentlicht: (2023)
IMASHRIMP: Automatic White Shrimp (Penaeus vannamei) Biometrical Analysis from Laboratory Images Using Computer Vision and Deep Learning
von: González, Abiam Remache, et al.
Veröffentlicht: (2025)
von: González, Abiam Remache, et al.
Veröffentlicht: (2025)
Towards a Generalizable Fusion Architecture for Multimodal Object Detection
von: Berjawi, Jad, et al.
Veröffentlicht: (2025)
von: Berjawi, Jad, et al.
Veröffentlicht: (2025)
A Simple Baseline for Streaming Video Understanding
von: Shen, Yujiao, et al.
Veröffentlicht: (2026)
von: Shen, Yujiao, et al.
Veröffentlicht: (2026)
EmoVerse: A MLLMs-Driven Emotion Representation Dataset for Interpretable Visual Emotion Analysis
von: Guo, Yijie, et al.
Veröffentlicht: (2025)
von: Guo, Yijie, et al.
Veröffentlicht: (2025)
Synthetic-Child: An AIGC-Based Synthetic Data Pipeline for Privacy-Preserving Child Posture Estimation
von: Zeng, Taowen
Veröffentlicht: (2026)
von: Zeng, Taowen
Veröffentlicht: (2026)
Exploring Surround-View Fisheye Camera 3D Object Detection
von: Li, Changcai, et al.
Veröffentlicht: (2025)
von: Li, Changcai, et al.
Veröffentlicht: (2025)
Car Object Counting and Position Estimation via Extension of the CLIP-EBC Framework
von: Jung, Seoik, et al.
Veröffentlicht: (2025)
von: Jung, Seoik, et al.
Veröffentlicht: (2025)
Pixel-Level Pavement Distress Assessment Using Instance Segmentation
von: Dewick, Logan, et al.
Veröffentlicht: (2026)
von: Dewick, Logan, et al.
Veröffentlicht: (2026)
A Reverse Causal Framework to Mitigate Spurious Correlations for Debiasing Scene Graph Generation
von: Sun, Shuzhou, et al.
Veröffentlicht: (2025)
von: Sun, Shuzhou, et al.
Veröffentlicht: (2025)
Action Anticipation from SoccerNet Football Video Broadcasts
von: Dalal, Mohamad, et al.
Veröffentlicht: (2025)
von: Dalal, Mohamad, et al.
Veröffentlicht: (2025)
A Recipe for Geometry-Aware 3D Mesh Transformers
von: Farazi, Mohammad, et al.
Veröffentlicht: (2024)
von: Farazi, Mohammad, et al.
Veröffentlicht: (2024)
Systematic Comparison of Projection Methods for Monocular 3D Human Pose Estimation on Fisheye Images
von: Käs, Stephanie, et al.
Veröffentlicht: (2025)
von: Käs, Stephanie, et al.
Veröffentlicht: (2025)
Is Single-View Mesh Reconstruction Ready for Robotics?
von: Nolte, Frederik, et al.
Veröffentlicht: (2025)
von: Nolte, Frederik, et al.
Veröffentlicht: (2025)
VDPP: Video Depth Post-Processing for Speed and Scalability
von: Yoon, Daewon, et al.
Veröffentlicht: (2026)
von: Yoon, Daewon, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025) -
Joint Learning of Depth, Pose, and Local Radiance Field for Large Scale Monocular 3D Reconstruction
von: Syed, Shahram Najam, et al.
Veröffentlicht: (2025) -
SimWorld: A Unified Benchmark for Simulator-Conditioned Scene Generation via World Model
von: Li, Xinqing, et al.
Veröffentlicht: (2025) -
Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
von: Mehta, Vinit, et al.
Veröffentlicht: (2025) -
Vi-SAFE: A Spatial-Temporal Framework for Efficient Violence Detection in Public Surveillance
von: Chang, Ligang, et al.
Veröffentlicht: (2025)