Video Understanding: From Geometry and Semantics to Unified Models
Fuente:
arXiv
Salvato in:
| Autori principali: | An, Zhaochong, Li, Zirui, Ye, Mingqiao, Qiao, Feng, Li, Jiaang, Wu, Zongwei, Thengane, Vishal, Li, Chengzu, Li, Lei, Van Gool, Luc, Sun, Guolei, Belongie, Serge |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Rethinking Few-shot 3D Point Cloud Semantic Segmentation
di: An, Zhaochong, et al.
Pubblicazione: (2024)
di: An, Zhaochong, et al.
Pubblicazione: (2024)
ChatMotion: A Multimodal Multi-Agent for Human Motion Analysis
di: Li, Lei, et al.
Pubblicazione: (2025)
di: Li, Lei, et al.
Pubblicazione: (2025)
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
di: Li, Chengzu, et al.
Pubblicazione: (2026)
di: Li, Chengzu, et al.
Pubblicazione: (2026)
Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model
di: An, Zhaochong, et al.
Pubblicazione: (2025)
di: An, Zhaochong, et al.
Pubblicazione: (2025)
Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation
di: An, Zhaochong, et al.
Pubblicazione: (2024)
di: An, Zhaochong, et al.
Pubblicazione: (2024)
Learning Local and Global Temporal Contexts for Video Semantic Segmentation
di: Sun, Guolei, et al.
Pubblicazione: (2022)
di: Sun, Guolei, et al.
Pubblicazione: (2022)
What if Othello-Playing Language Models Could See?
di: Chen, Xinyi, et al.
Pubblicazione: (2025)
di: Chen, Xinyi, et al.
Pubblicazione: (2025)
Revisiting the Perception-Distortion Trade-off with Spatial-Semantic Guided Super-Resolution
di: Wang, Dan, et al.
Pubblicazione: (2026)
di: Wang, Dan, et al.
Pubblicazione: (2026)
Gradient Correlation Subspace Learning against Catastrophic Forgetting
di: Dubnov, Tammuz, et al.
Pubblicazione: (2024)
di: Dubnov, Tammuz, et al.
Pubblicazione: (2024)
RGB-D Indiscernible Object Counting in Underwater Scenes
di: Sun, Guolei, et al.
Pubblicazione: (2023)
di: Sun, Guolei, et al.
Pubblicazione: (2023)
What You Have is What You Track: Adaptive and Robust Multimodal Tracking
di: Tan, Yuedong, et al.
Pubblicazione: (2025)
di: Tan, Yuedong, et al.
Pubblicazione: (2025)
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
di: Li, Jiaang, et al.
Pubblicazione: (2025)
di: Li, Jiaang, et al.
Pubblicazione: (2025)
XTrack: Multimodal Training Boosts RGB-X Video Object Trackers
di: Tan, Yuedong, et al.
Pubblicazione: (2024)
di: Tan, Yuedong, et al.
Pubblicazione: (2024)
Self-Explainable Affordance Learning with Embodied Caption
di: Zhang, Zhipeng, et al.
Pubblicazione: (2024)
di: Zhang, Zhipeng, et al.
Pubblicazione: (2024)
Evaluation of Cultural Competence of Vision-Language Models
di: Yadav, Srishti, et al.
Pubblicazione: (2025)
di: Yadav, Srishti, et al.
Pubblicazione: (2025)
VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward
di: An, Zhaochong, et al.
Pubblicazione: (2026)
di: An, Zhaochong, et al.
Pubblicazione: (2026)
SCOPE: Scene-Contextualized Incremental Few-Shot 3D Segmentation
di: Thengane, Vishal, et al.
Pubblicazione: (2026)
di: Thengane, Vishal, et al.
Pubblicazione: (2026)
EgoGaussian: Dynamic Scene Understanding from Egocentric Video with 3D Gaussian Splatting
di: Zhang, Daiwei, et al.
Pubblicazione: (2024)
di: Zhang, Daiwei, et al.
Pubblicazione: (2024)
Test-time Training for Hyperspectral Image Super-resolution
di: Li, Ke, et al.
Pubblicazione: (2024)
di: Li, Ke, et al.
Pubblicazione: (2024)
Online Informative Sampling using Semantic Features in Underwater Environments
di: Thengane, Shrutika Vishal, et al.
Pubblicazione: (2024)
di: Thengane, Shrutika Vishal, et al.
Pubblicazione: (2024)
Vanishing-Point-Guided Video Semantic Segmentation of Driving Scenes
di: Guo, Diandian, et al.
Pubblicazione: (2024)
di: Guo, Diandian, et al.
Pubblicazione: (2024)
Optimizing against Infeasible Inclusions from Data for Semantic Segmentation through Morphology
di: Basu, Shamik, et al.
Pubblicazione: (2024)
di: Basu, Shamik, et al.
Pubblicazione: (2024)
Foundational Models for 3D Point Clouds: A Survey and Outlook
di: Thengane, Vishal, et al.
Pubblicazione: (2025)
di: Thengane, Vishal, et al.
Pubblicazione: (2025)
Investigating the Effectiveness of Cross-Attention to Unlock Zero-Shot Editing of Text-to-Video Diffusion Models
di: Motamed, Saman, et al.
Pubblicazione: (2024)
di: Motamed, Saman, et al.
Pubblicazione: (2024)
Towards Open-Vocabulary Video Semantic Segmentation
di: Li, Xinhao, et al.
Pubblicazione: (2024)
di: Li, Xinhao, et al.
Pubblicazione: (2024)
Rethinking Global Context in Crowd Counting
di: Sun, Guolei, et al.
Pubblicazione: (2021)
di: Sun, Guolei, et al.
Pubblicazione: (2021)
Vision Transformers with Hierarchical Attention
di: Liu, Yun, et al.
Pubblicazione: (2021)
di: Liu, Yun, et al.
Pubblicazione: (2021)
Condition-Invariant Semantic Segmentation
di: Sakaridis, Christos, et al.
Pubblicazione: (2023)
di: Sakaridis, Christos, et al.
Pubblicazione: (2023)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
di: Li, Wenyan, et al.
Pubblicazione: (2024)
di: Li, Wenyan, et al.
Pubblicazione: (2024)
SLAck: Semantic, Location, and Appearance Aware Open-Vocabulary Tracking
di: Li, Siyuan, et al.
Pubblicazione: (2024)
di: Li, Siyuan, et al.
Pubblicazione: (2024)
Noise-Coded Illumination for Forensic and Photometric Video Analysis
di: Michael, Peter F., et al.
Pubblicazione: (2025)
di: Michael, Peter F., et al.
Pubblicazione: (2025)
Stitched Value Model for Diffusion Alignment
di: Go, Hyojun, et al.
Pubblicazione: (2026)
di: Go, Hyojun, et al.
Pubblicazione: (2026)
ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives
di: Fu, Yuqian, et al.
Pubblicazione: (2024)
di: Fu, Yuqian, et al.
Pubblicazione: (2024)
Causality Model for Semantic Understanding on Videos
di: Yicong, Li
Pubblicazione: (2025)
di: Yicong, Li
Pubblicazione: (2025)
CAFuser: Condition-Aware Multimodal Fusion for Robust Semantic Perception of Driving Scenes
di: Broedermann, Tim, et al.
Pubblicazione: (2024)
di: Broedermann, Tim, et al.
Pubblicazione: (2024)
Sun Off, Lights On: Photorealistic Monocular Nighttime Simulation for Robust Semantic Perception
di: Tzevelekakis, Konstantinos, et al.
Pubblicazione: (2024)
di: Tzevelekakis, Konstantinos, et al.
Pubblicazione: (2024)
Autonomous Vehicle Controllers From End-to-End Differentiable Simulation
di: Nachkov, Asen, et al.
Pubblicazione: (2024)
di: Nachkov, Asen, et al.
Pubblicazione: (2024)
Single-Model and Any-Modality for Video Object Tracking
di: Wu, Zongwei, et al.
Pubblicazione: (2023)
di: Wu, Zongwei, et al.
Pubblicazione: (2023)
Towards Online Real-Time Memory-based Video Inpainting Transformers
di: Thiry, Guillaume, et al.
Pubblicazione: (2024)
di: Thiry, Guillaume, et al.
Pubblicazione: (2024)
MMEarth: Exploring Multi-Modal Pretext Tasks For Geospatial Representation Learning
di: Nedungadi, Vishal, et al.
Pubblicazione: (2024)
di: Nedungadi, Vishal, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Rethinking Few-shot 3D Point Cloud Semantic Segmentation
di: An, Zhaochong, et al.
Pubblicazione: (2024) -
ChatMotion: A Multimodal Multi-Agent for Human Motion Analysis
di: Li, Lei, et al.
Pubblicazione: (2025) -
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
di: Li, Chengzu, et al.
Pubblicazione: (2026) -
Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model
di: An, Zhaochong, et al.
Pubblicazione: (2025) -
Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation
di: An, Zhaochong, et al.
Pubblicazione: (2024)