VideoPCDNet: Video Parsing and Prediction with Phase Correlation Networks
Fuente:
arXiv
Salvato in:
| Autori principali: | Vicente, Noel José Rodrigues, Lehner, Enrique, Villar-Corrales, Angel, Nogga, Jan, Behnke, Sven |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning
di: Villar-Corrales, Angel, et al.
Pubblicazione: (2025)
di: Villar-Corrales, Angel, et al.
Pubblicazione: (2025)
TextOCVP: Object-Centric Video Prediction with Language Guidance
di: Villar-Corrales, Angel, et al.
Pubblicazione: (2025)
di: Villar-Corrales, Angel, et al.
Pubblicazione: (2025)
MCDS-VSS: Moving Camera Dynamic Scene Video Semantic Segmentation by Filtering with Self-Supervised Geometry and Motion
di: Villar-Corrales, Angel, et al.
Pubblicazione: (2024)
di: Villar-Corrales, Angel, et al.
Pubblicazione: (2024)
Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking
di: Pätzold, Bastian, et al.
Pubblicazione: (2025)
di: Pätzold, Bastian, et al.
Pubblicazione: (2025)
OC-SOP: Enhancing Vision-Based 3D Semantic Occupancy Prediction by Object-Centric Awareness
di: Cao, Helin, et al.
Pubblicazione: (2025)
di: Cao, Helin, et al.
Pubblicazione: (2025)
SWA-SOP: Spatially-aware Window Attention for Semantic Occupancy Prediction in Autonomous Driving
di: Cao, Helin, et al.
Pubblicazione: (2025)
di: Cao, Helin, et al.
Pubblicazione: (2025)
DiffSSC: Semantic LiDAR Scan Completion using Denoising Diffusion Probabilistic Models
di: Cao, Helin, et al.
Pubblicazione: (2024)
di: Cao, Helin, et al.
Pubblicazione: (2024)
SLCF-Net: Sequential LiDAR-Camera Fusion for Semantic Scene Completion using a 3D Recurrent U-Net
di: Cao, Helin, et al.
Pubblicazione: (2024)
di: Cao, Helin, et al.
Pubblicazione: (2024)
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
di: Zhou, Jinxing, et al.
Pubblicazione: (2024)
di: Zhou, Jinxing, et al.
Pubblicazione: (2024)
LIAM: Multimodal Transformer for Language Instructions, Images, Actions and Semantic Maps
di: Wang, Yihao, et al.
Pubblicazione: (2025)
di: Wang, Yihao, et al.
Pubblicazione: (2025)
Semantic Parsing of Colonoscopy Videos with Multi-Label Temporal Networks
di: Kelner, Ori, et al.
Pubblicazione: (2023)
di: Kelner, Ori, et al.
Pubblicazione: (2023)
TKN: Transformer-based Keypoint Prediction Network For Real-time Video Prediction
di: Li, Haoran, et al.
Pubblicazione: (2023)
di: Li, Haoran, et al.
Pubblicazione: (2023)
CausalVE: Face Video Privacy Encryption via Causal Video Prediction
di: Huang, Yubo, et al.
Pubblicazione: (2024)
di: Huang, Yubo, et al.
Pubblicazione: (2024)
VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
di: Ji, Longbin, et al.
Pubblicazione: (2026)
di: Ji, Longbin, et al.
Pubblicazione: (2026)
VideoSAGE: Video Summarization with Graph Representation Learning
di: Chaves, Jose M. Rojas, et al.
Pubblicazione: (2024)
di: Chaves, Jose M. Rojas, et al.
Pubblicazione: (2024)
Slot-MPC: Goal-Conditioned Model Predictive Control with Object-Centric Representations
di: Spieler, Jonathan, et al.
Pubblicazione: (2026)
di: Spieler, Jonathan, et al.
Pubblicazione: (2026)
LoViT: Long Video Transformer for Surgical Phase Recognition
di: Liu, Yang, et al.
Pubblicazione: (2023)
di: Liu, Yang, et al.
Pubblicazione: (2023)
Reangle-A-Video: 4D Video Generation as Video-to-Video Translation
di: Jeong, Hyeonho, et al.
Pubblicazione: (2025)
di: Jeong, Hyeonho, et al.
Pubblicazione: (2025)
Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing
di: Zhao, Pengcheng, et al.
Pubblicazione: (2024)
di: Zhao, Pengcheng, et al.
Pubblicazione: (2024)
Continuous Video Process: Modeling Videos as Continuous Multi-Dimensional Processes for Video Prediction
di: Shrivastava, Gaurav, et al.
Pubblicazione: (2024)
di: Shrivastava, Gaurav, et al.
Pubblicazione: (2024)
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing
di: Wang, Langyu, et al.
Pubblicazione: (2024)
di: Wang, Langyu, et al.
Pubblicazione: (2024)
SIAM: A Simple Alternating Mixer for Video Prediction
di: Zheng, Xin, et al.
Pubblicazione: (2023)
di: Zheng, Xin, et al.
Pubblicazione: (2023)
Latent Video Prediction Learns Better World Models
di: Alrasheed, Ali J, et al.
Pubblicazione: (2026)
di: Alrasheed, Ali J, et al.
Pubblicazione: (2026)
Video Representation Learning with Joint-Embedding Predictive Architectures
di: Drozdov, Katrina, et al.
Pubblicazione: (2024)
di: Drozdov, Katrina, et al.
Pubblicazione: (2024)
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
di: Wang, Langyu, et al.
Pubblicazione: (2025)
di: Wang, Langyu, et al.
Pubblicazione: (2025)
LiDAR-based Registration against Georeferenced Models for Globally Consistent Allocentric Maps
di: Quenzel, Jan, et al.
Pubblicazione: (2024)
di: Quenzel, Jan, et al.
Pubblicazione: (2024)
EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation
di: He, Ruozhen, et al.
Pubblicazione: (2026)
di: He, Ruozhen, et al.
Pubblicazione: (2026)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
di: Wang, Shihao, et al.
Pubblicazione: (2025)
di: Wang, Shihao, et al.
Pubblicazione: (2025)
HyenaPixel: Global Image Context with Convolutions
di: Spravil, Julian, et al.
Pubblicazione: (2024)
di: Spravil, Julian, et al.
Pubblicazione: (2024)
FSRT: Facial Scene Representation Transformer for Face Reenactment from Factorized Appearance, Head-pose, and Facial Expression Features
di: Rochow, Andre, et al.
Pubblicazione: (2024)
di: Rochow, Andre, et al.
Pubblicazione: (2024)
Video Panels for Long Video Understanding
di: Doorenbos, Lars, et al.
Pubblicazione: (2025)
di: Doorenbos, Lars, et al.
Pubblicazione: (2025)
Intelligent Parsing: An Automated Parsing Framework for Extracting Design Semantics from E-commerce Creatives
di: Li, Guandong, et al.
Pubblicazione: (2023)
di: Li, Guandong, et al.
Pubblicazione: (2023)
Explicit Abstention Knobs for Predictable Reliability in Video Question Answering
di: Ortiz, Jorge
Pubblicazione: (2025)
di: Ortiz, Jorge
Pubblicazione: (2025)
OCK: Unsupervised Dynamic Video Prediction with Object-Centric Kinematics
di: Song, Yeon-Ji, et al.
Pubblicazione: (2024)
di: Song, Yeon-Ji, et al.
Pubblicazione: (2024)
DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder
di: Chen, Junyu, et al.
Pubblicazione: (2025)
di: Chen, Junyu, et al.
Pubblicazione: (2025)
Object-Shot Enhanced Grounding Network for Egocentric Video
di: Feng, Yisen, et al.
Pubblicazione: (2025)
di: Feng, Yisen, et al.
Pubblicazione: (2025)
Space-time Reinforcement Network for Video Object Segmentation
di: Chen, Yadang, et al.
Pubblicazione: (2024)
di: Chen, Yadang, et al.
Pubblicazione: (2024)
A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects
di: Xie, Guohuan, et al.
Pubblicazione: (2025)
di: Xie, Guohuan, et al.
Pubblicazione: (2025)
Reinforced Label Denoising for Weakly-Supervised Audio-Visual Video Parsing
di: Gao, Yongbiao, et al.
Pubblicazione: (2024)
di: Gao, Yongbiao, et al.
Pubblicazione: (2024)
Video-Infinity: Distributed Long Video Generation
di: Tan, Zhenxiong, et al.
Pubblicazione: (2024)
di: Tan, Zhenxiong, et al.
Pubblicazione: (2024)
Documenti analoghi
-
PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning
di: Villar-Corrales, Angel, et al.
Pubblicazione: (2025) -
TextOCVP: Object-Centric Video Prediction with Language Guidance
di: Villar-Corrales, Angel, et al.
Pubblicazione: (2025) -
MCDS-VSS: Moving Camera Dynamic Scene Video Semantic Segmentation by Filtering with Self-Supervised Geometry and Motion
di: Villar-Corrales, Angel, et al.
Pubblicazione: (2024) -
Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking
di: Pätzold, Bastian, et al.
Pubblicazione: (2025) -
OC-SOP: Enhancing Vision-Based 3D Semantic Occupancy Prediction by Object-Centric Awareness
di: Cao, Helin, et al.
Pubblicazione: (2025)