Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction
Fuente:
arXiv
Salvato in:
| Autori principali: | Karypidis, Efstathios, Gidaris, Spyros, Komodakis, Nikos |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Advancing Semantic Future Prediction through Multimodal Visual Sequence Transformers
di: Karypidis, Efstathios, et al.
Pubblicazione: (2025)
di: Karypidis, Efstathios, et al.
Pubblicazione: (2025)
DINO-Foresight: Looking into the Future with DINO
di: Karypidis, Efstathios, et al.
Pubblicazione: (2024)
di: Karypidis, Efstathios, et al.
Pubblicazione: (2024)
Boosting Generative Image Modeling via Joint Image-Feature Synthesis
di: Kouzelis, Theodoros, et al.
Pubblicazione: (2025)
di: Kouzelis, Theodoros, et al.
Pubblicazione: (2025)
Coevolving Representations in Joint Image-Feature Diffusion
di: Kouzelis, Theodoros, et al.
Pubblicazione: (2026)
di: Kouzelis, Theodoros, et al.
Pubblicazione: (2026)
Multi-Token Prediction Needs Registers
di: Gerontopoulos, Anastasios, et al.
Pubblicazione: (2025)
di: Gerontopoulos, Anastasios, et al.
Pubblicazione: (2025)
SPOT: Self-Training with Patch-Order Permutation for Object-Centric Learning with Autoregressive Transformers
di: Kakogeorgiou, Ioannis, et al.
Pubblicazione: (2023)
di: Kakogeorgiou, Ioannis, et al.
Pubblicazione: (2023)
MOCA: Self-supervised Representation Learning by Predicting Masked Online Codebook Assignments
di: Gidaris, Spyros, et al.
Pubblicazione: (2023)
di: Gidaris, Spyros, et al.
Pubblicazione: (2023)
Technical Report for the 5th CLVision Challenge at CVPR: Addressing the Class-Incremental with Repetition using Unlabeled Data -- 4th Place Solution
di: Moraiti, Panagiota, et al.
Pubblicazione: (2025)
di: Moraiti, Panagiota, et al.
Pubblicazione: (2025)
DIP: Unsupervised Dense In-Context Post-training of Visual Representations
di: Sirko-Galouchenko, Sophia, et al.
Pubblicazione: (2025)
di: Sirko-Galouchenko, Sophia, et al.
Pubblicazione: (2025)
Vanilla ViT for Automotive Point Cloud Semantic Segmentation
di: Puy, Gilles, et al.
Pubblicazione: (2026)
di: Puy, Gilles, et al.
Pubblicazione: (2026)
No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations
di: Simoncini, Walter, et al.
Pubblicazione: (2024)
di: Simoncini, Walter, et al.
Pubblicazione: (2024)
Drive&Segment: Unsupervised Semantic Segmentation of Urban Scenes via Cross-modal Distillation
di: Vobecky, Antonin, et al.
Pubblicazione: (2022)
di: Vobecky, Antonin, et al.
Pubblicazione: (2022)
Unsupervised Object Localization in the Era of Self-Supervised ViTs: A Survey
di: Siméoni, Oriane, et al.
Pubblicazione: (2023)
di: Siméoni, Oriane, et al.
Pubblicazione: (2023)
POP-3D: Open-Vocabulary 3D Occupancy Prediction from Images
di: Vobecky, Antonin, et al.
Pubblicazione: (2024)
di: Vobecky, Antonin, et al.
Pubblicazione: (2024)
Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation?
di: Aravanis, Tilemachos, et al.
Pubblicazione: (2026)
di: Aravanis, Tilemachos, et al.
Pubblicazione: (2026)
Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
di: Venkataramanan, Shashanka, et al.
Pubblicazione: (2025)
di: Venkataramanan, Shashanka, et al.
Pubblicazione: (2025)
Boosting Visual Instruction Tuning with Self-Supervised Guidance
di: Sirko-Galouchenko, Sophia, et al.
Pubblicazione: (2026)
di: Sirko-Galouchenko, Sophia, et al.
Pubblicazione: (2026)
OccFeat: Self-supervised Occupancy Feature Prediction for Pretraining BEV Segmentation Networks
di: Sirko-Galouchenko, Sophia, et al.
Pubblicazione: (2024)
di: Sirko-Galouchenko, Sophia, et al.
Pubblicazione: (2024)
ToNNO: Tomographic Reconstruction of a Neural Network's Output for Weakly Supervised Segmentation of 3D Medical Images
di: Schmidt-Mengin, Marius, et al.
Pubblicazione: (2024)
di: Schmidt-Mengin, Marius, et al.
Pubblicazione: (2024)
Set Prediction Guided by Semantic Concepts for Diverse Video Captioning
di: Lu, Yifan, et al.
Pubblicazione: (2023)
di: Lu, Yifan, et al.
Pubblicazione: (2023)
Exploiting Semantic and Pixel Representations for Ultra-Low Bitrate Image Compression
di: Wei, Hao, et al.
Pubblicazione: (2026)
di: Wei, Hao, et al.
Pubblicazione: (2026)
The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
di: Fan, Weichen, et al.
Pubblicazione: (2025)
di: Fan, Weichen, et al.
Pubblicazione: (2025)
Three Pillars improving Vision Foundation Model Distillation for Lidar
di: Puy, Gilles, et al.
Pubblicazione: (2023)
di: Puy, Gilles, et al.
Pubblicazione: (2023)
Hierarchical Neural Semantic Representation for 3D Semantic Correspondence
di: Du, Keyu, et al.
Pubblicazione: (2025)
di: Du, Keyu, et al.
Pubblicazione: (2025)
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
di: Lin, Bin, et al.
Pubblicazione: (2023)
di: Lin, Bin, et al.
Pubblicazione: (2023)
A High-Level Survey of Optical Remote Sensing
di: Koletsis, Panagiotis, et al.
Pubblicazione: (2026)
di: Koletsis, Panagiotis, et al.
Pubblicazione: (2026)
SeeClear: Semantic Distillation Enhances Pixel Condensation for Video Super-Resolution
di: Tang, Qi, et al.
Pubblicazione: (2024)
di: Tang, Qi, et al.
Pubblicazione: (2024)
Video Compression with Hierarchical Temporal Neural Representation
di: Zhu, Jun, et al.
Pubblicazione: (2026)
di: Zhu, Jun, et al.
Pubblicazione: (2026)
Attention, Please! Revisiting Attentive Probing Through the Lens of Efficiency
di: Psomas, Bill, et al.
Pubblicazione: (2025)
di: Psomas, Bill, et al.
Pubblicazione: (2025)
GaussianVideo: Efficient Video Representation via Hierarchical Gaussian Splatting
di: Bond, Andrew, et al.
Pubblicazione: (2025)
di: Bond, Andrew, et al.
Pubblicazione: (2025)
From Semantics to Pixels: Coarse-to-Fine Masked Autoencoders for Hierarchical Visual Understanding
di: Xiang, Wenzhao, et al.
Pubblicazione: (2026)
di: Xiang, Wenzhao, et al.
Pubblicazione: (2026)
Pixel-Accurate Epipolar Guided Matching
di: Nasypanyi, Oleksii, et al.
Pubblicazione: (2026)
di: Nasypanyi, Oleksii, et al.
Pubblicazione: (2026)
Pixel Sentence Representation Learning
di: Xiao, Chenghao, et al.
Pubblicazione: (2024)
di: Xiao, Chenghao, et al.
Pubblicazione: (2024)
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning
di: Jin, Peng, et al.
Pubblicazione: (2024)
di: Jin, Peng, et al.
Pubblicazione: (2024)
PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion Representation
di: Zhao, Junchuan, et al.
Pubblicazione: (2026)
di: Zhao, Junchuan, et al.
Pubblicazione: (2026)
SEAL: Semantic Attention Learning for Long Video Representation
di: Wang, Lan, et al.
Pubblicazione: (2024)
di: Wang, Lan, et al.
Pubblicazione: (2024)
Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction
di: Wei, Yujie, et al.
Pubblicazione: (2026)
di: Wei, Yujie, et al.
Pubblicazione: (2026)
CoordFlow: Coordinate Flow for Pixel-wise Neural Video Representation
di: Silver, Daniel, et al.
Pubblicazione: (2025)
di: Silver, Daniel, et al.
Pubblicazione: (2025)
Language-Guided Graph Representation Learning for Video Summarization
di: Li, Wenrui, et al.
Pubblicazione: (2025)
di: Li, Wenrui, et al.
Pubblicazione: (2025)
CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation
di: Sun, Lin, et al.
Pubblicazione: (2024)
di: Sun, Lin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Advancing Semantic Future Prediction through Multimodal Visual Sequence Transformers
di: Karypidis, Efstathios, et al.
Pubblicazione: (2025) -
DINO-Foresight: Looking into the Future with DINO
di: Karypidis, Efstathios, et al.
Pubblicazione: (2024) -
Boosting Generative Image Modeling via Joint Image-Feature Synthesis
di: Kouzelis, Theodoros, et al.
Pubblicazione: (2025) -
Coevolving Representations in Joint Image-Feature Diffusion
di: Kouzelis, Theodoros, et al.
Pubblicazione: (2026) -
Multi-Token Prediction Needs Registers
di: Gerontopoulos, Anastasios, et al.
Pubblicazione: (2025)