VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Han, Nagarajan, Tushar, Ballas, Nicolas, Assran, Mido, Komeili, Mojtaba, Bansal, Mohit, Sinha, Koustuv |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
von: Krojer, Benno, et al.
Veröffentlicht: (2025)
von: Krojer, Benno, et al.
Veröffentlicht: (2025)
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
von: Mur-Labadia, Lorenzo, et al.
Veröffentlicht: (2026)
von: Mur-Labadia, Lorenzo, et al.
Veröffentlicht: (2026)
Learning Latent Action World Models In The Wild
von: Garrido, Quentin, et al.
Veröffentlicht: (2026)
von: Garrido, Quentin, et al.
Veröffentlicht: (2026)
Revisiting Feature Prediction for Learning Visual Representations from Video
von: Bardes, Adrien, et al.
Veröffentlicht: (2024)
von: Bardes, Adrien, et al.
Veröffentlicht: (2024)
Learning and Leveraging World Models in Visual Representation Learning
von: Garrido, Quentin, et al.
Veröffentlicht: (2024)
von: Garrido, Quentin, et al.
Veröffentlicht: (2024)
Scaling Language-Free Visual Representation Learning
von: Fan, David, et al.
Veröffentlicht: (2025)
von: Fan, David, et al.
Veröffentlicht: (2025)
Step Differences in Instructional Video
von: Nagarajan, Tushar, et al.
Veröffentlicht: (2024)
von: Nagarajan, Tushar, et al.
Veröffentlicht: (2024)
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
von: Assran, Mido, et al.
Veröffentlicht: (2025)
von: Assran, Mido, et al.
Veröffentlicht: (2025)
A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures
von: Terver, Basile, et al.
Veröffentlicht: (2026)
von: Terver, Basile, et al.
Veröffentlicht: (2026)
Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos
von: Majumder, Sagnik, et al.
Veröffentlicht: (2024)
von: Majumder, Sagnik, et al.
Veröffentlicht: (2024)
Detours for Navigating Instructional Videos
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2024)
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2024)
Inference-time Physics Alignment of Video Generative Models with Latent World Models
von: Yuan, Jianhao, et al.
Veröffentlicht: (2026)
von: Yuan, Jianhao, et al.
Veröffentlicht: (2026)
Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning
von: Dou, Zi-Yi, et al.
Veröffentlicht: (2024)
von: Dou, Zi-Yi, et al.
Veröffentlicht: (2024)
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
von: Wang, Zun, et al.
Veröffentlicht: (2025)
von: Wang, Zun, et al.
Veröffentlicht: (2025)
Multimodal Representation Learning by Alternating Unimodal Adaptation
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2023)
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
von: Garrido, Quentin, et al.
Veröffentlicht: (2025)
von: Garrido, Quentin, et al.
Veröffentlicht: (2025)
ExpertAF: Expert Actionable Feedback from Video
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2024)
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2024)
SPOC: Spatially-Progressing Object State Change Segmentation in Video
von: Mandikal, Priyanka, et al.
Veröffentlicht: (2025)
von: Mandikal, Priyanka, et al.
Veröffentlicht: (2025)
Multi-Scale and Multi-Layer Contrastive Learning for Domain Generalization
von: Ballas, Aristotelis, et al.
Veröffentlicht: (2023)
von: Ballas, Aristotelis, et al.
Veröffentlicht: (2023)
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
von: Lin, Han, et al.
Veröffentlicht: (2023)
von: Lin, Han, et al.
Veröffentlicht: (2023)
Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
von: Lin, Han, et al.
Veröffentlicht: (2025)
von: Lin, Han, et al.
Veröffentlicht: (2025)
Video ReCap: Recursive Captioning of Hour-Long Videos
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2024)
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2024)
BIMBA: Selective-Scan Compression for Long-Range Video Question Answering
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2025)
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2025)
Video Representation Learning with Joint-Embedding Predictive Architectures
von: Drozdov, Katrina, et al.
Veröffentlicht: (2024)
von: Drozdov, Katrina, et al.
Veröffentlicht: (2024)
SiLVR: A Simple Language-based Video Reasoning Framework
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
Interpretable Few-shot Learning with Online Attribute Selection
von: Zarei, Mohammad Reza, et al.
Veröffentlicht: (2022)
von: Zarei, Mohammad Reza, et al.
Veröffentlicht: (2022)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
von: Lavoie, Samuel, et al.
Veröffentlicht: (2024)
von: Lavoie, Samuel, et al.
Veröffentlicht: (2024)
Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos
von: Majumder, Sagnik, et al.
Veröffentlicht: (2024)
von: Majumder, Sagnik, et al.
Veröffentlicht: (2024)
A Step towards Automated and Generalizable Tactile Map Generation using Generative Adversarial Networks
von: Hobson, David G, et al.
Veröffentlicht: (2024)
von: Hobson, David G, et al.
Veröffentlicht: (2024)
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
von: Wang, Zun, et al.
Veröffentlicht: (2026)
von: Wang, Zun, et al.
Veröffentlicht: (2026)
Flow and Depth Assisted Video Prediction with Latent Transformer
von: Suleyman, Eliyas, et al.
Veröffentlicht: (2025)
von: Suleyman, Eliyas, et al.
Veröffentlicht: (2025)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
von: Wang, Zun, et al.
Veröffentlicht: (2024)
von: Wang, Zun, et al.
Veröffentlicht: (2024)
CycleMix: Mixing Source Domains for Domain Generalization in Style-Dependent Data
von: Ballas, Aristotelis, et al.
Veröffentlicht: (2024)
von: Ballas, Aristotelis, et al.
Veröffentlicht: (2024)
Learning Procedural-aware Video Representations through State-Grounded Hierarchy Unfolding
von: Zhao, Jinghan, et al.
Veröffentlicht: (2025)
von: Zhao, Jinghan, et al.
Veröffentlicht: (2025)
Flatness and Gradient Alignment Are Both Necessary: Spectral-Aware Gradient-Aligned Exploration for Multi-Distribution Learning
von: Ballas, Aristotelis, et al.
Veröffentlicht: (2026)
von: Ballas, Aristotelis, et al.
Veröffentlicht: (2026)
AMEGO: Active Memory from long EGOcentric videos
von: Goletto, Gabriele, et al.
Veröffentlicht: (2024)
von: Goletto, Gabriele, et al.
Veröffentlicht: (2024)
TactileEval: A Step Towards Automated Fine-Grained Evaluation and Editing of Tactile Graphics
von: Khan, Adnan, et al.
Veröffentlicht: (2026)
von: Khan, Adnan, et al.
Veröffentlicht: (2026)
V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising
von: Lin, Han, et al.
Veröffentlicht: (2026)
von: Lin, Han, et al.
Veröffentlicht: (2026)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
von: Wang, Ziyang, et al.
Veröffentlicht: (2024)
von: Wang, Ziyang, et al.
Veröffentlicht: (2024)
Stochastic positional embeddings improve masked image modeling
von: Bar, Amir, et al.
Veröffentlicht: (2023)
von: Bar, Amir, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
von: Krojer, Benno, et al.
Veröffentlicht: (2025) -
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
von: Mur-Labadia, Lorenzo, et al.
Veröffentlicht: (2026) -
Learning Latent Action World Models In The Wild
von: Garrido, Quentin, et al.
Veröffentlicht: (2026) -
Revisiting Feature Prediction for Learning Visual Representations from Video
von: Bardes, Adrien, et al.
Veröffentlicht: (2024) -
Learning and Leveraging World Models in Visual Representation Learning
von: Garrido, Quentin, et al.
Veröffentlicht: (2024)