Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xianhang, Huang, Chen, Li, Chun-Liang, Malach, Eran, Susskind, Josh, Thilak, Vimal, Littwin, Etai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Text-Conditional JEPA for Learning Semantically Rich Visual Representations
by: Huang, Chen, et al.
Published: (2026)
by: Huang, Chen, et al.
Published: (2026)
Enhancing JEPAs with Spatial Conditioning: Robust and Efficient Representation Learning
by: Littwin, Etai, et al.
Published: (2024)
by: Littwin, Etai, et al.
Published: (2024)
How JEPA Avoids Noisy Features: The Implicit Bias of Deep Linear Self Distillation Networks
by: Littwin, Etai, et al.
Published: (2024)
by: Littwin, Etai, et al.
Published: (2024)
To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models
by: Malach, Eran, et al.
Published: (2025)
by: Malach, Eran, et al.
Published: (2025)
Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving
by: Wang, Linhan, et al.
Published: (2026)
by: Wang, Linhan, et al.
Published: (2026)
Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD Generalization
by: Zang, Yuhang, et al.
Published: (2024)
by: Zang, Yuhang, et al.
Published: (2024)
Steering Video Diffusion Transformers with Massive Activations
by: Cheng, Xianhang, et al.
Published: (2026)
by: Cheng, Xianhang, et al.
Published: (2026)
Efficient Temporal Sentence Grounding in Videos with Multi-Teacher Knowledge Distillation
by: Liang, Renjie, et al.
Published: (2023)
by: Liang, Renjie, et al.
Published: (2023)
How PARTs assemble into wholes: Learning the relative composition of images
by: Ayoughi, Melika, et al.
Published: (2025)
by: Ayoughi, Melika, et al.
Published: (2025)
JEPA-VLA: Video Predictive Embedding is Needed for VLA Models
by: Miao, Shangchen, et al.
Published: (2026)
by: Miao, Shangchen, et al.
Published: (2026)
TAP-JEPA: Frozen Future-Latent Probing and Two-Stage Score Fusion for EPIC-KITCHENS-100 Action Anticipation
by: Wang, Chaoyang, et al.
Published: (2026)
by: Wang, Chaoyang, et al.
Published: (2026)
Rethinking Reward Signals in Video GRPO: When Scores Become Targets
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
Vanishing Gradients in Reinforcement Finetuning of Language Models
by: Razin, Noam, et al.
Published: (2023)
by: Razin, Noam, et al.
Published: (2023)
Learning Long-term Motion Embeddings for Efficient Kinematics Generation
by: Stracke, Nick, et al.
Published: (2026)
by: Stracke, Nick, et al.
Published: (2026)
TIR-Flow: Active Video Search and Reasoning with Frozen VLMs
by: Jin, Hongbo, et al.
Published: (2026)
by: Jin, Hongbo, et al.
Published: (2026)
Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models
by: Orlova, Svetlana, et al.
Published: (2026)
by: Orlova, Svetlana, et al.
Published: (2026)
Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
by: Gui, Ming, et al.
Published: (2025)
by: Gui, Ming, et al.
Published: (2025)
Unmasked Teacher: Towards Training-Efficient Video Foundation Models
by: Li, Kunchang, et al.
Published: (2023)
by: Li, Kunchang, et al.
Published: (2023)
Revisiting Adversarial Training at Scale
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
JEPA-T: Joint-Embedding Predictive Architecture with Text Fusion for Image Generation
by: Wan, Siheng, et al.
Published: (2025)
by: Wan, Siheng, et al.
Published: (2025)
SSL: A Self-similarity Loss for Improving Generative Image Super-resolution
by: Chen, Du, et al.
Published: (2024)
by: Chen, Du, et al.
Published: (2024)
Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP
by: Huang, Chen, et al.
Published: (2024)
by: Huang, Chen, et al.
Published: (2024)
What Happens Next? Next Scene Prediction with a Unified Video Model
by: Li, Xinjie, et al.
Published: (2025)
by: Li, Xinjie, et al.
Published: (2025)
InstanceGen: Image Generation with Instance-level Instructions
by: Sella, Etai, et al.
Published: (2025)
by: Sella, Etai, et al.
Published: (2025)
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
by: Mur-Labadia, Lorenzo, et al.
Published: (2026)
by: Mur-Labadia, Lorenzo, et al.
Published: (2026)
Text-Guided Video Masked Autoencoder
by: Fan, David, et al.
Published: (2024)
by: Fan, David, et al.
Published: (2024)
Rethinking Noise-Robust Training for Frozen Vision Foundation Models: A Cross-Dataset Benchmark with a Case Study of Small-Loss Failure
by: Li, Zitong, et al.
Published: (2026)
by: Li, Zitong, et al.
Published: (2026)
Rethinking Video Segmentation with Masked Video Consistency: Did the Model Learn as Intended?
by: Liang, Chen, et al.
Published: (2024)
by: Liang, Chen, et al.
Published: (2024)
DiffSign: AI-Assisted Generation of Customizable Sign Language Videos With Enhanced Realism
by: Krishnamurthy, Sudha, et al.
Published: (2024)
by: Krishnamurthy, Sudha, et al.
Published: (2024)
FROSTER: Frozen CLIP Is A Strong Teacher for Open-Vocabulary Action Recognition
by: Huang, Xiaohu, et al.
Published: (2024)
by: Huang, Xiaohu, et al.
Published: (2024)
Steering and Rectifying Latent Representation Manifolds in Frozen Multi-modal LLMs for Video Anomaly Detection
by: Cai, Zhaolin, et al.
Published: (2026)
by: Cai, Zhaolin, et al.
Published: (2026)
Factorized Latent Dynamics for Video JEPA: An Empirical Study of Auxiliary Objectives
by: Premi, Santosh
Published: (2026)
by: Premi, Santosh
Published: (2026)
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
by: Liu, Yanqing, et al.
Published: (2024)
by: Liu, Yanqing, et al.
Published: (2024)
OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning
by: Li, Xianhang, et al.
Published: (2025)
by: Li, Xianhang, et al.
Published: (2025)
FrozenSeg: Harmonizing Frozen Foundation Models for Open-Vocabulary Segmentation
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
Evaluation of deep learning architectures for wildlife object detection: A comparative study of ResNet and Inception
by: Amonga, Malach Obisa, et al.
Published: (2025)
by: Amonga, Malach Obisa, et al.
Published: (2025)
HyCoVAD: A Hybrid SSL-LLM Model for Complex Video Anomaly Detection
by: Hemmatyar, Mohammad Mahdi, et al.
Published: (2025)
by: Hemmatyar, Mohammad Mahdi, et al.
Published: (2025)
STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows
by: Gu, Jiatao, et al.
Published: (2025)
by: Gu, Jiatao, et al.
Published: (2025)
PP-SSL : Priority-Perception Self-Supervised Learning for Fine-Grained Recognition
by: Li, ShuaiHeng, et al.
Published: (2024)
by: Li, ShuaiHeng, et al.
Published: (2024)
Matryoshka Diffusion Models
by: Gu, Jiatao, et al.
Published: (2023)
by: Gu, Jiatao, et al.
Published: (2023)
Similar Items
-
Text-Conditional JEPA for Learning Semantically Rich Visual Representations
by: Huang, Chen, et al.
Published: (2026) -
Enhancing JEPAs with Spatial Conditioning: Robust and Efficient Representation Learning
by: Littwin, Etai, et al.
Published: (2024) -
How JEPA Avoids Noisy Features: The Implicit Bias of Deep Linear Self Distillation Networks
by: Littwin, Etai, et al.
Published: (2024) -
To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models
by: Malach, Eran, et al.
Published: (2025) -
Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving
by: Wang, Linhan, et al.
Published: (2026)