DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
Fuente:
arXiv
Saved in:
| Main Authors: | He, Xiangteng, Sakai, Shunsuke, Chandhok, Shivam, Beery, Sara, Yuan, Kun, Padoy, Nicolas, Hasegawa, Tatsuhito, Sigal, Leonid |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models
by: Sakai, Shunsuke, et al.
Published: (2025)
by: Sakai, Shunsuke, et al.
Published: (2025)
Analytical Softmax Temperature Setting from Feature Dimensions for Model- and Domain-Robust Classification
by: Hasegawa, Tatsuhito, et al.
Published: (2025)
by: Hasegawa, Tatsuhito, et al.
Published: (2025)
Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities
by: Chandhok, Shivam, et al.
Published: (2024)
by: Chandhok, Shivam, et al.
Published: (2024)
Noisy Deep Ensemble: Accelerating Deep Ensemble Learning via Noise Injection
by: Sakai, Shunsuke, et al.
Published: (2025)
by: Sakai, Shunsuke, et al.
Published: (2025)
Contrastive Learning-Enhanced Trajectory Matching for Small-Scale Dataset Distillation
by: Li, Wenmin, et al.
Published: (2025)
by: Li, Wenmin, et al.
Published: (2025)
DMT-JEPA: Discriminative Masked Targets for Joint-Embedding Predictive Architecture
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
MM-R$^3$: On (In-)Consistency of Vision-Language Models (VLMs)
by: Chou, Shih-Han, et al.
Published: (2024)
by: Chou, Shih-Han, et al.
Published: (2024)
Test-Time Consistency in Vision Language Models
by: Chou, Shih-Han, et al.
Published: (2025)
by: Chou, Shih-Han, et al.
Published: (2025)
SceneGPT: A Language Model for 3D Scene Understanding
by: Chandhok, Shivam
Published: (2024)
by: Chandhok, Shivam
Published: (2024)
Pool-Select-Refine: Allocation-Aware Generative Dataset Distillation with Soft-Label-Guided Latent Refinement
by: Li, Wenmin, et al.
Published: (2026)
by: Li, Wenmin, et al.
Published: (2026)
VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
by: Chen, Delong, et al.
Published: (2025)
by: Chen, Delong, et al.
Published: (2025)
Learning What Matters: Prioritized Concept Learning via Relative Error-driven Sample Selection
by: Chandhok, Shivam, et al.
Published: (2025)
by: Chandhok, Shivam, et al.
Published: (2025)
A-JEPA: Joint-Embedding Predictive Architecture Can Listen
by: Fei, Zhengcong, et al.
Published: (2023)
by: Fei, Zhengcong, et al.
Published: (2023)
JEPA for RL: Investigating Joint-Embedding Predictive Architectures for Reinforcement Learning
by: Kenneweg, Tristan, et al.
Published: (2025)
by: Kenneweg, Tristan, et al.
Published: (2025)
JEPA-T: Joint-Embedding Predictive Architecture with Text Fusion for Image Generation
by: Wan, Siheng, et al.
Published: (2025)
by: Wan, Siheng, et al.
Published: (2025)
Do Vision-Language Foundational models show Robust Visual Perception?
by: Chandhok, Shivam, et al.
Published: (2024)
by: Chandhok, Shivam, et al.
Published: (2024)
CoRDS: Coreset-based Representative and Diverse Selection for Streaming Video Understanding
by: Mahdizadeh, Ailar, et al.
Published: (2026)
by: Mahdizadeh, Ailar, et al.
Published: (2026)
US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound
by: Radhachandran, Ashwath, et al.
Published: (2026)
by: Radhachandran, Ashwath, et al.
Published: (2026)
UR-JEPA: Uniform Rectifiability as a Regularizer for Joint-Embedding Predictive Architectures
by: Le, Triet M.
Published: (2026)
by: Le, Triet M.
Published: (2026)
Rectified LpJEPA: Joint-Embedding Predictive Architectures with Sparse and Maximum-Entropy Representations
by: Kuang, Yilun, et al.
Published: (2026)
by: Kuang, Yilun, et al.
Published: (2026)
Point-JEPA: A Joint Embedding Predictive Architecture for Self-Supervised Learning on Point Cloud
by: Saito, Ayumu, et al.
Published: (2024)
by: Saito, Ayumu, et al.
Published: (2024)
RadJEPA: Radiology Encoder for Chest X-Rays via Joint Embedding Predictive Architecture
by: Khan, Anas Anwarul Haq, et al.
Published: (2026)
by: Khan, Anas Anwarul Haq, et al.
Published: (2026)
CNN-JEPA: Self-Supervised Pretraining Convolutional Neural Networks Using Joint Embedding Predictive Architecture
by: Kalapos, András, et al.
Published: (2024)
by: Kalapos, András, et al.
Published: (2024)
Easy Ensemble: Simple Deep Ensemble Learning for Sensor-Based Human Activity Recognition
by: Hasegawa, Tatsuhito, et al.
Published: (2022)
by: Hasegawa, Tatsuhito, et al.
Published: (2022)
SCALE-VLP: Soft-Weighted Contrastive Volumetric Vision-Language Pre-training with Spatial-Knowledge Semantics
by: Mahdizadeh, Ailar, et al.
Published: (2025)
by: Mahdizadeh, Ailar, et al.
Published: (2025)
HQ-JEPA: Hybrid Quantum Joint-Embedding Predictive Architecture for Cross-Modal Remote Sensing Representation Learning
by: Hossain, Md Aminur, et al.
Published: (2026)
by: Hossain, Md Aminur, et al.
Published: (2026)
3D-JEPA: A Joint Embedding Predictive Architecture for 3D Self-Supervised Representation Learning
by: Hu, Naiwen, et al.
Published: (2024)
by: Hu, Naiwen, et al.
Published: (2024)
CrossJEPA: Cross-Modal Joint-Embedding Predictive Architecture for Efficient 3D Representation Learning from 2D Images
by: Perera, Avishka, et al.
Published: (2025)
by: Perera, Avishka, et al.
Published: (2025)
LADMIM: Logical Anomaly Detection with Masked Image Modeling in Discrete Latent Space
by: Sakai, Shunsuke, et al.
Published: (2024)
by: Sakai, Shunsuke, et al.
Published: (2024)
CR-JEPA: Cross-Modal Joint-Embedding Predictive Learning for Remote Sensing Image Retrieval
by: Hossain, Md Aminur, et al.
Published: (2026)
by: Hossain, Md Aminur, et al.
Published: (2026)
Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities
by: Chandhok, Shivam, et al.
Published: (2025)
by: Chandhok, Shivam, et al.
Published: (2025)
JEPA-VLA: Video Predictive Embedding is Needed for VLA Models
by: Miao, Shangchen, et al.
Published: (2026)
by: Miao, Shangchen, et al.
Published: (2026)
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
by: Luo, Jiayun, et al.
Published: (2025)
by: Luo, Jiayun, et al.
Published: (2025)
From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature
by: Yuan, Kun, et al.
Published: (2025)
by: Yuan, Kun, et al.
Published: (2025)
MTS-JEPA: Multi-Resolution Joint-Embedding Predictive Architecture for Time-Series Anomaly Prediction
by: He, Yanan, et al.
Published: (2026)
by: He, Yanan, et al.
Published: (2026)
HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition
by: Yuan, Kun, et al.
Published: (2024)
by: Yuan, Kun, et al.
Published: (2024)
Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation
by: Yuan, Kun, et al.
Published: (2024)
by: Yuan, Kun, et al.
Published: (2024)
JEPA4Rec: Learning Effective Language Representations for Sequential Recommendation via Joint Embedding Predictive Architecture
by: Nguyen, Minh-Anh, et al.
Published: (2025)
by: Nguyen, Minh-Anh, et al.
Published: (2025)
T-JEPA: A Joint-Embedding Predictive Architecture for Trajectory Similarity Computation
by: Li, Lihuan, et al.
Published: (2024)
by: Li, Lihuan, et al.
Published: (2024)
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
by: Tuncay, Ludovic, et al.
Published: (2025)
by: Tuncay, Ludovic, et al.
Published: (2025)
Similar Items
-
InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models
by: Sakai, Shunsuke, et al.
Published: (2025) -
Analytical Softmax Temperature Setting from Feature Dimensions for Model- and Domain-Robust Classification
by: Hasegawa, Tatsuhito, et al.
Published: (2025) -
Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities
by: Chandhok, Shivam, et al.
Published: (2024) -
Noisy Deep Ensemble: Accelerating Deep Ensemble Learning via Noise Injection
by: Sakai, Shunsuke, et al.
Published: (2025) -
Contrastive Learning-Enhanced Trajectory Matching for Small-Scale Dataset Distillation
by: Li, Wenmin, et al.
Published: (2025)