Social-JEPA: Emergent Geometric Isomorphism
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Haoran, Wang, Youjin, Duan, Yi, Fu, Rong, Zhao, Dianyu, Fan, Sicheng, Cao, Shuaishuai, Guo, Wentao, Zhou, Xiao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving
by: Wang, Linhan, et al.
Published: (2026)
by: Wang, Linhan, et al.
Published: (2026)
Texo: Formula Recognition within 20M Parameters
by: Mao, Sicheng
Published: (2026)
by: Mao, Sicheng
Published: (2026)
MammoDINO: Anatomically Aware Self-Supervision for Mammographic Images
by: Zhou, Sicheng, et al.
Published: (2025)
by: Zhou, Sicheng, et al.
Published: (2025)
Self-Supervised JEPA-based World Models for LiDAR Occupancy Completion and Forecasting
by: Zhu, Haoran, et al.
Published: (2026)
by: Zhu, Haoran, et al.
Published: (2026)
BRo-JEPA: Learning Modular Arithmetic in Latent Space
by: Jha, Divyansh, et al.
Published: (2026)
by: Jha, Divyansh, et al.
Published: (2026)
Triple-domain Feature Learning with Frequency-aware Memory Enhancement for Moving Infrared Small Target Detection
by: Duan, Weiwei, et al.
Published: (2024)
by: Duan, Weiwei, et al.
Published: (2024)
Brain-JEPA: Brain Dynamics Foundation Model with Gradient Positioning and Spatiotemporal Masking
by: Dong, Zijian, et al.
Published: (2024)
by: Dong, Zijian, et al.
Published: (2024)
Isomorphic Pruning for Vision Models
by: Fang, Gongfan, et al.
Published: (2024)
by: Fang, Gongfan, et al.
Published: (2024)
Two-Stream Interactive Joint Learning of Scene Parsing and Geometric Vision Tasks
by: Tang, Guanfeng, et al.
Published: (2026)
by: Tang, Guanfeng, et al.
Published: (2026)
Beyond Generative Priors: Minority Sampling with JEPA-Guided Diffusion
by: Park, Sol, et al.
Published: (2026)
by: Park, Sol, et al.
Published: (2026)
TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
by: Han, Yi, et al.
Published: (2025)
by: Han, Yi, et al.
Published: (2025)
Temporal vs. Spatial: Comparing DINOv3 and V-JEPA2 Feature Representations for Video Action Analysis
by: Kodathala, Sai Varun, et al.
Published: (2025)
by: Kodathala, Sai Varun, et al.
Published: (2025)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
by: Li, Yuying, et al.
Published: (2025)
by: Li, Yuying, et al.
Published: (2025)
Debiasing Diffusion Priors via 3D Attention for Consistent Gaussian Splatting
by: Jin, Shilong, et al.
Published: (2025)
by: Jin, Shilong, et al.
Published: (2025)
ReaSon: Reinforced Causal Search with Information Bottleneck for Video Understanding
by: Zhou, Yuan, et al.
Published: (2025)
by: Zhou, Yuan, et al.
Published: (2025)
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
by: Balestriero, Randall, et al.
Published: (2025)
by: Balestriero, Randall, et al.
Published: (2025)
seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World Models
by: Ghaemi, Hafez, et al.
Published: (2025)
by: Ghaemi, Hafez, et al.
Published: (2025)
US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound
by: Radhachandran, Ashwath, et al.
Published: (2026)
by: Radhachandran, Ashwath, et al.
Published: (2026)
Beyond Boundaries: Leveraging Vision Foundation Models for Source-Free Object Detection
by: Yao, Huizai, et al.
Published: (2025)
by: Yao, Huizai, et al.
Published: (2025)
LLMI3D: MLLM-based 3D Perception from a Single 2D Image
by: Yang, Fan, et al.
Published: (2024)
by: Yang, Fan, et al.
Published: (2024)
A$^2$LC: Active and Automated Label Correction for Semantic Segmentation
by: Jeon, Youjin, et al.
Published: (2025)
by: Jeon, Youjin, et al.
Published: (2025)
ConsDreamer: Advancing Multi-View Consistency for Zero-Shot Text-to-3D Generation
by: Zhou, Yuan, et al.
Published: (2025)
by: Zhou, Yuan, et al.
Published: (2025)
Online Monitoring Framework for Automotive Time Series Data using JEPA Embeddings
by: Fertig, Alexander, et al.
Published: (2026)
by: Fertig, Alexander, et al.
Published: (2026)
UR-JEPA: Uniform Rectifiability as a Regularizer for Joint-Embedding Predictive Architectures
by: Le, Triet M.
Published: (2026)
by: Le, Triet M.
Published: (2026)
Ovis-U1 Technical Report
by: Wang, Guo-Hua, et al.
Published: (2025)
by: Wang, Guo-Hua, et al.
Published: (2025)
HiT-JEPA: A Hierarchical Self-supervised Trajectory Embedding Framework for Similarity Computation
by: Li, Lihuan, et al.
Published: (2025)
by: Li, Lihuan, et al.
Published: (2025)
HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator
by: Yang, Fan, et al.
Published: (2024)
by: Yang, Fan, et al.
Published: (2024)
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
by: Assran, Mido, et al.
Published: (2025)
by: Assran, Mido, et al.
Published: (2025)
VT-Former: An Exploratory Study on Vehicle Trajectory Prediction for Highway Surveillance through Graph Isomorphism and Transformer
by: Pazho, Armin Danesh, et al.
Published: (2023)
by: Pazho, Armin Danesh, et al.
Published: (2023)
How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?
by: Lee, Yujian, et al.
Published: (2026)
by: Lee, Yujian, et al.
Published: (2026)
D2Fusion: Dual-domain Fusion with Feature Superposition for Deepfake Detection
by: Qiu, Xueqi, et al.
Published: (2025)
by: Qiu, Xueqi, et al.
Published: (2025)
GERA: Geometric Embedding for Efficient Point Registration Analysis
by: Li, Geng, et al.
Published: (2024)
by: Li, Geng, et al.
Published: (2024)
LightZeroNav: Zero-Shot Vision Language Navigation in Continuous Environments Based on Lightweight VLMs
by: Luo, Kun, et al.
Published: (2026)
by: Luo, Kun, et al.
Published: (2026)
Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning
by: Yeh, Chun-Hsiao, et al.
Published: (2026)
by: Yeh, Chun-Hsiao, et al.
Published: (2026)
Uncertainty-aware Efficient Subgraph Isomorphism using Graph Topology
by: Kusari, Arpan, et al.
Published: (2022)
by: Kusari, Arpan, et al.
Published: (2022)
MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
by: Pang, Yuqi, et al.
Published: (2025)
by: Pang, Yuqi, et al.
Published: (2025)
Lost in UNet: Improving Infrared Small Target Detection by Underappreciated Local Features
by: Quan, Wuzhou, et al.
Published: (2024)
by: Quan, Wuzhou, et al.
Published: (2024)
FloorSAM: SAM-Guided Floorplan Reconstruction with Semantic-Geometric Fusion
by: Ye, Han, et al.
Published: (2025)
by: Ye, Han, et al.
Published: (2025)
Tangram: Benchmark for Evaluating Geometric Element Recognition in Large Multimodal Models
by: Zhang, Chao, et al.
Published: (2024)
by: Zhang, Chao, et al.
Published: (2024)
AdaTP: Attention-Debiased Token Pruning for Video Large Language Models
by: Sun, Fengyuan, et al.
Published: (2025)
by: Sun, Fengyuan, et al.
Published: (2025)
Similar Items
-
Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving
by: Wang, Linhan, et al.
Published: (2026) -
Texo: Formula Recognition within 20M Parameters
by: Mao, Sicheng
Published: (2026) -
MammoDINO: Anatomically Aware Self-Supervision for Mammographic Images
by: Zhou, Sicheng, et al.
Published: (2025) -
Self-Supervised JEPA-based World Models for LiDAR Occupancy Completion and Forecasting
by: Zhu, Haoran, et al.
Published: (2026) -
BRo-JEPA: Learning Modular Arithmetic in Latent Space
by: Jha, Divyansh, et al.
Published: (2026)