Social-JEPA: Emergent Geometric Isomorphism
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Haoran, Wang, Youjin, Duan, Yi, Fu, Rong, Zhao, Dianyu, Fan, Sicheng, Cao, Shuaishuai, Guo, Wentao, Zhou, Xiao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving
von: Wang, Linhan, et al.
Veröffentlicht: (2026)
von: Wang, Linhan, et al.
Veröffentlicht: (2026)
Texo: Formula Recognition within 20M Parameters
von: Mao, Sicheng
Veröffentlicht: (2026)
von: Mao, Sicheng
Veröffentlicht: (2026)
MammoDINO: Anatomically Aware Self-Supervision for Mammographic Images
von: Zhou, Sicheng, et al.
Veröffentlicht: (2025)
von: Zhou, Sicheng, et al.
Veröffentlicht: (2025)
Self-Supervised JEPA-based World Models for LiDAR Occupancy Completion and Forecasting
von: Zhu, Haoran, et al.
Veröffentlicht: (2026)
von: Zhu, Haoran, et al.
Veröffentlicht: (2026)
BRo-JEPA: Learning Modular Arithmetic in Latent Space
von: Jha, Divyansh, et al.
Veröffentlicht: (2026)
von: Jha, Divyansh, et al.
Veröffentlicht: (2026)
Triple-domain Feature Learning with Frequency-aware Memory Enhancement for Moving Infrared Small Target Detection
von: Duan, Weiwei, et al.
Veröffentlicht: (2024)
von: Duan, Weiwei, et al.
Veröffentlicht: (2024)
Brain-JEPA: Brain Dynamics Foundation Model with Gradient Positioning and Spatiotemporal Masking
von: Dong, Zijian, et al.
Veröffentlicht: (2024)
von: Dong, Zijian, et al.
Veröffentlicht: (2024)
Isomorphic Pruning for Vision Models
von: Fang, Gongfan, et al.
Veröffentlicht: (2024)
von: Fang, Gongfan, et al.
Veröffentlicht: (2024)
Two-Stream Interactive Joint Learning of Scene Parsing and Geometric Vision Tasks
von: Tang, Guanfeng, et al.
Veröffentlicht: (2026)
von: Tang, Guanfeng, et al.
Veröffentlicht: (2026)
Beyond Generative Priors: Minority Sampling with JEPA-Guided Diffusion
von: Park, Sol, et al.
Veröffentlicht: (2026)
von: Park, Sol, et al.
Veröffentlicht: (2026)
TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
von: Han, Yi, et al.
Veröffentlicht: (2025)
von: Han, Yi, et al.
Veröffentlicht: (2025)
Temporal vs. Spatial: Comparing DINOv3 and V-JEPA2 Feature Representations for Video Action Analysis
von: Kodathala, Sai Varun, et al.
Veröffentlicht: (2025)
von: Kodathala, Sai Varun, et al.
Veröffentlicht: (2025)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
von: Li, Yuying, et al.
Veröffentlicht: (2025)
von: Li, Yuying, et al.
Veröffentlicht: (2025)
Debiasing Diffusion Priors via 3D Attention for Consistent Gaussian Splatting
von: Jin, Shilong, et al.
Veröffentlicht: (2025)
von: Jin, Shilong, et al.
Veröffentlicht: (2025)
ReaSon: Reinforced Causal Search with Information Bottleneck for Video Understanding
von: Zhou, Yuan, et al.
Veröffentlicht: (2025)
von: Zhou, Yuan, et al.
Veröffentlicht: (2025)
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
von: Balestriero, Randall, et al.
Veröffentlicht: (2025)
von: Balestriero, Randall, et al.
Veröffentlicht: (2025)
seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World Models
von: Ghaemi, Hafez, et al.
Veröffentlicht: (2025)
von: Ghaemi, Hafez, et al.
Veröffentlicht: (2025)
US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound
von: Radhachandran, Ashwath, et al.
Veröffentlicht: (2026)
von: Radhachandran, Ashwath, et al.
Veröffentlicht: (2026)
Beyond Boundaries: Leveraging Vision Foundation Models for Source-Free Object Detection
von: Yao, Huizai, et al.
Veröffentlicht: (2025)
von: Yao, Huizai, et al.
Veröffentlicht: (2025)
LLMI3D: MLLM-based 3D Perception from a Single 2D Image
von: Yang, Fan, et al.
Veröffentlicht: (2024)
von: Yang, Fan, et al.
Veröffentlicht: (2024)
A$^2$LC: Active and Automated Label Correction for Semantic Segmentation
von: Jeon, Youjin, et al.
Veröffentlicht: (2025)
von: Jeon, Youjin, et al.
Veröffentlicht: (2025)
ConsDreamer: Advancing Multi-View Consistency for Zero-Shot Text-to-3D Generation
von: Zhou, Yuan, et al.
Veröffentlicht: (2025)
von: Zhou, Yuan, et al.
Veröffentlicht: (2025)
Online Monitoring Framework for Automotive Time Series Data using JEPA Embeddings
von: Fertig, Alexander, et al.
Veröffentlicht: (2026)
von: Fertig, Alexander, et al.
Veröffentlicht: (2026)
UR-JEPA: Uniform Rectifiability as a Regularizer for Joint-Embedding Predictive Architectures
von: Le, Triet M.
Veröffentlicht: (2026)
von: Le, Triet M.
Veröffentlicht: (2026)
Ovis-U1 Technical Report
von: Wang, Guo-Hua, et al.
Veröffentlicht: (2025)
von: Wang, Guo-Hua, et al.
Veröffentlicht: (2025)
HiT-JEPA: A Hierarchical Self-supervised Trajectory Embedding Framework for Similarity Computation
von: Li, Lihuan, et al.
Veröffentlicht: (2025)
von: Li, Lihuan, et al.
Veröffentlicht: (2025)
HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator
von: Yang, Fan, et al.
Veröffentlicht: (2024)
von: Yang, Fan, et al.
Veröffentlicht: (2024)
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
von: Assran, Mido, et al.
Veröffentlicht: (2025)
von: Assran, Mido, et al.
Veröffentlicht: (2025)
How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?
von: Lee, Yujian, et al.
Veröffentlicht: (2026)
von: Lee, Yujian, et al.
Veröffentlicht: (2026)
VT-Former: An Exploratory Study on Vehicle Trajectory Prediction for Highway Surveillance through Graph Isomorphism and Transformer
von: Pazho, Armin Danesh, et al.
Veröffentlicht: (2023)
von: Pazho, Armin Danesh, et al.
Veröffentlicht: (2023)
D2Fusion: Dual-domain Fusion with Feature Superposition for Deepfake Detection
von: Qiu, Xueqi, et al.
Veröffentlicht: (2025)
von: Qiu, Xueqi, et al.
Veröffentlicht: (2025)
GERA: Geometric Embedding for Efficient Point Registration Analysis
von: Li, Geng, et al.
Veröffentlicht: (2024)
von: Li, Geng, et al.
Veröffentlicht: (2024)
LightZeroNav: Zero-Shot Vision Language Navigation in Continuous Environments Based on Lightweight VLMs
von: Luo, Kun, et al.
Veröffentlicht: (2026)
von: Luo, Kun, et al.
Veröffentlicht: (2026)
Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning
von: Yeh, Chun-Hsiao, et al.
Veröffentlicht: (2026)
von: Yeh, Chun-Hsiao, et al.
Veröffentlicht: (2026)
Uncertainty-aware Efficient Subgraph Isomorphism using Graph Topology
von: Kusari, Arpan, et al.
Veröffentlicht: (2022)
von: Kusari, Arpan, et al.
Veröffentlicht: (2022)
MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
von: Pang, Yuqi, et al.
Veröffentlicht: (2025)
von: Pang, Yuqi, et al.
Veröffentlicht: (2025)
Lost in UNet: Improving Infrared Small Target Detection by Underappreciated Local Features
von: Quan, Wuzhou, et al.
Veröffentlicht: (2024)
von: Quan, Wuzhou, et al.
Veröffentlicht: (2024)
FloorSAM: SAM-Guided Floorplan Reconstruction with Semantic-Geometric Fusion
von: Ye, Han, et al.
Veröffentlicht: (2025)
von: Ye, Han, et al.
Veröffentlicht: (2025)
Tangram: Benchmark for Evaluating Geometric Element Recognition in Large Multimodal Models
von: Zhang, Chao, et al.
Veröffentlicht: (2024)
von: Zhang, Chao, et al.
Veröffentlicht: (2024)
AdaTP: Attention-Debiased Token Pruning for Video Large Language Models
von: Sun, Fengyuan, et al.
Veröffentlicht: (2025)
von: Sun, Fengyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving
von: Wang, Linhan, et al.
Veröffentlicht: (2026) -
Texo: Formula Recognition within 20M Parameters
von: Mao, Sicheng
Veröffentlicht: (2026) -
MammoDINO: Anatomically Aware Self-Supervision for Mammographic Images
von: Zhou, Sicheng, et al.
Veröffentlicht: (2025) -
Self-Supervised JEPA-based World Models for LiDAR Occupancy Completion and Forecasting
von: Zhu, Haoran, et al.
Veröffentlicht: (2026) -
BRo-JEPA: Learning Modular Arithmetic in Latent Space
von: Jha, Divyansh, et al.
Veröffentlicht: (2026)