ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Haichao, Li, Yijiang, He, Shwai, Nagarajan, Tushar, Chen, Mingfei, Lu, Jianglin, Li, Ang, Fu, Yun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cooperative Perception: A Resource-Efficient Framework for Multi-Drone 3D Scene Reconstruction Using Federated Diffusion and NeRF
by: Pourmandi, Massoud
Published: (2025)
by: Pourmandi, Massoud
Published: (2025)
Learning Accurate Whole-body Throwing with High-frequency Residual Policy and Pullback Tube Acceleration
by: Ma, Yuntao, et al.
Published: (2025)
by: Ma, Yuntao, et al.
Published: (2025)
Towards Ubiquitous Mapping and Localization for Dynamic Indoor Environments
by: Djerroud, Halim, et al.
Published: (2026)
by: Djerroud, Halim, et al.
Published: (2026)
Key-Scan-Based Mobile Robot Navigation: Integrated Mapping, Planning, and Control using Graphs of Scan Regions
by: Latha, Dharshan Bashkaran, et al.
Published: (2024)
by: Latha, Dharshan Bashkaran, et al.
Published: (2024)
Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features
by: Chahine, Makram, et al.
Published: (2024)
by: Chahine, Makram, et al.
Published: (2024)
CADE 2.5 - ZeResFDG: Frequency-Decoupled, Rescaled and Zero-Projected Guidance for SD/SDXL Latent Diffusion Models
by: Rychkovskiy, Denis
Published: (2025)
by: Rychkovskiy, Denis
Published: (2025)
Out-of-Sight Embodied Agents: Multimodal Tracking, Sensor Fusion, and Trajectory Forecasting
by: Zhang, Haichao, et al.
Published: (2025)
by: Zhang, Haichao, et al.
Published: (2025)
Beyond RGB: Leveraging Vision Transformers for Thermal Weapon Segmentation
by: Kambhatla, Akhila, et al.
Published: (2025)
by: Kambhatla, Akhila, et al.
Published: (2025)
QSilk: Micrograin Stabilization and Adaptive Quantile Clipping for Detail-Friendly Latent Diffusion
by: Rychkovskiy, Denis
Published: (2025)
by: Rychkovskiy, Denis
Published: (2025)
Gr-IoU: Ground-Intersection over Union for Robust Multi-Object Tracking with 3D Geometric Constraints
by: Toida, Keisuke, et al.
Published: (2024)
by: Toida, Keisuke, et al.
Published: (2024)
Sat-JEPA-Diff: Bridging Self-Supervised Learning and Generative Diffusion for Remote Sensing
by: Komurcu, Kursat, et al.
Published: (2026)
by: Komurcu, Kursat, et al.
Published: (2026)
Agentic UAVs: LLM-Driven Autonomy with Integrated Tool-Calling and Cognitive Reasoning
by: Koubaa, Anis, et al.
Published: (2025)
by: Koubaa, Anis, et al.
Published: (2025)
The Impact of Class Uncertainty Propagation in Perception-Based Motion Planning
by: Shah, Jibran Iqbal, et al.
Published: (2026)
by: Shah, Jibran Iqbal, et al.
Published: (2026)
Learning Adaptive Neural Teleoperation for Humanoid Robots: From Inverse Kinematics to End-to-End Control
by: Atamuradov, Sanjar
Published: (2025)
by: Atamuradov, Sanjar
Published: (2025)
YOLO Ensemble for UAV-based Multispectral Defect Detection in Wind Turbine Components
by: Svystun, Serhii, et al.
Published: (2025)
by: Svystun, Serhii, et al.
Published: (2025)
BG-YOLO: A Bidirectional-Guided Method for Underwater Object Detection
by: Zhang, Jian, et al.
Published: (2024)
by: Zhang, Jian, et al.
Published: (2024)
Experimental Evaluation of Road-Crossing Decisions by Autonomous Wheelchairs against Environmental Factors
by: Corradini, Franca, et al.
Published: (2024)
by: Corradini, Franca, et al.
Published: (2024)
CulinaryCut-VLAP: A Vision-Language-Action-Physics Framework for Food Cutting via a Force-Aware Material Point Method
by: Koh, Hyunseo, et al.
Published: (2026)
by: Koh, Hyunseo, et al.
Published: (2026)
ElasticFlow: One-Step Physics-Consistent Policy with Elastic Time Horizons for Language-Guided Manipulation
by: Chen, Kewei, et al.
Published: (2026)
by: Chen, Kewei, et al.
Published: (2026)
EventFlow: Real-Time Neuromorphic Event-Driven Classification of Two-Phase Boiling Flow Regimes
by: Chang, Sanghyeon, et al.
Published: (2025)
by: Chang, Sanghyeon, et al.
Published: (2025)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
by: Raoufi, Behnam, et al.
Published: (2025)
by: Raoufi, Behnam, et al.
Published: (2025)
Method of UAV Inspection of Photovoltaic Modules Using Thermal and RGB Data Fusion
by: Lysyi, Andrii, et al.
Published: (2025)
by: Lysyi, Andrii, et al.
Published: (2025)
SemanticFeels: Semantic Labeling during In-Hand Manipulation
by: Khalil, Anas Al Shikh, et al.
Published: (2026)
by: Khalil, Anas Al Shikh, et al.
Published: (2026)
A Landmark-Aware Visual Navigation Dataset
by: Johnson, Faith, et al.
Published: (2024)
by: Johnson, Faith, et al.
Published: (2024)
Measuring What Matters: Scenario-Driven Evaluation for Trajectory Predictors in Autonomous Driving
by: Da, Longchao, et al.
Published: (2025)
by: Da, Longchao, et al.
Published: (2025)
PathFormer: A Transformer with 3D Grid Constraints for Digital Twin Robot-Arm Trajectory Generation
by: Alanazi, Ahmed, et al.
Published: (2025)
by: Alanazi, Ahmed, et al.
Published: (2025)
MVTamperBench: Evaluating Robustness of Vision-Language Models
by: Agarwal, Amit, et al.
Published: (2024)
by: Agarwal, Amit, et al.
Published: (2024)
Safe Road-Crossing by Autonomous Wheelchairs: a Novel Dataset and its Experimental Evaluation
by: Grigioni, Carlo, et al.
Published: (2024)
by: Grigioni, Carlo, et al.
Published: (2024)
Thermal RGB Fusion for Micro-UAV Wildfire Perimeter Tracking with Minimal Comms
by: Erkalkan, Ercan, et al.
Published: (2025)
by: Erkalkan, Ercan, et al.
Published: (2025)
S3Simulator: A benchmarking Side Scan Sonar Simulator dataset for Underwater Image Analysis
by: S, Kamal Basha, et al.
Published: (2024)
by: S, Kamal Basha, et al.
Published: (2024)
ExpReS-VLA: Specializing Vision-Language-Action Models Through Experience Replay and Retrieval
by: Syed, Shahram Najam, et al.
Published: (2025)
by: Syed, Shahram Najam, et al.
Published: (2025)
Do Generative Metrics Predict YOLO Performance? An Evaluation Across Models, Augmentation Ratios, and Dataset Complexity
by: Marian, Vasile, et al.
Published: (2026)
by: Marian, Vasile, et al.
Published: (2026)
Dense Video Understanding with Gated Residual Tokenization
by: Zhang, Haichao, et al.
Published: (2025)
by: Zhang, Haichao, et al.
Published: (2025)
Learning coordinated badminton skills for legged manipulators
by: Ma, Yuntao, et al.
Published: (2025)
by: Ma, Yuntao, et al.
Published: (2025)
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
by: Chen, Kewei, et al.
Published: (2025)
by: Chen, Kewei, et al.
Published: (2025)
RoboMoRe: LLM-based Robot Co-design via Joint Optimization of Morphology and Reward
by: Fang, Jiawei, et al.
Published: (2025)
by: Fang, Jiawei, et al.
Published: (2025)
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
by: Chen, Kewei, et al.
Published: (2026)
by: Chen, Kewei, et al.
Published: (2026)
Task and Motion Planning in Hierarchical 3D Scene Graphs
by: Ray, Aaron, et al.
Published: (2024)
by: Ray, Aaron, et al.
Published: (2024)
COBRA-PPM: A Causal Bayesian Reasoning Architecture Using Probabilistic Programming for Robot Manipulation Under Uncertainty
by: Cannizzaro, Ricardo, et al.
Published: (2024)
by: Cannizzaro, Ricardo, et al.
Published: (2024)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
by: Qesaraku, Bjorna, et al.
Published: (2025)
by: Qesaraku, Bjorna, et al.
Published: (2025)
Similar Items
-
Cooperative Perception: A Resource-Efficient Framework for Multi-Drone 3D Scene Reconstruction Using Federated Diffusion and NeRF
by: Pourmandi, Massoud
Published: (2025) -
Learning Accurate Whole-body Throwing with High-frequency Residual Policy and Pullback Tube Acceleration
by: Ma, Yuntao, et al.
Published: (2025) -
Towards Ubiquitous Mapping and Localization for Dynamic Indoor Environments
by: Djerroud, Halim, et al.
Published: (2026) -
Key-Scan-Based Mobile Robot Navigation: Integrated Mapping, Planning, and Control using Graphs of Scan Regions
by: Latha, Dharshan Bashkaran, et al.
Published: (2024) -
Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features
by: Chahine, Makram, et al.
Published: (2024)