The Power of Next-Frame Prediction for Learning Physical Laws
Fuente:
arXiv
Guardado en:
| Autores principales: | Winterbottom, Thomas, Hudson, G. Thomas, Kluvanec, Daniel, Slack, Dean, Sterling, Jamie, Shentu, Junjie, Xiao, Chenghao, Zhou, Zheming, Moubayed, Noura Al |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Tricks and Plug-ins for Gradient Boosting in Image Classification
por: Fang, Biyi, et al.
Publicado: (2025)
por: Fang, Biyi, et al.
Publicado: (2025)
VA-$π$: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
por: Liao, Xinyao, et al.
Publicado: (2025)
por: Liao, Xinyao, et al.
Publicado: (2025)
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
por: Chen, Kewei, et al.
Publicado: (2025)
por: Chen, Kewei, et al.
Publicado: (2025)
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
por: Chen, Kewei, et al.
Publicado: (2026)
por: Chen, Kewei, et al.
Publicado: (2026)
Unpacking Hateful Memes: Presupposed Context and False Claims
por: Cai, Weibin, et al.
Publicado: (2025)
por: Cai, Weibin, et al.
Publicado: (2025)
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
por: Tu, Songjun, et al.
Publicado: (2025)
por: Tu, Songjun, et al.
Publicado: (2025)
CulinaryCut-VLAP: A Vision-Language-Action-Physics Framework for Food Cutting via a Force-Aware Material Point Method
por: Koh, Hyunseo, et al.
Publicado: (2026)
por: Koh, Hyunseo, et al.
Publicado: (2026)
Training a Student Expert via Semi-Supervised Foundation Model Distillation
por: Taghavi, Pardis, et al.
Publicado: (2026)
por: Taghavi, Pardis, et al.
Publicado: (2026)
Physics-informed Variational Autoencoders for Improved Robustness to Environmental Factors of Variation
por: Thoreau, Romain, et al.
Publicado: (2022)
por: Thoreau, Romain, et al.
Publicado: (2022)
SemanticFeels: Semantic Labeling during In-Hand Manipulation
por: Khalil, Anas Al Shikh, et al.
Publicado: (2026)
por: Khalil, Anas Al Shikh, et al.
Publicado: (2026)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
por: Qesaraku, Bjorna, et al.
Publicado: (2025)
por: Qesaraku, Bjorna, et al.
Publicado: (2025)
Do Generative Metrics Predict YOLO Performance? An Evaluation Across Models, Augmentation Ratios, and Dataset Complexity
por: Marian, Vasile, et al.
Publicado: (2026)
por: Marian, Vasile, et al.
Publicado: (2026)
Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling
por: Jung, Seoik, et al.
Publicado: (2025)
por: Jung, Seoik, et al.
Publicado: (2025)
Training for X-Ray Vision: Amodal Segmentation, Amodal Content Completion, and View-Invariant Object Representation from Multi-Camera Video
por: Moore, Alexander, et al.
Publicado: (2025)
por: Moore, Alexander, et al.
Publicado: (2025)
Cooperative Perception: A Resource-Efficient Framework for Multi-Drone 3D Scene Reconstruction Using Federated Diffusion and NeRF
por: Pourmandi, Massoud
Publicado: (2025)
por: Pourmandi, Massoud
Publicado: (2025)
Visible and Hyperspectral Imaging for Quality Assessment of Milk: Property Characterisation and Identification
por: Martinelli, Massimo, et al.
Publicado: (2026)
por: Martinelli, Massimo, et al.
Publicado: (2026)
Balanced conic rectified flow
por: Kim, Shin Seong, et al.
Publicado: (2025)
por: Kim, Shin Seong, et al.
Publicado: (2025)
Spectral Integrated Gradients for Coarse-to-Fine Feature Attribution
por: Kim, Soyeon, et al.
Publicado: (2026)
por: Kim, Soyeon, et al.
Publicado: (2026)
Manifold-Aligned Guided Integrated Gradients for Reliable Feature Attribution
por: Kim, Soyeon, et al.
Publicado: (2026)
por: Kim, Soyeon, et al.
Publicado: (2026)
Method of UAV Inspection of Photovoltaic Modules Using Thermal and RGB Data Fusion
por: Lysyi, Andrii, et al.
Publicado: (2025)
por: Lysyi, Andrii, et al.
Publicado: (2025)
ExpReS-VLA: Specializing Vision-Language-Action Models Through Experience Replay and Retrieval
por: Syed, Shahram Najam, et al.
Publicado: (2025)
por: Syed, Shahram Najam, et al.
Publicado: (2025)
Supervised Learning Has a Necessary Geometric Blind Spot: Theory, Consequences, and Minimal Repair
por: Rajput, Vishal
Publicado: (2026)
por: Rajput, Vishal
Publicado: (2026)
AUTHENTICATION: Identifying Rare Failure Modes in Autonomous Vehicle Perception Systems using Adversarially Guided Diffusion Models
por: Zarei, Mohammad, et al.
Publicado: (2025)
por: Zarei, Mohammad, et al.
Publicado: (2025)
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
por: Sáez, Arnau Igualde, et al.
Publicado: (2025)
por: Sáez, Arnau Igualde, et al.
Publicado: (2025)
Rethinking Visual Intelligence: Insights from Video Pretraining
por: Acuaviva, Pablo, et al.
Publicado: (2025)
por: Acuaviva, Pablo, et al.
Publicado: (2025)
Multimodal Generative AI for Story Point Estimation in Software Development
por: Islam, Mohammad Rubyet, et al.
Publicado: (2025)
por: Islam, Mohammad Rubyet, et al.
Publicado: (2025)
Contrastive Consolidation of Top-Down Modulations Achieves Sparsely Supervised Continual Learning
por: Tran, Viet Anh Khoa, et al.
Publicado: (2025)
por: Tran, Viet Anh Khoa, et al.
Publicado: (2025)
IDOL: Instant Photorealistic 3D Human Creation from a Single Image
por: Zhuang, Yiyu, et al.
Publicado: (2024)
por: Zhuang, Yiyu, et al.
Publicado: (2024)
Sat-JEPA-Diff: Bridging Self-Supervised Learning and Generative Diffusion for Remote Sensing
por: Komurcu, Kursat, et al.
Publicado: (2026)
por: Komurcu, Kursat, et al.
Publicado: (2026)
ForAug: Recombining Foregrounds and Backgrounds to Improve Vision Transformer Training with Bias Mitigation
por: Nauen, Tobias Christian, et al.
Publicado: (2025)
por: Nauen, Tobias Christian, et al.
Publicado: (2025)
A Survey on Vision-Language-Action Models for Embodied AI
por: Ma, Yueen, et al.
Publicado: (2024)
por: Ma, Yueen, et al.
Publicado: (2024)
Surg$Σ$: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence
por: Zeng, Zhitao, et al.
Publicado: (2026)
por: Zeng, Zhitao, et al.
Publicado: (2026)
Learning Association via Track-Detection Matching for Multi-Object Tracking
por: Adžemović, Momir
Publicado: (2025)
por: Adžemović, Momir
Publicado: (2025)
Convolutional Model Trees
por: Armstrong, William Ward, et al.
Publicado: (2025)
por: Armstrong, William Ward, et al.
Publicado: (2025)
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
por: Fu, Tianyu, et al.
Publicado: (2024)
por: Fu, Tianyu, et al.
Publicado: (2024)
Akasha 2: Hamiltonian State Space Duality and Visual-Language Joint Embedding Predictive Architectur
por: Meziani, Yani
Publicado: (2026)
por: Meziani, Yani
Publicado: (2026)
Clarification as Supervision: Reinforcement Learning for Vision-Language Interfaces
por: Gkountouras, John, et al.
Publicado: (2025)
por: Gkountouras, John, et al.
Publicado: (2025)
Robust Noise Attenuation via Adaptive Pooling of Transformer Outputs
por: Brothers, Greyson
Publicado: (2025)
por: Brothers, Greyson
Publicado: (2025)
Complex Facial Expression Recognition Using Deep Knowledge Distillation of Basic Features
por: Maiden, Angus, et al.
Publicado: (2023)
por: Maiden, Angus, et al.
Publicado: (2023)
TACIT: Transformation-Aware Capturing of Implicit Thought
por: Nobrega, Daniel
Publicado: (2026)
por: Nobrega, Daniel
Publicado: (2026)
Ejemplares similares
-
Tricks and Plug-ins for Gradient Boosting in Image Classification
por: Fang, Biyi, et al.
Publicado: (2025) -
VA-$π$: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
por: Liao, Xinyao, et al.
Publicado: (2025) -
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
por: Chen, Kewei, et al.
Publicado: (2025) -
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
por: Chen, Kewei, et al.
Publicado: (2026) -
Unpacking Hateful Memes: Presupposed Context and False Claims
por: Cai, Weibin, et al.
Publicado: (2025)