Physics Knowledge in Frontier Models: A Diagnostic Study of Failure Modes
Fuente:
arXiv
Saved in:
| Main Authors: | Bagdonaviciute, Ieva, Vineet, Vibhav |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Navigating Hallucinations for Reasoning of Unintentional Activities
by: Grover, Shresth, et al.
Published: (2024)
by: Grover, Shresth, et al.
Published: (2024)
HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
by: Azad, Shehreen, et al.
Published: (2025)
by: Azad, Shehreen, et al.
Published: (2025)
StreamReady: Learning What to Answer and When in Long Streaming Videos
by: Azad, Shehreen, et al.
Published: (2026)
by: Azad, Shehreen, et al.
Published: (2026)
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
by: Kumar, Akash, et al.
Published: (2025)
by: Kumar, Akash, et al.
Published: (2025)
OmViD: Omni-supervised active learning for video action detection
by: Rana, Aayush, et al.
Published: (2025)
by: Rana, Aayush, et al.
Published: (2025)
Understanding Depth and Height Perception in Large Visual-Language Models
by: Azad, Shehreen, et al.
Published: (2024)
by: Azad, Shehreen, et al.
Published: (2024)
On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes
by: Modi, Rajat, et al.
Published: (2024)
by: Modi, Rajat, et al.
Published: (2024)
CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
by: Grover, Shresth, et al.
Published: (2025)
by: Grover, Shresth, et al.
Published: (2025)
PEEKABOO: Interactive Video Generation via Masked-Diffusion
by: Jain, Yash, et al.
Published: (2023)
by: Jain, Yash, et al.
Published: (2023)
Robustness Analysis on Foundational Segmentation Models
by: Schiappa, Madeline Chantry, et al.
Published: (2023)
by: Schiappa, Madeline Chantry, et al.
Published: (2023)
Grounding Task Assistance with Multimodal Cues from a Single Demonstration
by: Sarch, Gabriel, et al.
Published: (2025)
by: Sarch, Gabriel, et al.
Published: (2025)
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames
by: Ravi, Sahithya, et al.
Published: (2025)
by: Ravi, Sahithya, et al.
Published: (2025)
Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models
by: Wang, Jiayu, et al.
Published: (2024)
by: Wang, Jiayu, et al.
Published: (2024)
OS-Marathon: Benchmarking Computer-Use Agents on Long-Horizon Repetitive Tasks
by: Wu, Jing, et al.
Published: (2026)
by: Wu, Jing, et al.
Published: (2026)
Simplifying Knowledge Transfer in Pretrained Models
by: Jain, Siddharth, et al.
Published: (2025)
by: Jain, Siddharth, et al.
Published: (2025)
PRIME: Prioritizing Interpretability in Failure Mode Extraction
by: Rezaei, Keivan, et al.
Published: (2023)
by: Rezaei, Keivan, et al.
Published: (2023)
Can Video Large Multimodal Models Think Like Doubters-or Double-Down: A Study on Defeasible Video Entailment
by: Zhang, Yue, et al.
Published: (2025)
by: Zhang, Yue, et al.
Published: (2025)
FineGRAIN: Evaluating Failure Modes of Text-to-Image Models with Vision Language Model Judges
by: Hayes, Kevin David, et al.
Published: (2025)
by: Hayes, Kevin David, et al.
Published: (2025)
DreamDistribution: Learning Prompt Distribution for Diverse In-distribution Generation
by: Zhao, Brian Nlong, et al.
Published: (2023)
by: Zhao, Brian Nlong, et al.
Published: (2023)
A Study of Failure Modes in Two-Stage Human-Object Interaction Detection
by: Wang, Lemeng, et al.
Published: (2026)
by: Wang, Lemeng, et al.
Published: (2026)
MM-GEN: Enhancing Task Performance Through Targeted Multimodal Data Curation
by: Joshi, Siddharth, et al.
Published: (2025)
by: Joshi, Siddharth, et al.
Published: (2025)
More Images, More Problems? A Controlled Analysis of VLM Failure Modes
by: Das, Anurag, et al.
Published: (2026)
by: Das, Anurag, et al.
Published: (2026)
Discovering Failure Modes in Vision-Language Models using RL
by: Jain, Kanishk, et al.
Published: (2026)
by: Jain, Kanishk, et al.
Published: (2026)
How does My Model Fail? Automatic Identification and Interpretation of Physical Plausibility Failure Modes with Matryoshka Transcoders
by: Tang, Yiming, et al.
Published: (2025)
by: Tang, Yiming, et al.
Published: (2025)
Scanner-Agnostic MRI Harmonization via SSIM-Guided Disentanglement
by: Caldera, Luca, et al.
Published: (2025)
by: Caldera, Luca, et al.
Published: (2025)
Understanding the Failure Modes of Out-of-Distribution Generalization
by: Nagarajan, Vaishnavh, et al.
Published: (2020)
by: Nagarajan, Vaishnavh, et al.
Published: (2020)
TIDE: Training Locally Interpretable Domain Generalization Models Enables Test-time Correction
by: Agarwal, Aishwarya, et al.
Published: (2024)
by: Agarwal, Aishwarya, et al.
Published: (2024)
Images Speak Louder Than Scores: Failure Mode Escape for Enhancing Generative Quality
by: Shao, Jie, et al.
Published: (2025)
by: Shao, Jie, et al.
Published: (2025)
WorldBench: Disambiguating Physics for Diagnostic Evaluation of World Models
by: Upadhyay, Rishi, et al.
Published: (2026)
by: Upadhyay, Rishi, et al.
Published: (2026)
Unearthing Skill-Level Insights for Understanding Trade-Offs of Foundation Models
by: Moayeri, Mazda, et al.
Published: (2024)
by: Moayeri, Mazda, et al.
Published: (2024)
Towards Unbiased and Robust Spatio-Temporal Scene Graph Generation and Anticipation
by: Peddi, Rohith, et al.
Published: (2024)
by: Peddi, Rohith, et al.
Published: (2024)
One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations
by: Balakrishnan, Ravikumar, et al.
Published: (2026)
by: Balakrishnan, Ravikumar, et al.
Published: (2026)
PLUTO-4: Frontier Pathology Foundation Models
by: Padigela, Harshith, et al.
Published: (2025)
by: Padigela, Harshith, et al.
Published: (2025)
NVILA: Efficient Frontier Visual Language Models
by: Liu, Zhijian, et al.
Published: (2024)
by: Liu, Zhijian, et al.
Published: (2024)
Maximally Useful and Minimally Redundant: The Key to Self Supervised Learning for Imbalanced Data
by: Sharma, Yash Kumar, et al.
Published: (2025)
by: Sharma, Yash Kumar, et al.
Published: (2025)
Fara-7B: An Efficient Agentic Model for Computer Use
by: Awadallah, Ahmed, et al.
Published: (2025)
by: Awadallah, Ahmed, et al.
Published: (2025)
Advancing Egocentric Video Question Answering with Multimodal Large Language Models
by: Patel, Alkesh, et al.
Published: (2025)
by: Patel, Alkesh, et al.
Published: (2025)
A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding
by: Zhang, Yue, et al.
Published: (2026)
by: Zhang, Yue, et al.
Published: (2026)
FoundBioNet: A Foundation-Based Model for IDH Genotyping of Glioma from Multi-Parametric MRI
by: Farahani, Somayeh, et al.
Published: (2025)
by: Farahani, Somayeh, et al.
Published: (2025)
AMVICC: A Novel Benchmark for Cross-Modal Failure Mode Profiling for VLMs and IGMs
by: Basappa, Aahana, et al.
Published: (2026)
by: Basappa, Aahana, et al.
Published: (2026)
Similar Items
-
Navigating Hallucinations for Reasoning of Unintentional Activities
by: Grover, Shresth, et al.
Published: (2024) -
HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
by: Azad, Shehreen, et al.
Published: (2025) -
StreamReady: Learning What to Answer and When in Long Streaming Videos
by: Azad, Shehreen, et al.
Published: (2026) -
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
by: Kumar, Akash, et al.
Published: (2025) -
OmViD: Omni-supervised active learning for video action detection
by: Rana, Aayush, et al.
Published: (2025)