ContPhy: Continuum Physical Concept Learning and Reasoning from Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Zhicheng, Yan, Xin, Chen, Zhenfang, Wang, Jingzhou, Lim, Qin Zhi Eddie, Tenenbaum, Joshua B., Gan, Chuang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
STAR: A Benchmark for Situated Reasoning in Real-World Videos
by: Wu, Bo, et al.
Published: (2024)
by: Wu, Bo, et al.
Published: (2024)
Compositional Physical Reasoning of Objects and Events from Videos
by: Chen, Zhenfang, et al.
Published: (2024)
by: Chen, Zhenfang, et al.
Published: (2024)
PhyScensis: Physics-Augmented LLM Agents for Complex Physical Scene Arrangement
by: Wang, Yian, et al.
Published: (2026)
by: Wang, Yian, et al.
Published: (2026)
Neuro-Symbolic Concepts
by: Mao, Jiayuan, et al.
Published: (2025)
by: Mao, Jiayuan, et al.
Published: (2025)
SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World Knowledge
by: Wang, Andong, et al.
Published: (2024)
by: Wang, Andong, et al.
Published: (2024)
TextPSG: Panoptic Scene Graph Generation from Textual Descriptions
by: Zhao, Chengyang, et al.
Published: (2023)
by: Zhao, Chengyang, et al.
Published: (2023)
VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction
by: Liang, Jiarong, et al.
Published: (2026)
by: Liang, Jiarong, et al.
Published: (2026)
VideoPhy: Evaluating Physical Commonsense for Video Generation
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
Physically Compatible 3D Object Modeling from a Single Image
by: Guo, Minghao, et al.
Published: (2024)
by: Guo, Minghao, et al.
Published: (2024)
"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
by: Gu, Jing, et al.
Published: (2025)
by: Gu, Jing, et al.
Published: (2025)
PhyRPR: Training-Free Physics-Constrained Video Generation
by: Zhao, Yibo, et al.
Published: (2026)
by: Zhao, Yibo, et al.
Published: (2026)
PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
by: Zhan, Yu-Wei, et al.
Published: (2025)
by: Zhan, Yu-Wei, et al.
Published: (2025)
PhyEduVideo: A Benchmark for Evaluating Text-to-Video Models for Physics Education
by: M, Megha Mariam K., et al.
Published: (2026)
by: M, Megha Mariam K., et al.
Published: (2026)
PhyGround: Benchmarking Physical Reasoning in Generative World Models
by: Lin, Juyi, et al.
Published: (2026)
by: Lin, Juyi, et al.
Published: (2026)
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
by: Bansal, Hritik, et al.
Published: (2025)
by: Bansal, Hritik, et al.
Published: (2025)
Building Cooperative Embodied Agents Modularly with Large Language Models
by: Zhang, Hongxin, et al.
Published: (2023)
by: Zhang, Hongxin, et al.
Published: (2023)
PhyWorld: Physics-Faithful World Model for Video Generation
by: Zhao, Pu, et al.
Published: (2026)
by: Zhao, Pu, et al.
Published: (2026)
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
by: Cai, Yuanhao, et al.
Published: (2025)
by: Cai, Yuanhao, et al.
Published: (2025)
HAZARD Challenge: Embodied Decision Making in Dynamically Changing Environments
by: Zhou, Qinhong, et al.
Published: (2024)
by: Zhou, Qinhong, et al.
Published: (2024)
FlexAttention for Efficient High-Resolution Vision-Language Models
by: Li, Junyan, et al.
Published: (2024)
by: Li, Junyan, et al.
Published: (2024)
ProPhy: Progressive Physical Alignment for Dynamic World Simulation
by: Wang, Zijun, et al.
Published: (2025)
by: Wang, Zijun, et al.
Published: (2025)
PhyPrompt: RL-based Prompt Refinement for Physically Plausible Text-to-Video Generation
by: Wu, Shang, et al.
Published: (2026)
by: Wu, Shang, et al.
Published: (2026)
PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI
by: Yang, Yandan, et al.
Published: (2024)
by: Yang, Yandan, et al.
Published: (2024)
PhyCo: Learning Controllable Physical Priors for Generative Motion
by: Narayanan, Sriram, et al.
Published: (2026)
by: Narayanan, Sriram, et al.
Published: (2026)
UniPhy: Learning a Unified Constitutive Model for Inverse Physics Simulation
by: Mittal, Himangi, et al.
Published: (2025)
by: Mittal, Himangi, et al.
Published: (2025)
VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models
by: Zhang, Zhicheng, et al.
Published: (2025)
by: Zhang, Zhicheng, et al.
Published: (2025)
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
by: Huang, Yidong, et al.
Published: (2026)
by: Huang, Yidong, et al.
Published: (2026)
VCA: Video Curious Agent for Long Video Understanding
by: Yang, Zeyuan, et al.
Published: (2024)
by: Yang, Zeyuan, et al.
Published: (2024)
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
by: Qin, Jialong, et al.
Published: (2025)
by: Qin, Jialong, et al.
Published: (2025)
Phy-CoSF: Physics-Guided Continuous Spectral Fields Reconstruction and Super-Resolution for Snapshot Compressive Imaging
by: Chen, Wudi, et al.
Published: (2026)
by: Chen, Wudi, et al.
Published: (2026)
M-PhyGs: Multi-Material Object Dynamics from Video
by: Wada, Norika, et al.
Published: (2025)
by: Wada, Norika, et al.
Published: (2025)
PhyCritic: Multimodal Critic Models for Physical AI
by: Xiong, Tianyi, et al.
Published: (2026)
by: Xiong, Tianyi, et al.
Published: (2026)
PhyRecon: Physically Plausible Neural Scene Reconstruction
by: Ni, Junfeng, et al.
Published: (2024)
by: Ni, Junfeng, et al.
Published: (2024)
PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation
by: Xue, Qiyao, et al.
Published: (2024)
by: Xue, Qiyao, et al.
Published: (2024)
PhiP-G: Physics-Guided Text-to-3D Compositional Scene Generation
by: Li, Qixuan, et al.
Published: (2025)
by: Li, Qixuan, et al.
Published: (2025)
Phy-Diff: Physics-guided Hourglass Diffusion Model for Diffusion MRI Synthesis
by: Zhang, Juanhua, et al.
Published: (2024)
by: Zhang, Juanhua, et al.
Published: (2024)
Learning Iterative Reasoning through Energy Diffusion
by: Du, Yilun, et al.
Published: (2024)
by: Du, Yilun, et al.
Published: (2024)
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models
by: Souza, Rafael, et al.
Published: (2024)
by: Souza, Rafael, et al.
Published: (2024)
PhyGaP: Physically-Grounded Gaussians with Polarization Cues
by: Wu, Jiale, et al.
Published: (2026)
by: Wu, Jiale, et al.
Published: (2026)
Potential Based Diffusion Motion Planning
by: Luo, Yunhao, et al.
Published: (2024)
by: Luo, Yunhao, et al.
Published: (2024)
Similar Items
-
STAR: A Benchmark for Situated Reasoning in Real-World Videos
by: Wu, Bo, et al.
Published: (2024) -
Compositional Physical Reasoning of Objects and Events from Videos
by: Chen, Zhenfang, et al.
Published: (2024) -
PhyScensis: Physics-Augmented LLM Agents for Complex Physical Scene Arrangement
by: Wang, Yian, et al.
Published: (2026) -
Neuro-Symbolic Concepts
by: Mao, Jiayuan, et al.
Published: (2025) -
SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World Knowledge
by: Wang, Andong, et al.
Published: (2024)