VideoPhy: Evaluating Physical Commonsense for Video Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Bansal, Hritik, Lin, Zongyu, Xie, Tianyi, Zong, Zeshun, Yarom, Michal, Bitton, Yonatan, Jiang, Chenfanfu, Sun, Yizhou, Chang, Kai-Wei, Grover, Aditya |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
di: Bansal, Hritik, et al.
Pubblicazione: (2025)
di: Bansal, Hritik, et al.
Pubblicazione: (2025)
TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
di: Bansal, Hritik, et al.
Pubblicazione: (2024)
di: Bansal, Hritik, et al.
Pubblicazione: (2024)
Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis
di: Ramos, Vasco, et al.
Pubblicazione: (2024)
di: Ramos, Vasco, et al.
Pubblicazione: (2024)
A Convex Formulation of Frictional Contact for the Material Point Method and Rigid Bodies
di: Zong, Zeshun, et al.
Pubblicazione: (2024)
di: Zong, Zeshun, et al.
Pubblicazione: (2024)
Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models
di: Bansal, Hritik, et al.
Pubblicazione: (2023)
di: Bansal, Hritik, et al.
Pubblicazione: (2023)
PhysMotion: Physics-Grounded Dynamics From a Single Image
di: Tan, Xiyang, et al.
Pubblicazione: (2024)
di: Tan, Xiyang, et al.
Pubblicazione: (2024)
PhysGaussian: Physics-Integrated 3D Gaussians for Generative Dynamics
di: Xie, Tianyi, et al.
Pubblicazione: (2023)
di: Xie, Tianyi, et al.
Pubblicazione: (2023)
Atlas3D: Physically Constrained Self-Supporting Text-to-3D for Simulation and Fabrication
di: Chen, Yunuo, et al.
Pubblicazione: (2024)
di: Chen, Yunuo, et al.
Pubblicazione: (2024)
A Convex Formulation of Material Points and Rigid Bodies with GPU-Accelerated Async-Coupling for Interactive Simulation
di: Yu, Chang, et al.
Pubblicazione: (2025)
di: Yu, Chang, et al.
Pubblicazione: (2025)
Comparing Bad Apples to Good Oranges: Aligning Large Language Models via Joint Preference Optimization
di: Bansal, Hritik, et al.
Pubblicazione: (2024)
di: Bansal, Hritik, et al.
Pubblicazione: (2024)
PhyEduVideo: A Benchmark for Evaluating Text-to-Video Models for Physics Education
di: M, Megha Mariam K., et al.
Pubblicazione: (2026)
di: M, Megha Mariam K., et al.
Pubblicazione: (2026)
GRIP: A General Robotic Incremental Potential Contact Simulation Dataset for Unified Deformable-Rigid Coupled Grasping
di: Ma, Siyu, et al.
Pubblicazione: (2025)
di: Ma, Siyu, et al.
Pubblicazione: (2025)
When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning
di: Singhi, Nishad, et al.
Pubblicazione: (2025)
di: Singhi, Nishad, et al.
Pubblicazione: (2025)
Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators
di: Bansal, Hritik, et al.
Pubblicazione: (2025)
di: Bansal, Hritik, et al.
Pubblicazione: (2025)
PhysAnimator: Physics-Guided Generative Cartoon Animation
di: Xie, Tianyi, et al.
Pubblicazione: (2025)
di: Xie, Tianyi, et al.
Pubblicazione: (2025)
ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models
di: Wadhawan, Rohan, et al.
Pubblicazione: (2024)
di: Wadhawan, Rohan, et al.
Pubblicazione: (2024)
MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants
di: Bansal, Hritik, et al.
Pubblicazione: (2024)
di: Bansal, Hritik, et al.
Pubblicazione: (2024)
Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
di: Zohar, Orr, et al.
Pubblicazione: (2024)
di: Zohar, Orr, et al.
Pubblicazione: (2024)
Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models
di: Bitton-Guetta, Nitzan, et al.
Pubblicazione: (2024)
di: Bitton-Guetta, Nitzan, et al.
Pubblicazione: (2024)
PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models
di: Meng, Fanqing, et al.
Pubblicazione: (2024)
di: Meng, Fanqing, et al.
Pubblicazione: (2024)
HoneyBee: Data Recipes for Vision-Language Reasoners
di: Bansal, Hritik, et al.
Pubblicazione: (2025)
di: Bansal, Hritik, et al.
Pubblicazione: (2025)
Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks
di: Bordalo, João, et al.
Pubblicazione: (2024)
di: Bordalo, João, et al.
Pubblicazione: (2024)
Survey of Bias In Text-to-Image Generation: Definition, Evaluation, and Mitigation
di: Wan, Yixin, et al.
Pubblicazione: (2024)
di: Wan, Yixin, et al.
Pubblicazione: (2024)
AnimaMimic: Imitating 3D Animation from Video Priors
di: Xie, Tianyi, et al.
Pubblicazione: (2025)
di: Xie, Tianyi, et al.
Pubblicazione: (2025)
Embedded IPC: Fast and Intersection-free Simulation in Reduced Subspace for Robot Manipulation
di: Du, Wenxin, et al.
Pubblicazione: (2024)
di: Du, Wenxin, et al.
Pubblicazione: (2024)
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
di: Cao, Meng, et al.
Pubblicazione: (2024)
di: Cao, Meng, et al.
Pubblicazione: (2024)
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
di: Huang, Yidong, et al.
Pubblicazione: (2026)
di: Huang, Yidong, et al.
Pubblicazione: (2026)
LaViDa: A Large Diffusion Language Model for Multimodal Understanding
di: Li, Shufan, et al.
Pubblicazione: (2025)
di: Li, Shufan, et al.
Pubblicazione: (2025)
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits
di: Yosef, Ron, et al.
Pubblicazione: (2025)
di: Yosef, Ron, et al.
Pubblicazione: (2025)
"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
di: Gu, Jing, et al.
Pubblicazione: (2025)
di: Gu, Jing, et al.
Pubblicazione: (2025)
Gaussian Splashing: Unified Particles for Versatile Motion Synthesis and Rendering
di: Feng, Yutao, et al.
Pubblicazione: (2024)
di: Feng, Yutao, et al.
Pubblicazione: (2024)
SparseCL: Sparse Contrastive Learning for Contradiction Retrieval
di: Xu, Haike, et al.
Pubblicazione: (2024)
di: Xu, Haike, et al.
Pubblicazione: (2024)
PhyWorld: Physics-Faithful World Model for Video Generation
di: Zhao, Pu, et al.
Pubblicazione: (2026)
di: Zhao, Pu, et al.
Pubblicazione: (2026)
PhyRPR: Training-Free Physics-Constrained Video Generation
di: Zhao, Yibo, et al.
Pubblicazione: (2026)
di: Zhao, Yibo, et al.
Pubblicazione: (2026)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
di: Zhang, Yue, et al.
Pubblicazione: (2026)
di: Zhang, Yue, et al.
Pubblicazione: (2026)
Dress-1-to-3: Single Image to Simulation-Ready 3D Outfit with Diffusion Prior and Differentiable Physics
di: Li, Xuan, et al.
Pubblicazione: (2025)
di: Li, Xuan, et al.
Pubblicazione: (2025)
Substepping the Material Point Method
di: Jiang, Chenfanfu
Pubblicazione: (2025)
di: Jiang, Chenfanfu
Pubblicazione: (2025)
ContPhy: Continuum Physical Concept Learning and Reasoning from Videos
di: Zheng, Zhicheng, et al.
Pubblicazione: (2024)
di: Zheng, Zhicheng, et al.
Pubblicazione: (2024)
Scaling transformer neural networks for skillful and reliable medium-range weather forecasting
di: Nguyen, Tung, et al.
Pubblicazione: (2023)
di: Nguyen, Tung, et al.
Pubblicazione: (2023)
RefVNLI: Towards Scalable Evaluation of Subject-driven Text-to-image Generation
di: Slobodkin, Aviv, et al.
Pubblicazione: (2025)
di: Slobodkin, Aviv, et al.
Pubblicazione: (2025)
Documenti analoghi
-
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
di: Bansal, Hritik, et al.
Pubblicazione: (2025) -
TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
di: Bansal, Hritik, et al.
Pubblicazione: (2024) -
Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis
di: Ramos, Vasco, et al.
Pubblicazione: (2024) -
A Convex Formulation of Frictional Contact for the Material Point Method and Rigid Bodies
di: Zong, Zeshun, et al.
Pubblicazione: (2024) -
Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models
di: Bansal, Hritik, et al.
Pubblicazione: (2023)