PhyCritic: Multimodal Critic Models for Physical AI
Fuente:
arXiv
Guardado en:
| Autores principales: | Xiong, Tianyi, Wang, Shihao, Liu, Guilin, Dong, Yi, Li, Ming, Huang, Heng, Kautz, Jan, Yu, Zhiding |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LLaVA-Critic: Learning to Evaluate Multimodal Models
por: Xiong, Tianyi, et al.
Publicado: (2024)
por: Xiong, Tianyi, et al.
Publicado: (2024)
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
por: Man, Yunze, et al.
Publicado: (2025)
por: Man, Yunze, et al.
Publicado: (2025)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
por: Wang, Shihao, et al.
Publicado: (2025)
por: Wang, Shihao, et al.
Publicado: (2025)
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
por: Huang, De-An, et al.
Publicado: (2025)
por: Huang, De-An, et al.
Publicado: (2025)
Slow-Fast Architecture for Video Multi-Modal Large Language Models
por: Shi, Min, et al.
Publicado: (2025)
por: Shi, Min, et al.
Publicado: (2025)
StreamChat: Chatting with Streaming Video
por: Liu, Jihao, et al.
Publicado: (2024)
por: Liu, Jihao, et al.
Publicado: (2024)
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
por: Wang, Xiyao, et al.
Publicado: (2025)
por: Wang, Xiyao, et al.
Publicado: (2025)
STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
por: Jiang, Jindong, et al.
Publicado: (2025)
por: Jiang, Jindong, et al.
Publicado: (2025)
Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation
por: Li, Zhenxin, et al.
Publicado: (2024)
por: Li, Zhenxin, et al.
Publicado: (2024)
LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
por: Man, Yunze, et al.
Publicado: (2025)
por: Man, Yunze, et al.
Publicado: (2025)
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
por: Zhao, Yue, et al.
Publicado: (2025)
por: Zhao, Yue, et al.
Publicado: (2025)
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
por: Wang, Shihao, et al.
Publicado: (2025)
por: Wang, Shihao, et al.
Publicado: (2025)
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
por: Wang, Shihao, et al.
Publicado: (2024)
por: Wang, Shihao, et al.
Publicado: (2024)
Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
por: Shi, Min, et al.
Publicado: (2024)
por: Shi, Min, et al.
Publicado: (2024)
Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline
por: Chen, Guo, et al.
Publicado: (2026)
por: Chen, Guo, et al.
Publicado: (2026)
Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?
por: Li, Zhiqi, et al.
Publicado: (2023)
por: Li, Zhiqi, et al.
Publicado: (2023)
The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment
por: Ouyang, Ziheng, et al.
Publicado: (2025)
por: Ouyang, Ziheng, et al.
Publicado: (2025)
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
por: Wang, Shihao, et al.
Publicado: (2026)
por: Wang, Shihao, et al.
Publicado: (2026)
Stateful Token Reduction for Long-Video Hybrid VLMs
por: Jiang, Jindong, et al.
Publicado: (2026)
por: Jiang, Jindong, et al.
Publicado: (2026)
Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models
por: Chen, Guo, et al.
Publicado: (2025)
por: Chen, Guo, et al.
Publicado: (2025)
LITA: Language Instructed Temporal-Localization Assistant
por: Huang, De-An, et al.
Publicado: (2024)
por: Huang, De-An, et al.
Publicado: (2024)
"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
por: Gu, Jing, et al.
Publicado: (2025)
por: Gu, Jing, et al.
Publicado: (2025)
PhyRecon: Physically Plausible Neural Scene Reconstruction
por: Ni, Junfeng, et al.
Publicado: (2024)
por: Ni, Junfeng, et al.
Publicado: (2024)
AIDE: Agentically Improve Visual Language Model with Domain Experts
por: Chiu, Ming-Chang, et al.
Publicado: (2025)
por: Chiu, Ming-Chang, et al.
Publicado: (2025)
DiffiT: Diffusion Vision Transformers for Image Generation
por: Hatamizadeh, Ali, et al.
Publicado: (2023)
por: Hatamizadeh, Ali, et al.
Publicado: (2023)
PhyTracker: An Online Tracker for Phytoplankton
por: Yu, Yang, et al.
Publicado: (2024)
por: Yu, Yang, et al.
Publicado: (2024)
PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments
por: Zhou, Weijie, et al.
Publicado: (2025)
por: Zhou, Weijie, et al.
Publicado: (2025)
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
por: Zhang, Di, et al.
Publicado: (2024)
por: Zhang, Di, et al.
Publicado: (2024)
PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models
por: Meng, Fanqing, et al.
Publicado: (2024)
por: Meng, Fanqing, et al.
Publicado: (2024)
Hydra-NeXt: Robust Closed-Loop Driving with Open-Loop Training
por: Li, Zhenxin, et al.
Publicado: (2025)
por: Li, Zhenxin, et al.
Publicado: (2025)
PhyCustom: Towards Realistic Physical Customization in Text-to-Image Generation
por: Wu, Fan, et al.
Publicado: (2025)
por: Wu, Fan, et al.
Publicado: (2025)
PhyDAE: Physics-Guided Degradation-Adaptive Experts for All-in-One Remote Sensing Image Restoration
por: Dong, Zhe, et al.
Publicado: (2025)
por: Dong, Zhe, et al.
Publicado: (2025)
PhyRPR: Training-Free Physics-Constrained Video Generation
por: Zhao, Yibo, et al.
Publicado: (2026)
por: Zhao, Yibo, et al.
Publicado: (2026)
PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
por: Zhan, Yu-Wei, et al.
Publicado: (2025)
por: Zhan, Yu-Wei, et al.
Publicado: (2025)
Modality-Balancing Preference Optimization of Large Multimodal Models by Adversarial Negative Mining
por: Liu, Chenxi, et al.
Publicado: (2025)
por: Liu, Chenxi, et al.
Publicado: (2025)
Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following
por: Xiong, Tianyi, et al.
Publicado: (2025)
por: Xiong, Tianyi, et al.
Publicado: (2025)
ProPhy: Progressive Physical Alignment for Dynamic World Simulation
por: Wang, Zijun, et al.
Publicado: (2025)
por: Wang, Zijun, et al.
Publicado: (2025)
UniPhy: Learning a Unified Constitutive Model for Inverse Physics Simulation
por: Mittal, Himangi, et al.
Publicado: (2025)
por: Mittal, Himangi, et al.
Publicado: (2025)
PhyMAGIC: Physical Motion-Aware Generative Inference with Confidence-guided LLM
por: Meng, Siwei, et al.
Publicado: (2025)
por: Meng, Siwei, et al.
Publicado: (2025)
VideoPhy: Evaluating Physical Commonsense for Video Generation
por: Bansal, Hritik, et al.
Publicado: (2024)
por: Bansal, Hritik, et al.
Publicado: (2024)
Ejemplares similares
-
LLaVA-Critic: Learning to Evaluate Multimodal Models
por: Xiong, Tianyi, et al.
Publicado: (2024) -
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
por: Man, Yunze, et al.
Publicado: (2025) -
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
por: Wang, Shihao, et al.
Publicado: (2025) -
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
por: Huang, De-An, et al.
Publicado: (2025) -
Slow-Fast Architecture for Video Multi-Modal Large Language Models
por: Shi, Min, et al.
Publicado: (2025)