VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization
Fuente:
arXiv
Salvato in:
| Autori principali: | Cheng, Junhao, Hou, Liang, Zhong, Tianxiong, Tao, Xin, Wan, Pengfei, Gai, Kun, Liao, Jing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO
di: Cheng, Junhao, et al.
Pubblicazione: (2025)
di: Cheng, Junhao, et al.
Pubblicazione: (2025)
Scaling Image and Video Generation via Test-Time Evolutionary Search
di: He, Haoran, et al.
Pubblicazione: (2025)
di: He, Haoran, et al.
Pubblicazione: (2025)
VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information Assumption
di: Zhong, Tianxiong, et al.
Pubblicazione: (2025)
di: Zhong, Tianxiong, et al.
Pubblicazione: (2025)
Decoupling Complexity from Scale in Latent Diffusion Model
di: Zhong, Tianxiong, et al.
Pubblicazione: (2025)
di: Zhong, Tianxiong, et al.
Pubblicazione: (2025)
MTV-Inpaint: Multi-Task Long Video Inpainting
di: Yang, Shiyuan, et al.
Pubblicazione: (2025)
di: Yang, Shiyuan, et al.
Pubblicazione: (2025)
Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional Encodings
di: Hou, Liang, et al.
Pubblicazione: (2025)
di: Hou, Liang, et al.
Pubblicazione: (2025)
ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning
di: Huang, Yuzhou, et al.
Pubblicazione: (2025)
di: Huang, Yuzhou, et al.
Pubblicazione: (2025)
Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
di: Cheng, Junhao, et al.
Pubblicazione: (2025)
di: Cheng, Junhao, et al.
Pubblicazione: (2025)
3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation
di: Fang, Zhixue, et al.
Pubblicazione: (2026)
di: Fang, Zhixue, et al.
Pubblicazione: (2026)
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing
di: Hu, Jiahao, et al.
Pubblicazione: (2024)
di: Hu, Jiahao, et al.
Pubblicazione: (2024)
GridProbe: Posterior-Probing for Adaptive Test-Time Compute in Long-Video VLMs
di: Eltahir, Mohamed, et al.
Pubblicazione: (2026)
di: Eltahir, Mohamed, et al.
Pubblicazione: (2026)
Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion
di: Yang, Shiyuan, et al.
Pubblicazione: (2024)
di: Yang, Shiyuan, et al.
Pubblicazione: (2024)
A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
di: Wu, Shengqiong, et al.
Pubblicazione: (2025)
di: Wu, Shengqiong, et al.
Pubblicazione: (2025)
UniVideo: Unified Understanding, Generation, and Editing for Videos
di: Wei, Cong, et al.
Pubblicazione: (2025)
di: Wei, Cong, et al.
Pubblicazione: (2025)
MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives
di: Ji, Sihui, et al.
Pubblicazione: (2025)
di: Ji, Sihui, et al.
Pubblicazione: (2025)
Can Vision-Language Models be a Good Guesser? Exploring VLMs for Times and Location Reasoning
di: Zhang, Gengyuan, et al.
Pubblicazione: (2023)
di: Zhang, Gengyuan, et al.
Pubblicazione: (2023)
VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning
di: Cai, Minghong, et al.
Pubblicazione: (2025)
di: Cai, Minghong, et al.
Pubblicazione: (2025)
PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning
di: Ji, Sihui, et al.
Pubblicazione: (2025)
di: Ji, Sihui, et al.
Pubblicazione: (2025)
Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
di: Zhang, Xintong, et al.
Pubblicazione: (2025)
di: Zhang, Xintong, et al.
Pubblicazione: (2025)
VMoBA: Mixture-of-Block Attention for Video Diffusion Models
di: Wu, Jianzong, et al.
Pubblicazione: (2025)
di: Wu, Jianzong, et al.
Pubblicazione: (2025)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
di: Li, Weiming, et al.
Pubblicazione: (2025)
di: Li, Weiming, et al.
Pubblicazione: (2025)
Stable Velocity: A Variance Perspective on Flow Matching
di: Yang, Donglin, et al.
Pubblicazione: (2026)
di: Yang, Donglin, et al.
Pubblicazione: (2026)
Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards
di: Gambashidze, Alexander, et al.
Pubblicazione: (2025)
di: Gambashidze, Alexander, et al.
Pubblicazione: (2025)
Monet: Reasoning in Latent Visual Space Beyond Images and Language
di: Wang, Qixun, et al.
Pubblicazione: (2025)
di: Wang, Qixun, et al.
Pubblicazione: (2025)
Mitigating the Noise Shift for Denoising Generative Models via Noise Awareness Guidance
di: Zhong, Jincheng, et al.
Pubblicazione: (2025)
di: Zhong, Jincheng, et al.
Pubblicazione: (2025)
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
di: Yao, Linli, et al.
Pubblicazione: (2026)
di: Yao, Linli, et al.
Pubblicazione: (2026)
RelightMaster: Precise Video Relighting with Multi-plane Light Images
di: Bian, Weikang, et al.
Pubblicazione: (2025)
di: Bian, Weikang, et al.
Pubblicazione: (2025)
TIR-Flow: Active Video Search and Reasoning with Frozen VLMs
di: Jin, Hongbo, et al.
Pubblicazione: (2026)
di: Jin, Hongbo, et al.
Pubblicazione: (2026)
FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
di: Ju, Xuan, et al.
Pubblicazione: (2025)
di: Ju, Xuan, et al.
Pubblicazione: (2025)
CamCloneMaster: Enabling Reference-based Camera Control for Video Generation
di: Luo, Yawen, et al.
Pubblicazione: (2025)
di: Luo, Yawen, et al.
Pubblicazione: (2025)
VRIQ: Benchmarking and Analyzing Visual-Reasoning IQ of VLMs
di: Khezresmaeilzadeh, Tina, et al.
Pubblicazione: (2026)
di: Khezresmaeilzadeh, Tina, et al.
Pubblicazione: (2026)
A Survey of Interactive Generative Video
di: Yu, Jiwen, et al.
Pubblicazione: (2025)
di: Yu, Jiwen, et al.
Pubblicazione: (2025)
UNIC: Unified In-Context Video Editing
di: Ye, Zixuan, et al.
Pubblicazione: (2025)
di: Ye, Zixuan, et al.
Pubblicazione: (2025)
VINO: A Unified Visual Generator with Interleaved OmniModal Context
di: Chen, Junyi, et al.
Pubblicazione: (2026)
di: Chen, Junyi, et al.
Pubblicazione: (2026)
Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling
di: Wang, Yuan, et al.
Pubblicazione: (2026)
di: Wang, Yuan, et al.
Pubblicazione: (2026)
VideoPro: Adaptive Program Reasoning for Long Video Understanding
di: Li, Chenglin, et al.
Pubblicazione: (2025)
di: Li, Chenglin, et al.
Pubblicazione: (2025)
Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks
di: Yang, Cheng, et al.
Pubblicazione: (2025)
di: Yang, Cheng, et al.
Pubblicazione: (2025)
Imbalance in Balance: Online Concept Balancing in Generation Models
di: Shi, Yukai, et al.
Pubblicazione: (2025)
di: Shi, Yukai, et al.
Pubblicazione: (2025)
ANTS: Adaptive Negative Textual Space Shaping for OOD Detection via Test-Time MLLM Understanding and Reasoning
di: Zhu, Wenjie, et al.
Pubblicazione: (2025)
di: Zhu, Wenjie, et al.
Pubblicazione: (2025)
AUTO: Adaptive Outlier Optimization for Test-Time OOD Detection
di: Yang, Puning, et al.
Pubblicazione: (2023)
di: Yang, Puning, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO
di: Cheng, Junhao, et al.
Pubblicazione: (2025) -
Scaling Image and Video Generation via Test-Time Evolutionary Search
di: He, Haoran, et al.
Pubblicazione: (2025) -
VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information Assumption
di: Zhong, Tianxiong, et al.
Pubblicazione: (2025) -
Decoupling Complexity from Scale in Latent Diffusion Model
di: Zhong, Tianxiong, et al.
Pubblicazione: (2025) -
MTV-Inpaint: Multi-Task Long Video Inpainting
di: Yang, Shiyuan, et al.
Pubblicazione: (2025)