Building a Precise Video Language with Human-AI Oversight

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lin, Zhiqiu, Mitra, Chancharik, Cen, Siyuan, Li, Isaac, Huang, Yuhan, Ling, Yu Tong Tiffany, Wang, Hewei, Pi, Irene, Zhu, Shihang, Rao, Ryan, Liu, George, Li, Jiaxi, Li, Ruojin, Han, Yili, Du, Yilun, Ramanan, Deva
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911625076801536
author Lin, Zhiqiu
Mitra, Chancharik
Cen, Siyuan
Li, Isaac
Huang, Yuhan
Ling, Yu Tong Tiffany
Wang, Hewei
Pi, Irene
Zhu, Shihang
Rao, Ryan
Liu, George
Li, Jiaxi
Li, Ruojin
Han, Yili
Du, Yilun
Ramanan, Deva
author_facet Lin, Zhiqiu
Mitra, Chancharik
Cen, Siyuan
Li, Isaac
Huang, Yuhan
Ling, Yu Tong Tiffany
Wang, Hewei
Pi, Irene
Zhu, Shihang
Rao, Ryan
Liu, George
Li, Jiaxi
Li, Ruojin
Han, Yili
Du, Yilun
Ramanan, Deva
contents Video-language models (VLMs) learn to reason about the dynamic visual world through natural language. We introduce a suite of open datasets, benchmarks, and recipes for scalable oversight that enable precise video captioning. First, we define a structured specification for describing subjects, scenes, motion, spatial, and camera dynamics, grounded by hundreds of carefully defined visual primitives developed with professional video creators such as filmmakers. Next, to curate high-quality captions, we introduce CHAI (Critique-based Human-AI Oversight), a framework where trained experts critique and revise model-generated pre-captions into improved post-captions. This division of labor improves annotation accuracy and efficiency by offloading text generation to models, allowing humans to better focus on verification. Additionally, these critiques and preferences between pre- and post-captions provide rich supervision for improving open-source models (Qwen3-VL) on caption generation, reward modeling, and critique generation through SFT, DPO, and inference-time scaling. Our ablations show that critique quality in precision, recall, and constructiveness, ensured by our oversight framework, directly governs downstream performance. With modest expert supervision, the resulting model outperforms closed-source models such as Gemini-3.1-Pro. Finally, we apply our approach to re-caption large-scale professional videos (e.g., films, commercials, games) and fine-tune video generation models such as Wan to better follow detailed prompts of up to 400 words, achieving finer control over cinematography including camera motion, angle, lens, focus, point of view, and framing. Our results show that precise specification and human-AI oversight are key to professional-level video understanding and generation. Data and code are available on our project page: https://linzhiqiu.github.io/papers/chai/
format Preprint
id arxiv_https___arxiv_org_abs_2604_21718
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Building a Precise Video Language with Human-AI Oversight
Lin, Zhiqiu
Mitra, Chancharik
Cen, Siyuan
Li, Isaac
Huang, Yuhan
Ling, Yu Tong Tiffany
Wang, Hewei
Pi, Irene
Zhu, Shihang
Rao, Ryan
Liu, George
Li, Jiaxi
Li, Ruojin
Han, Yili
Du, Yilun
Ramanan, Deva
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Multimedia
Video-language models (VLMs) learn to reason about the dynamic visual world through natural language. We introduce a suite of open datasets, benchmarks, and recipes for scalable oversight that enable precise video captioning. First, we define a structured specification for describing subjects, scenes, motion, spatial, and camera dynamics, grounded by hundreds of carefully defined visual primitives developed with professional video creators such as filmmakers. Next, to curate high-quality captions, we introduce CHAI (Critique-based Human-AI Oversight), a framework where trained experts critique and revise model-generated pre-captions into improved post-captions. This division of labor improves annotation accuracy and efficiency by offloading text generation to models, allowing humans to better focus on verification. Additionally, these critiques and preferences between pre- and post-captions provide rich supervision for improving open-source models (Qwen3-VL) on caption generation, reward modeling, and critique generation through SFT, DPO, and inference-time scaling. Our ablations show that critique quality in precision, recall, and constructiveness, ensured by our oversight framework, directly governs downstream performance. With modest expert supervision, the resulting model outperforms closed-source models such as Gemini-3.1-Pro. Finally, we apply our approach to re-caption large-scale professional videos (e.g., films, commercials, games) and fine-tune video generation models such as Wan to better follow detailed prompts of up to 400 words, achieving finer control over cinematography including camera motion, angle, lens, focus, point of view, and framing. Our results show that precise specification and human-AI oversight are key to professional-level video understanding and generation. Data and code are available on our project page: https://linzhiqiu.github.io/papers/chai/
title Building a Precise Video Language with Human-AI Oversight
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Multimedia
url https://arxiv.org/abs/2604.21718