Building a Precise Video Language with Human-AI Oversight
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Zhiqiu, Mitra, Chancharik, Cen, Siyuan, Li, Isaac, Huang, Yuhan, Ling, Yu Tong Tiffany, Wang, Hewei, Pi, Irene, Zhu, Shihang, Rao, Ryan, Liu, George, Li, Jiaxi, Li, Ruojin, Han, Yili, Du, Yilun, Ramanan, Deva |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Understanding Camera Motions in Any Video
by: Lin, Zhiqiu, et al.
Published: (2025)
by: Lin, Zhiqiu, et al.
Published: (2025)
Language Models as Black-Box Optimizers for Vision-Language Models
by: Liu, Shihong, et al.
Published: (2023)
by: Liu, Shihong, et al.
Published: (2023)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
by: Lin, Zhiqiu, et al.
Published: (2024)
by: Lin, Zhiqiu, et al.
Published: (2024)
GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation
by: Li, Baiqi, et al.
Published: (2024)
by: Li, Baiqi, et al.
Published: (2024)
ResearchPulse: Building Method-Experiment Chains through Multi-Document Scientific Inference
by: Chen, Qi, et al.
Published: (2025)
by: Chen, Qi, et al.
Published: (2025)
Dependency Structure Augmented Contextual Scoping Framework for Multimodal Aspect-Based Sentiment Analysis
by: Liu, Hao, et al.
Published: (2025)
by: Liu, Hao, et al.
Published: (2025)
T2VParser: Adaptive Decomposition Tokens for Partial Alignment in Text to Video Retrieval
by: Li, Yili, et al.
Published: (2025)
by: Li, Yili, et al.
Published: (2025)
Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
by: Waheed, Abdul, et al.
Published: (2025)
by: Waheed, Abdul, et al.
Published: (2025)
Talking-to-Build: How LLM-Assisted Interface Shapes Player Performance and Experience in Minecraft
by: Sun, Xin, et al.
Published: (2025)
by: Sun, Xin, et al.
Published: (2025)
CPSL: Representing Volumetric Video via Content-Promoted Scene Layers
by: Hu, Kaiyuan, et al.
Published: (2025)
by: Hu, Kaiyuan, et al.
Published: (2025)
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
by: Mitra, Chancharik, et al.
Published: (2024)
by: Mitra, Chancharik, et al.
Published: (2024)
TimeNeRF: Building Generalizable Neural Radiance Fields across Time from Few-Shot Input Views
by: Hung, Hsiang-Hui, et al.
Published: (2025)
by: Hung, Hsiang-Hui, et al.
Published: (2025)
Shorter Is Different: Characterizing the Dynamics of Short-Form Video Platforms
by: Chen, Zhilong, et al.
Published: (2024)
by: Chen, Zhilong, et al.
Published: (2024)
MIND Your Reasoning: A Meta-Cognitive Intuitive-Reflective Network for Dual-Reasoning in Multimodal Stance Detection
by: Wang, Bingbing, et al.
Published: (2025)
by: Wang, Bingbing, et al.
Published: (2025)
MMAPS: End-to-End Multi-Grained Multi-Modal Attribute-Aware Product Summarization
by: Chen, Tao, et al.
Published: (2023)
by: Chen, Tao, et al.
Published: (2023)
Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization
by: Fu, Yuhan, et al.
Published: (2024)
by: Fu, Yuhan, et al.
Published: (2024)
A Survey of Information Disorder on Video-Sharing Platforms
by: Li, Meiyu, et al.
Published: (2025)
by: Li, Meiyu, et al.
Published: (2025)
LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
MMPKUBase: A Comprehensive and High-quality Chinese Multi-modal Knowledge Graph
by: Yi, Xuan, et al.
Published: (2024)
by: Yi, Xuan, et al.
Published: (2024)
Remember Past, Anticipate Future: Learning Continual Multimodal Misinformation Detectors
by: Wang, Bing, et al.
Published: (2025)
by: Wang, Bing, et al.
Published: (2025)
SEA: Low-Resource Safety Alignment for Multimodal Large Language Models via Synthetic Embeddings
by: Lu, Weikai, et al.
Published: (2025)
by: Lu, Weikai, et al.
Published: (2025)
SACRED: A Faithful Annotated Multimedia Multimodal Multilingual Dataset for Classifying Connectedness Types in Online Spirituality
by: Guan, Qinghao, et al.
Published: (2026)
by: Guan, Qinghao, et al.
Published: (2026)
Retrieval-Augmented Multimodal Model for Fake News Detection
by: Li, Yiheng, et al.
Published: (2026)
by: Li, Yiheng, et al.
Published: (2026)
MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
by: Zhang, Yuang, et al.
Published: (2024)
by: Zhang, Yuang, et al.
Published: (2024)
P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark
by: Sun, Tao, et al.
Published: (2025)
by: Sun, Tao, et al.
Published: (2025)
Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding
by: Li, Zhu, et al.
Published: (2025)
by: Li, Zhu, et al.
Published: (2025)
Collaborative Evolution: Multi-Round Learning Between Large and Small Language Models for Emergent Fake News Detection
by: Zhou, Ziyi, et al.
Published: (2025)
by: Zhou, Ziyi, et al.
Published: (2025)
AIM: Let Any Multi-modal Large Language Models Embrace Efficient In-Context Learning
by: Gao, Jun, et al.
Published: (2024)
by: Gao, Jun, et al.
Published: (2024)
Image Conductor: Precision Control for Interactive Video Synthesis
by: Li, Yaowei, et al.
Published: (2024)
by: Li, Yaowei, et al.
Published: (2024)
Applying LLM-Powered Virtual Humans to Child Interviews in Child-Centered Design
by: Li, Linshi, et al.
Published: (2025)
by: Li, Linshi, et al.
Published: (2025)
Mutual Information-based Representations Disentanglement for Unaligned Multimodal Language Sequences
by: Qian, Fan, et al.
Published: (2024)
by: Qian, Fan, et al.
Published: (2024)
Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors
by: Li, Jia, et al.
Published: (2025)
by: Li, Jia, et al.
Published: (2025)
Enhancing Multimodal Entity and Relation Extraction with Variational Information Bottleneck
by: Cui, Shiyao, et al.
Published: (2023)
by: Cui, Shiyao, et al.
Published: (2023)
Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment
by: Li, Jia, et al.
Published: (2025)
by: Li, Jia, et al.
Published: (2025)
Hierarchical Aligned Multimodal Learning for NER on Tweet Posts
by: Liu, Peipei, et al.
Published: (2023)
by: Liu, Peipei, et al.
Published: (2023)
Magnifier Prompt: Tackling Multimodal Hallucination via Extremely Simple Instructions
by: Fu, Yuhan, et al.
Published: (2024)
by: Fu, Yuhan, et al.
Published: (2024)
Towards Event Extraction from Speech with Contextual Clues
by: Kang, Jingqi, et al.
Published: (2024)
by: Kang, Jingqi, et al.
Published: (2024)
Explainable Multimodal Emotion Recognition
by: Lian, Zheng, et al.
Published: (2023)
by: Lian, Zheng, et al.
Published: (2023)
MotionPro: A Precise Motion Controller for Image-to-Video Generation
by: Zhang, Zhongwei, et al.
Published: (2025)
by: Zhang, Zhongwei, et al.
Published: (2025)
SpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description
by: Jin, Zeyu, et al.
Published: (2024)
by: Jin, Zeyu, et al.
Published: (2024)
Similar Items
-
Towards Understanding Camera Motions in Any Video
by: Lin, Zhiqiu, et al.
Published: (2025) -
Language Models as Black-Box Optimizers for Vision-Language Models
by: Liu, Shihong, et al.
Published: (2023) -
Evaluating Text-to-Visual Generation with Image-to-Text Generation
by: Lin, Zhiqiu, et al.
Published: (2024) -
GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation
by: Li, Baiqi, et al.
Published: (2024) -
ResearchPulse: Building Method-Experiment Chains through Multi-Document Scientific Inference
by: Chen, Qi, et al.
Published: (2025)