VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kulkarni, Yogesh, Fazli, Pooyan |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
VideoSAVi: Self-Aligned Video Language Models without Human Supervision
par: Kulkarni, Yogesh, et autres
Publié: (2024)
par: Kulkarni, Yogesh, et autres
Publié: (2024)
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
par: Kulkarni, Yogesh, et autres
Publié: (2025)
par: Kulkarni, Yogesh, et autres
Publié: (2025)
AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
par: Kulkarni, Yogesh, et autres
Publié: (2025)
par: Kulkarni, Yogesh, et autres
Publié: (2025)
ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
par: Li, Chaoyu, et autres
Publié: (2025)
par: Li, Chaoyu, et autres
Publié: (2025)
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
par: Li, Chaoyu, et autres
Publié: (2024)
par: Li, Chaoyu, et autres
Publié: (2024)
VideoA11y: Method and Dataset for Accessible Video Description
par: Li, Chaoyu, et autres
Publié: (2025)
par: Li, Chaoyu, et autres
Publié: (2025)
ChartQA-X: Generating Explanations for Visual Chart Reasoning
par: Hegde, Shamanthak, et autres
Publié: (2025)
par: Hegde, Shamanthak, et autres
Publié: (2025)
Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
par: Kumar, Yogesh
Publié: (2025)
par: Kumar, Yogesh
Publié: (2025)
VideoLLM Benchmarks and Evaluation: A Survey
par: Kumar, Yogesh
Publié: (2025)
par: Kumar, Yogesh
Publié: (2025)
FrameOracle: Learning What to See and How Much to See in Videos
par: Li, Chaoyu, et autres
Publié: (2025)
par: Li, Chaoyu, et autres
Publié: (2025)
CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation
par: Li, Chaoyu, et autres
Publié: (2026)
par: Li, Chaoyu, et autres
Publié: (2026)
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs
par: Wang, Xiaodong, et autres
Publié: (2025)
par: Wang, Xiaodong, et autres
Publié: (2025)
Diffusion-APO: Trajectory-Aware Direct Preference Alignment for Video Diffusion Transformers
par: Zhu, Jingyuan, et autres
Publié: (2026)
par: Zhu, Jingyuan, et autres
Publié: (2026)
VideoDPO: Omni-Preference Alignment for Video Diffusion Generation
par: Liu, Runtao, et autres
Publié: (2024)
par: Liu, Runtao, et autres
Publié: (2024)
Learning to Localize Actions in Instructional Videos with LLM-Based Multi-Pathway Text-Video Alignment
par: Chen, Yuxiao, et autres
Publié: (2024)
par: Chen, Yuxiao, et autres
Publié: (2024)
PASTA: Towards Flexible and Efficient HDR Imaging Via Progressively Aggregated Spatio-Temporal Alignment
par: Liu, Xiaoning, et autres
Publié: (2024)
par: Liu, Xiaoning, et autres
Publié: (2024)
VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?
par: Tang, Yolo Y., et autres
Publié: (2024)
par: Tang, Yolo Y., et autres
Publié: (2024)
McSc: Motion-Corrective Preference Alignment for Video Generation with Self-Critic Hierarchical Reasoning
par: Yang, Qiushi, et autres
Publié: (2025)
par: Yang, Qiushi, et autres
Publié: (2025)
Aligning Moments in Time using Video Queries
par: Kumar, Yogesh, et autres
Publié: (2025)
par: Kumar, Yogesh, et autres
Publié: (2025)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
par: Ji, Yatai, et autres
Publié: (2024)
par: Ji, Yatai, et autres
Publié: (2024)
Stable Mean Teacher for Semi-supervised Video Action Detection
par: Kumar, Akash, et autres
Publié: (2024)
par: Kumar, Akash, et autres
Publié: (2024)
HuViDPO:Enhancing Video Generation through Direct Preference Optimization for Human-Centric Alignment
par: Jiang, Lifan, et autres
Publié: (2025)
par: Jiang, Lifan, et autres
Publié: (2025)
Can Text-to-Video Generation help Video-Language Alignment?
par: Zanella, Luca, et autres
Publié: (2025)
par: Zanella, Luca, et autres
Publié: (2025)
StreamReady: Learning What to Answer and When in Long Streaming Videos
par: Azad, Shehreen, et autres
Publié: (2026)
par: Azad, Shehreen, et autres
Publié: (2026)
MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation
par: Tong, Haibo, et autres
Publié: (2025)
par: Tong, Haibo, et autres
Publié: (2025)
From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models
par: Cai, Zefan, et autres
Publié: (2025)
par: Cai, Zefan, et autres
Publié: (2025)
Reg-DPO: SFT-Regularized Direct Preference Optimization with GT-Pair for Improving Video Generation
par: Du, Jie, et autres
Publié: (2025)
par: Du, Jie, et autres
Publié: (2025)
HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
par: Azad, Shehreen, et autres
Publié: (2025)
par: Azad, Shehreen, et autres
Publié: (2025)
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
par: Kumar, Akash, et autres
Publié: (2025)
par: Kumar, Akash, et autres
Publié: (2025)
ViLA: Efficient Video-Language Alignment for Video Question Answering
par: Wang, Xijun, et autres
Publié: (2023)
par: Wang, Xijun, et autres
Publié: (2023)
Foundation Models for Video Understanding: A Survey
par: Madan, Neelu, et autres
Publié: (2024)
par: Madan, Neelu, et autres
Publié: (2024)
Alignment-free Raw Video Demoireing
par: Xu, Shuning, et autres
Publié: (2024)
par: Xu, Shuning, et autres
Publié: (2024)
OnlineVPO: Align Video Diffusion Model with Online Video-Centric Preference Optimization
par: Zhang, Jiacheng, et autres
Publié: (2024)
par: Zhang, Jiacheng, et autres
Publié: (2024)
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
par: Kumar, Akash, et autres
Publié: (2025)
par: Kumar, Akash, et autres
Publié: (2025)
VideoLLM-online: Online Video Large Language Model for Streaming Video
par: Chen, Joya, et autres
Publié: (2024)
par: Chen, Joya, et autres
Publié: (2024)
VideoLifter: Lifting Videos to 3D with Fast Hierarchical Stereo Alignment
par: Cong, Wenyan, et autres
Publié: (2025)
par: Cong, Wenyan, et autres
Publié: (2025)
Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models
par: Yi, Jinhui, et autres
Publié: (2024)
par: Yi, Jinhui, et autres
Publié: (2024)
Temporal Object-Aware Vision Transformer for Few-Shot Video Object Detection
par: Kumar, Yogesh, et autres
Publié: (2025)
par: Kumar, Yogesh, et autres
Publié: (2025)
ParallelVLM: Lossless Video-LLM Acceleration with Visual Alignment Aware Parallel Speculative Decoding
par: Kong, Quan, et autres
Publié: (2026)
par: Kong, Quan, et autres
Publié: (2026)
Discriminator-Free Direct Preference Optimization for Video Diffusion
par: Cheng, Haoran, et autres
Publié: (2025)
par: Cheng, Haoran, et autres
Publié: (2025)
Documents similaires
-
VideoSAVi: Self-Aligned Video Language Models without Human Supervision
par: Kulkarni, Yogesh, et autres
Publié: (2024) -
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
par: Kulkarni, Yogesh, et autres
Publié: (2025) -
AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
par: Kulkarni, Yogesh, et autres
Publié: (2025) -
ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
par: Li, Chaoyu, et autres
Publié: (2025) -
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
par: Li, Chaoyu, et autres
Publié: (2024)