Saved in:
| Main Authors: | Kulkarni, Yogesh, Fazli, Pooyan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2504.14096 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoSAVi: Self-Aligned Video Language Models without Human Supervision
by: Kulkarni, Yogesh, et al.
Published: (2024)
by: Kulkarni, Yogesh, et al.
Published: (2024)
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
by: Kulkarni, Yogesh, et al.
Published: (2025)
by: Kulkarni, Yogesh, et al.
Published: (2025)
AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
by: Kulkarni, Yogesh, et al.
Published: (2025)
by: Kulkarni, Yogesh, et al.
Published: (2025)
ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
by: Li, Chaoyu, et al.
Published: (2025)
by: Li, Chaoyu, et al.
Published: (2025)
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
by: Li, Chaoyu, et al.
Published: (2024)
by: Li, Chaoyu, et al.
Published: (2024)
VideoA11y: Method and Dataset for Accessible Video Description
by: Li, Chaoyu, et al.
Published: (2025)
by: Li, Chaoyu, et al.
Published: (2025)
ChartQA-X: Generating Explanations for Visual Chart Reasoning
by: Hegde, Shamanthak, et al.
Published: (2025)
by: Hegde, Shamanthak, et al.
Published: (2025)
FrameOracle: Learning What to See and How Much to See in Videos
by: Li, Chaoyu, et al.
Published: (2025)
by: Li, Chaoyu, et al.
Published: (2025)
Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
by: Kumar, Yogesh
Published: (2025)
by: Kumar, Yogesh
Published: (2025)
VideoLLM Benchmarks and Evaluation: A Survey
by: Kumar, Yogesh
Published: (2025)
by: Kumar, Yogesh
Published: (2025)
CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation
by: Li, Chaoyu, et al.
Published: (2026)
by: Li, Chaoyu, et al.
Published: (2026)
VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?
by: Tang, Yolo Y., et al.
Published: (2024)
by: Tang, Yolo Y., et al.
Published: (2024)
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs
by: Wang, Xiaodong, et al.
Published: (2025)
by: Wang, Xiaodong, et al.
Published: (2025)
VideoDPO: Omni-Preference Alignment for Video Diffusion Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
PASTA: Towards Flexible and Efficient HDR Imaging Via Progressively Aggregated Spatio-Temporal Alignment
by: Liu, Xiaoning, et al.
Published: (2024)
by: Liu, Xiaoning, et al.
Published: (2024)
Diffusion-APO: Trajectory-Aware Direct Preference Alignment for Video Diffusion Transformers
by: Zhu, Jingyuan, et al.
Published: (2026)
by: Zhu, Jingyuan, et al.
Published: (2026)
Learning to Localize Actions in Instructional Videos with LLM-Based Multi-Pathway Text-Video Alignment
by: Chen, Yuxiao, et al.
Published: (2024)
by: Chen, Yuxiao, et al.
Published: (2024)
Aligning Moments in Time using Video Queries
by: Kumar, Yogesh, et al.
Published: (2025)
by: Kumar, Yogesh, et al.
Published: (2025)
Stable Mean Teacher for Semi-supervised Video Action Detection
by: Kumar, Akash, et al.
Published: (2024)
by: Kumar, Akash, et al.
Published: (2024)
McSc: Motion-Corrective Preference Alignment for Video Generation with Self-Critic Hierarchical Reasoning
by: Yang, Qiushi, et al.
Published: (2025)
by: Yang, Qiushi, et al.
Published: (2025)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
by: Ji, Yatai, et al.
Published: (2024)
by: Ji, Yatai, et al.
Published: (2024)
StreamReady: Learning What to Answer and When in Long Streaming Videos
by: Azad, Shehreen, et al.
Published: (2026)
by: Azad, Shehreen, et al.
Published: (2026)
Reg-DPO: SFT-Regularized Direct Preference Optimization with GT-Pair for Improving Video Generation
by: Du, Jie, et al.
Published: (2025)
by: Du, Jie, et al.
Published: (2025)
HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
by: Azad, Shehreen, et al.
Published: (2025)
by: Azad, Shehreen, et al.
Published: (2025)
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
by: Kumar, Akash, et al.
Published: (2025)
by: Kumar, Akash, et al.
Published: (2025)
HuViDPO:Enhancing Video Generation through Direct Preference Optimization for Human-Centric Alignment
by: Jiang, Lifan, et al.
Published: (2025)
by: Jiang, Lifan, et al.
Published: (2025)
From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models
by: Cai, Zefan, et al.
Published: (2025)
by: Cai, Zefan, et al.
Published: (2025)
OSCaR: Object State Captioning and State Change Representation
by: Nguyen, Nguyen, et al.
Published: (2024)
by: Nguyen, Nguyen, et al.
Published: (2024)
Foundation Models for Video Understanding: A Survey
by: Madan, Neelu, et al.
Published: (2024)
by: Madan, Neelu, et al.
Published: (2024)
Can Text-to-Video Generation help Video-Language Alignment?
by: Zanella, Luca, et al.
Published: (2025)
by: Zanella, Luca, et al.
Published: (2025)
Temporal Object-Aware Vision Transformer for Few-Shot Video Object Detection
by: Kumar, Yogesh, et al.
Published: (2025)
by: Kumar, Yogesh, et al.
Published: (2025)
EZ-CLIP: Efficient Zeroshot Video Action Recognition
by: Ahmad, Shahzad, et al.
Published: (2023)
by: Ahmad, Shahzad, et al.
Published: (2023)
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
by: Kumar, Akash, et al.
Published: (2025)
by: Kumar, Akash, et al.
Published: (2025)
MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation
by: Tong, Haibo, et al.
Published: (2025)
by: Tong, Haibo, et al.
Published: (2025)
ViLA: Efficient Video-Language Alignment for Video Question Answering
by: Wang, Xijun, et al.
Published: (2023)
by: Wang, Xijun, et al.
Published: (2023)
Semi-supervised Active Learning for Video Action Detection
by: Singh, Ayush, et al.
Published: (2023)
by: Singh, Ayush, et al.
Published: (2023)
Alignment-free Raw Video Demoireing
by: Xu, Shuning, et al.
Published: (2024)
by: Xu, Shuning, et al.
Published: (2024)
OnlineVPO: Align Video Diffusion Model with Online Video-Centric Preference Optimization
by: Zhang, Jiacheng, et al.
Published: (2024)
by: Zhang, Jiacheng, et al.
Published: (2024)
VideoLifter: Lifting Videos to 3D with Fast Hierarchical Stereo Alignment
by: Cong, Wenyan, et al.
Published: (2025)
by: Cong, Wenyan, et al.
Published: (2025)
Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models
by: Yi, Jinhui, et al.
Published: (2024)
by: Yi, Jinhui, et al.
Published: (2024)
Similar Items
-
VideoSAVi: Self-Aligned Video Language Models without Human Supervision
by: Kulkarni, Yogesh, et al.
Published: (2024) -
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
by: Kulkarni, Yogesh, et al.
Published: (2025) -
AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
by: Kulkarni, Yogesh, et al.
Published: (2025) -
ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
by: Li, Chaoyu, et al.
Published: (2025) -
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
by: Li, Chaoyu, et al.
Published: (2024)