VideoSAVi: Self-Aligned Video Language Models without Human Supervision
Fuente:
arXiv
Saved in:
| Main Authors: | Kulkarni, Yogesh, Fazli, Pooyan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment
by: Kulkarni, Yogesh, et al.
Published: (2025)
by: Kulkarni, Yogesh, et al.
Published: (2025)
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
by: Kulkarni, Yogesh, et al.
Published: (2025)
by: Kulkarni, Yogesh, et al.
Published: (2025)
AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
by: Kulkarni, Yogesh, et al.
Published: (2025)
by: Kulkarni, Yogesh, et al.
Published: (2025)
ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
by: Li, Chaoyu, et al.
Published: (2025)
by: Li, Chaoyu, et al.
Published: (2025)
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
by: Li, Chaoyu, et al.
Published: (2024)
by: Li, Chaoyu, et al.
Published: (2024)
VideoA11y: Method and Dataset for Accessible Video Description
by: Li, Chaoyu, et al.
Published: (2025)
by: Li, Chaoyu, et al.
Published: (2025)
ChartQA-X: Generating Explanations for Visual Chart Reasoning
by: Hegde, Shamanthak, et al.
Published: (2025)
by: Hegde, Shamanthak, et al.
Published: (2025)
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
by: Kumar, Akash, et al.
Published: (2025)
by: Kumar, Akash, et al.
Published: (2025)
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
by: Kumar, Akash, et al.
Published: (2025)
by: Kumar, Akash, et al.
Published: (2025)
Aligning Moments in Time using Video Queries
by: Kumar, Yogesh, et al.
Published: (2025)
by: Kumar, Yogesh, et al.
Published: (2025)
SAViL-Det: Semantic-Aware Vision-Language Model for Multi-Script Text Detection
by: Zighem, Mohammed-En-Nadhir, et al.
Published: (2025)
by: Zighem, Mohammed-En-Nadhir, et al.
Published: (2025)
InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
by: Wang, Chenting, et al.
Published: (2025)
by: Wang, Chenting, et al.
Published: (2025)
FrameOracle: Learning What to See and How Much to See in Videos
by: Li, Chaoyu, et al.
Published: (2025)
by: Li, Chaoyu, et al.
Published: (2025)
Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
by: Kumar, Yogesh
Published: (2025)
by: Kumar, Yogesh
Published: (2025)
An Empirical Study of Accuracy-Robustness Tradeoff and Training Efficiency in Self-Supervised Learning
by: Ghofrani, Fatemeh, et al.
Published: (2025)
by: Ghofrani, Fatemeh, et al.
Published: (2025)
VideoSSR: Video Self-Supervised Reinforcement Learning
by: He, Zefeng, et al.
Published: (2025)
by: He, Zefeng, et al.
Published: (2025)
CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation
by: Li, Chaoyu, et al.
Published: (2026)
by: Li, Chaoyu, et al.
Published: (2026)
Video Depth without Video Models
by: Ke, Bingxin, et al.
Published: (2024)
by: Ke, Bingxin, et al.
Published: (2024)
Aligning Anime Video Generation with Human Feedback
by: Zhu, Bingwen, et al.
Published: (2025)
by: Zhu, Bingwen, et al.
Published: (2025)
Video-Bench: Human-Aligned Video Generation Benchmark
by: Han, Hui, et al.
Published: (2025)
by: Han, Hui, et al.
Published: (2025)
Human-Aligned Generative Perception: Bridging Psychophysics and Generative Models
by: Titikhsha, Antara, et al.
Published: (2025)
by: Titikhsha, Antara, et al.
Published: (2025)
Behavioral Geometric Supervision Aligns Video Foundation Models with Human Social Perception
by: Garcia, Kathy, et al.
Published: (2025)
by: Garcia, Kathy, et al.
Published: (2025)
Language-Guided Self-Supervised Video Summarization Using Text Semantic Matching Considering the Diversity of the Video
by: Sugihara, Tomoya, et al.
Published: (2024)
by: Sugihara, Tomoya, et al.
Published: (2024)
VideoClusterNet: Self-Supervised and Adaptive Face Clustering For Videos
by: Walawalkar, Devesh, et al.
Published: (2024)
by: Walawalkar, Devesh, et al.
Published: (2024)
Aligning Effective Tokens with Video Anomaly in Large Language Models
by: Chen, Yingxian, et al.
Published: (2025)
by: Chen, Yingxian, et al.
Published: (2025)
SelfHVD: Self-Supervised Handheld Video Deblurring
by: Xu, Honglei, et al.
Published: (2025)
by: Xu, Honglei, et al.
Published: (2025)
How Effective are Self-Supervised Models for Contact Identification in Videos
by: Gunawardhana, Malitha, et al.
Published: (2024)
by: Gunawardhana, Malitha, et al.
Published: (2024)
Self-Supervised Video Desmoking for Laparoscopic Surgery
by: Wu, Renlong, et al.
Published: (2024)
by: Wu, Renlong, et al.
Published: (2024)
Self-Supervised Animal Identification for Long Videos
by: Fang, Xuyang, et al.
Published: (2026)
by: Fang, Xuyang, et al.
Published: (2026)
Finding Optimal Video Moment without Training: Gaussian Boundary Optimization for Weakly Supervised Video Grounding
by: Kim, Sunoh, et al.
Published: (2026)
by: Kim, Sunoh, et al.
Published: (2026)
GT-SVJ: Generative-Transformer-Based Self-Supervised Video Judge For Efficient Video Reward Modeling
by: Shekhar, Shivanshu, et al.
Published: (2026)
by: Shekhar, Shivanshu, et al.
Published: (2026)
Advancing Video Self-Supervised Learning via Image Foundation Models
by: Wu, Jingwei, et al.
Published: (2025)
by: Wu, Jingwei, et al.
Published: (2025)
Reshoot-Anything: A Self-Supervised Model for In-the-Wild Video Reshooting
by: Paliwal, Avinash, et al.
Published: (2026)
by: Paliwal, Avinash, et al.
Published: (2026)
SPKLIP: Aligning Spike Video Streams with Natural Language
by: Gao, Yongchang, et al.
Published: (2025)
by: Gao, Yongchang, et al.
Published: (2025)
VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale
by: Kulkarni, Parth Parag, et al.
Published: (2026)
by: Kulkarni, Parth Parag, et al.
Published: (2026)
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders
by: Ahamed, Shihab Aaqil, et al.
Published: (2025)
by: Ahamed, Shihab Aaqil, et al.
Published: (2025)
Joint Self-Supervised Video Alignment and Action Segmentation
by: Ali, Ali Shah, et al.
Published: (2025)
by: Ali, Ali Shah, et al.
Published: (2025)
ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO
by: Ahn, Daechul, et al.
Published: (2024)
by: Ahn, Daechul, et al.
Published: (2024)
VideoLLM Benchmarks and Evaluation: A Survey
by: Kumar, Yogesh
Published: (2025)
by: Kumar, Yogesh
Published: (2025)
FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning
by: Hu, Shiyu, et al.
Published: (2024)
by: Hu, Shiyu, et al.
Published: (2024)
Similar Items
-
VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment
by: Kulkarni, Yogesh, et al.
Published: (2025) -
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
by: Kulkarni, Yogesh, et al.
Published: (2025) -
AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
by: Kulkarni, Yogesh, et al.
Published: (2025) -
ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
by: Li, Chaoyu, et al.
Published: (2025) -
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
by: Li, Chaoyu, et al.
Published: (2024)