GT-SVJ: Generative-Transformer-Based Self-Supervised Video Judge For Efficient Video Reward Modeling
Fuente:
arXiv
Guardado en:
| Autores principales: | Shekhar, Shivanshu, Bhattacharya, Uttaran, Addanki, Raghavendra, Tanjim, Mehrab, Sarkhel, Somdeb, Zhang, Tong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SKALD: Learning-Based Shot Assembly for Coherent Multi-Shot Video Creation
por: Lu, Chen Yi, et al.
Publicado: (2025)
por: Lu, Chen Yi, et al.
Publicado: (2025)
X-Reflect: Cross-Reflection Prompting for Multimodal Recommendation
por: Lyu, Hanjia, et al.
Publicado: (2024)
por: Lyu, Hanjia, et al.
Publicado: (2024)
SEE-DPO: Self Entropy Enhanced Direct Preference Optimization
por: Shekhar, Shivanshu, et al.
Publicado: (2024)
por: Shekhar, Shivanshu, et al.
Publicado: (2024)
On Explaining Visual Captioning with Hybrid Markov Logic Networks
por: Shah, Monika, et al.
Publicado: (2025)
por: Shah, Monika, et al.
Publicado: (2025)
Disentangling Fine-Tuning from Pre-Training in Visual Captioning with Hybrid Markov Logic
por: Shah, Monika, et al.
Publicado: (2025)
por: Shah, Monika, et al.
Publicado: (2025)
Rethinking Failure Attribution in Multi-Agent Systems: A Multi-Perspective Benchmark and Evaluation
por: In, Yeonjun, et al.
Publicado: (2026)
por: In, Yeonjun, et al.
Publicado: (2026)
Analyzing the Sensitivity of Vision Language Models in Visual Question Answering
por: Shah, Monika, et al.
Publicado: (2025)
por: Shah, Monika, et al.
Publicado: (2025)
Efficient and Robust Registration on the 3D Special Euclidean Group
por: Bhattacharya, Uttaran, et al.
Publicado: (2019)
por: Bhattacharya, Uttaran, et al.
Publicado: (2019)
HighlightMe: Detecting Highlights from Human-Centric Videos
por: Bhattacharya, Uttaran, et al.
Publicado: (2021)
por: Bhattacharya, Uttaran, et al.
Publicado: (2021)
Show Me What I Like: Detecting User-Specific Video Highlights Using Content-Based Multi-Head Attention
por: Bhattacharya, Uttaran, et al.
Publicado: (2022)
por: Bhattacharya, Uttaran, et al.
Publicado: (2022)
VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding
por: Waheed, Abdul, et al.
Publicado: (2025)
por: Waheed, Abdul, et al.
Publicado: (2025)
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
por: Slyman, Eric, et al.
Publicado: (2025)
por: Slyman, Eric, et al.
Publicado: (2025)
Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation
por: Lu, Yunhong, et al.
Publicado: (2025)
por: Lu, Yunhong, et al.
Publicado: (2025)
Speech2UnifiedExpressions: Synchronous Synthesis of Co-Speech Affective Face and Body Expressions from Affordable Inputs
por: Bhattacharya, Uttaran, et al.
Publicado: (2024)
por: Bhattacharya, Uttaran, et al.
Publicado: (2024)
VideoSSR: Video Self-Supervised Reinforcement Learning
por: He, Zefeng, et al.
Publicado: (2025)
por: He, Zefeng, et al.
Publicado: (2025)
MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation
por: Tong, Haibo, et al.
Publicado: (2025)
por: Tong, Haibo, et al.
Publicado: (2025)
Video Generation Models Are Good Latent Reward Models
por: Mi, Xiaoyue, et al.
Publicado: (2025)
por: Mi, Xiaoyue, et al.
Publicado: (2025)
VideoSAVi: Self-Aligned Video Language Models without Human Supervision
por: Kulkarni, Yogesh, et al.
Publicado: (2024)
por: Kulkarni, Yogesh, et al.
Publicado: (2024)
Geo-Align: Video Generation Alignment via Metric Geometry Reward
por: Li, Zizun, et al.
Publicado: (2026)
por: Li, Zizun, et al.
Publicado: (2026)
InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
por: Wang, Chenting, et al.
Publicado: (2025)
por: Wang, Chenting, et al.
Publicado: (2025)
VideoClusterNet: Self-Supervised and Adaptive Face Clustering For Videos
por: Walawalkar, Devesh, et al.
Publicado: (2024)
por: Walawalkar, Devesh, et al.
Publicado: (2024)
SelfHVD: Self-Supervised Handheld Video Deblurring
por: Xu, Honglei, et al.
Publicado: (2025)
por: Xu, Honglei, et al.
Publicado: (2025)
How Effective are Self-Supervised Models for Contact Identification in Videos
por: Gunawardhana, Malitha, et al.
Publicado: (2024)
por: Gunawardhana, Malitha, et al.
Publicado: (2024)
WorldModelBench: Judging Video Generation Models As World Models
por: Li, Dacheng, et al.
Publicado: (2025)
por: Li, Dacheng, et al.
Publicado: (2025)
Self-Supervised Video Desmoking for Laparoscopic Surgery
por: Wu, Renlong, et al.
Publicado: (2024)
por: Wu, Renlong, et al.
Publicado: (2024)
Self-Supervised Animal Identification for Long Videos
por: Fang, Xuyang, et al.
Publicado: (2026)
por: Fang, Xuyang, et al.
Publicado: (2026)
VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding
por: Zhang, Zhihong, et al.
Publicado: (2025)
por: Zhang, Zhihong, et al.
Publicado: (2025)
Efficient Self-Supervised Video Hashing with Selective State Spaces
por: Wang, Jinpeng, et al.
Publicado: (2024)
por: Wang, Jinpeng, et al.
Publicado: (2024)
Reward-Forcing: Autoregressive Video Generation with Reward Feedback
por: Zhang, Jingran, et al.
Publicado: (2026)
por: Zhang, Jingran, et al.
Publicado: (2026)
SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
por: Chen, Junsong, et al.
Publicado: (2025)
por: Chen, Junsong, et al.
Publicado: (2025)
Video-Based Reward Modeling for Computer-Use Agents
por: Song, Linxin, et al.
Publicado: (2026)
por: Song, Linxin, et al.
Publicado: (2026)
Video Consistency Distance: Enhancing Temporal Consistency for Image-to-Video Generation via Reward-Based Fine-Tuning
por: Aoshima, Takehiro, et al.
Publicado: (2025)
por: Aoshima, Takehiro, et al.
Publicado: (2025)
A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations
por: Deshpande, Vijeta, et al.
Publicado: (2025)
por: Deshpande, Vijeta, et al.
Publicado: (2025)
Advancing Video Self-Supervised Learning via Image Foundation Models
por: Wu, Jingwei, et al.
Publicado: (2025)
por: Wu, Jingwei, et al.
Publicado: (2025)
Reshoot-Anything: A Self-Supervised Model for In-the-Wild Video Reshooting
por: Paliwal, Avinash, et al.
Publicado: (2026)
por: Paliwal, Avinash, et al.
Publicado: (2026)
SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
por: Lian, Jiesong, et al.
Publicado: (2025)
por: Lian, Jiesong, et al.
Publicado: (2025)
Sharpening Lightweight Models for Generalized Polyp Segmentation: A Boundary Guided Distillation from Foundation Models
por: Agnihotri, Shivanshu, et al.
Publicado: (2026)
por: Agnihotri, Shivanshu, et al.
Publicado: (2026)
Depth-Wise Representation Development Under Blockwise Self-Supervised Learning for Video Vision Transformers
por: Römer, Jonas, et al.
Publicado: (2026)
por: Römer, Jonas, et al.
Publicado: (2026)
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders
por: Ahamed, Shihab Aaqil, et al.
Publicado: (2025)
por: Ahamed, Shihab Aaqil, et al.
Publicado: (2025)
Joint Self-Supervised Video Alignment and Action Segmentation
por: Ali, Ali Shah, et al.
Publicado: (2025)
por: Ali, Ali Shah, et al.
Publicado: (2025)
Ejemplares similares
-
SKALD: Learning-Based Shot Assembly for Coherent Multi-Shot Video Creation
por: Lu, Chen Yi, et al.
Publicado: (2025) -
X-Reflect: Cross-Reflection Prompting for Multimodal Recommendation
por: Lyu, Hanjia, et al.
Publicado: (2024) -
SEE-DPO: Self Entropy Enhanced Direct Preference Optimization
por: Shekhar, Shivanshu, et al.
Publicado: (2024) -
On Explaining Visual Captioning with Hybrid Markov Logic Networks
por: Shah, Monika, et al.
Publicado: (2025) -
Disentangling Fine-Tuning from Pre-Training in Visual Captioning with Hybrid Markov Logic
por: Shah, Monika, et al.
Publicado: (2025)