VCEval: Rethinking What is a Good Educational Video and How to Automatically Evaluate It
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Xiaoxuan, Gu, Zhouhong, Jiang, Sihang, Li, Zhixu, Feng, Hongwei, Xiao, Yanghua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating Semantic Variation in Text-to-Image Synthesis: A Causal Perspective
von: Zhu, Xiangru, et al.
Veröffentlicht: (2024)
von: Zhu, Xiangru, et al.
Veröffentlicht: (2024)
MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models
von: Guo, Zile, et al.
Veröffentlicht: (2026)
von: Guo, Zile, et al.
Veröffentlicht: (2026)
Rethinking Video with a Universal Event-Based Representation
von: Freeman, Andrew
Veröffentlicht: (2024)
von: Freeman, Andrew
Veröffentlicht: (2024)
Multimodal Fake News Video Explanation: Dataset, Analysis and Evaluation
von: Chen, Lizhi, et al.
Veröffentlicht: (2025)
von: Chen, Lizhi, et al.
Veröffentlicht: (2025)
PolySmart @ TRECVid 2024 Medical Video Question Answering
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
von: Cai, Lingling, et al.
Veröffentlicht: (2024)
von: Cai, Lingling, et al.
Veröffentlicht: (2024)
AutoMV: An Automatic Multi-Agent System for Music Video Generation
von: Tang, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Tang, Xiaoxuan, et al.
Veröffentlicht: (2025)
Look, Compare and Draw: Differential Query Transformer for Automatic Oil Painting
von: Liu, Lingyu, et al.
Veröffentlicht: (2026)
von: Liu, Lingyu, et al.
Veröffentlicht: (2026)
ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
Hierarchical Action Recognition: A Contrastive Video-Language Approach with Hierarchical Interactions
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
PolySmart @ TRECVid 2024 Video Captioning (VTT)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation
von: Shi, Haoyuan, et al.
Veröffentlicht: (2026)
von: Shi, Haoyuan, et al.
Veröffentlicht: (2026)
Adaptive 3D Gaussian Splatting Video Streaming
von: Gong, Han, et al.
Veröffentlicht: (2025)
von: Gong, Han, et al.
Veröffentlicht: (2025)
A Human-Annotated Video Dataset for Training and Evaluation of 360-Degree Video Summarization Methods
von: Kontostathis, Ioannis, et al.
Veröffentlicht: (2024)
von: Kontostathis, Ioannis, et al.
Veröffentlicht: (2024)
Looking Backward: Streaming Video-to-Video Translation with Feature Banks
von: Liang, Feng, et al.
Veröffentlicht: (2024)
von: Liang, Feng, et al.
Veröffentlicht: (2024)
VC-Bench: Pioneering the Video Connecting Benchmark with a Dataset and Evaluation Metrics
von: Yin, Zhiyu, et al.
Veröffentlicht: (2026)
von: Yin, Zhiyu, et al.
Veröffentlicht: (2026)
Bridging the Pose-Semantic Gap: A Cascade Framework for Text-Based Person Anomaly Search
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
When and How to Cut Classical Concerts? A Multimodal Automated Video Editing Approach
von: Gonzálbez-Biosca, Daniel, et al.
Veröffentlicht: (2025)
von: Gonzálbez-Biosca, Daniel, et al.
Veröffentlicht: (2025)
FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis
von: Liang, Feng, et al.
Veröffentlicht: (2023)
von: Liang, Feng, et al.
Veröffentlicht: (2023)
Scaling and Masking: A New Paradigm of Data Sampling for Image and Video Quality Assessment
von: Liu, Yongxu, et al.
Veröffentlicht: (2024)
von: Liu, Yongxu, et al.
Veröffentlicht: (2024)
Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation
von: Liu, Lingyu, et al.
Veröffentlicht: (2026)
von: Liu, Lingyu, et al.
Veröffentlicht: (2026)
NeR-SC: Adapting Neural Video Representation to Screen Content
von: Shi, Ruohan, et al.
Veröffentlicht: (2026)
von: Shi, Ruohan, et al.
Veröffentlicht: (2026)
ESVQA: Perceptual Quality Assessment of Egocentric Spatial Videos
von: Zhu, Xilei, et al.
Veröffentlicht: (2024)
von: Zhu, Xilei, et al.
Veröffentlicht: (2024)
Viewport Prediction for Volumetric Video Streaming by Exploring Video Saliency and Trajectory Information
von: Li, Jie, et al.
Veröffentlicht: (2023)
von: Li, Jie, et al.
Veröffentlicht: (2023)
Arena: A Patch-of-Interest ViT Inference Acceleration System for Edge-Assisted Video Analytics
von: Peng, Haosong, et al.
Veröffentlicht: (2024)
von: Peng, Haosong, et al.
Veröffentlicht: (2024)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2025)
von: Zhou, Sheng, et al.
Veröffentlicht: (2025)
Deep Shape-Texture Statistics for Completely Blind Image Quality Evaluation
von: Li, Yixuan, et al.
Veröffentlicht: (2024)
von: Li, Yixuan, et al.
Veröffentlicht: (2024)
Rethinking Multi-view Representation Learning via Distilled Disentangling
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)
StreamingEval: A Unified Evaluation Protocol towards Realistic Streaming Video Understanding
von: Tang, Guowei, et al.
Veröffentlicht: (2026)
von: Tang, Guowei, et al.
Veröffentlicht: (2026)
SSNVC: Single Stream Neural Video Compression with Implicit Temporal Information
von: Wang, Feng, et al.
Veröffentlicht: (2024)
von: Wang, Feng, et al.
Veröffentlicht: (2024)
SoccerHigh: A Benchmark Dataset for Automatic Soccer Video Summarization
von: Díaz-Juan, Artur, et al.
Veröffentlicht: (2025)
von: Díaz-Juan, Artur, et al.
Veröffentlicht: (2025)
Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-tuning
von: Xiao, Junhao, et al.
Veröffentlicht: (2026)
von: Xiao, Junhao, et al.
Veröffentlicht: (2026)
Hybrid Local-Global Context Learning for Neural Video Compression
von: Zhai, Yongqi, et al.
Veröffentlicht: (2024)
von: Zhai, Yongqi, et al.
Veröffentlicht: (2024)
Learning Video Context as Interleaved Multimodal Sequences
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
Exposure Completing for Temporally Consistent Neural High Dynamic Range Video Rendering
von: Cui, Jiahao, et al.
Veröffentlicht: (2024)
von: Cui, Jiahao, et al.
Veröffentlicht: (2024)
AIS 2024 Challenge on Video Quality Assessment of User-Generated Content: Methods and Results
von: Conde, Marcos V., et al.
Veröffentlicht: (2024)
von: Conde, Marcos V., et al.
Veröffentlicht: (2024)
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
WVSC: Wireless Video Semantic Communication with Multi-frame Compensation
von: Xie, Bingyan, et al.
Veröffentlicht: (2025)
von: Xie, Bingyan, et al.
Veröffentlicht: (2025)
ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models
von: Qu, Mengxue, et al.
Veröffentlicht: (2024)
von: Qu, Mengxue, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Evaluating Semantic Variation in Text-to-Image Synthesis: A Causal Perspective
von: Zhu, Xiangru, et al.
Veröffentlicht: (2024) -
MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models
von: Guo, Zile, et al.
Veröffentlicht: (2026) -
Rethinking Video with a Universal Event-Based Representation
von: Freeman, Andrew
Veröffentlicht: (2024) -
Multimodal Fake News Video Explanation: Dataset, Analysis and Evaluation
von: Chen, Lizhi, et al.
Veröffentlicht: (2025) -
PolySmart @ TRECVid 2024 Medical Video Question Answering
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)