Rethinking Human Evaluation Protocol for Text-to-Video Models: Enhancing Reliability,Reproducibility, and Practicality
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Tianle, Ma, Langtian, Yan, Yuchen, Zhang, Yuchen, Wang, Kai, Yang, Yue, Guo, Ziyao, Shao, Wenqi, You, Yang, Qiao, Yu, Luo, Ping, Zhang, Kaipeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Two Trades is not Baffled: Condensing Graph via Crafting Rational Gradient Matching
von: Zhang, Tianle, et al.
Veröffentlicht: (2024)
von: Zhang, Tianle, et al.
Veröffentlicht: (2024)
Enhance-A-Video: Better Generated Video for Free
von: Luo, Yang, et al.
Veröffentlicht: (2025)
von: Luo, Yang, et al.
Veröffentlicht: (2025)
DiffAgent: Fast and Accurate Text-to-Image API Selection with Large Language Model
von: Zhao, Lirui, et al.
Veröffentlicht: (2024)
von: Zhao, Lirui, et al.
Veröffentlicht: (2024)
ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Capability for Large Vision-Language Models
von: Liu, Shuo, et al.
Veröffentlicht: (2024)
von: Liu, Shuo, et al.
Veröffentlicht: (2024)
PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
Dynamic Multimodal Evaluation with Flexible Complexity by Vision-Language Bootstrapping
von: Yang, Yue, et al.
Veröffentlicht: (2024)
von: Yang, Yue, et al.
Veröffentlicht: (2024)
Diffree: Text-Guided Shape Free Object Inpainting with Diffusion Model
von: Zhao, Lirui, et al.
Veröffentlicht: (2024)
von: Zhao, Lirui, et al.
Veröffentlicht: (2024)
MLLMs-Augmented Visual-Language Representation Learning
von: Liu, Yanqing, et al.
Veröffentlicht: (2023)
von: Liu, Yanqing, et al.
Veröffentlicht: (2023)
Navigating Complexity: Toward Lossless Graph Condensation via Expanding Window Matching
von: Zhang, Yuchen, et al.
Veröffentlicht: (2024)
von: Zhang, Yuchen, et al.
Veröffentlicht: (2024)
Efficient High-Resolution Visual Representation Learning with State Space Model for Human Pose Estimation
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
Position: Towards Implicit Prompt For Text-To-Image Models
von: Yang, Yue, et al.
Veröffentlicht: (2024)
von: Yang, Yue, et al.
Veröffentlicht: (2024)
B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
von: Xie, Yuxuan, et al.
Veröffentlicht: (2024)
von: Xie, Yuxuan, et al.
Veröffentlicht: (2024)
Towards Lossless Dataset Distillation via Difficulty-Aligned Trajectory Matching
von: Guo, Ziyao, et al.
Veröffentlicht: (2023)
von: Guo, Ziyao, et al.
Veröffentlicht: (2023)
ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
Open-Vocabulary Animal Keypoint Detection with Semantic-feature Matching
von: Zhang, Hao, et al.
Veröffentlicht: (2023)
von: Zhang, Hao, et al.
Veröffentlicht: (2023)
Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
Towards Holistic Language-video Representation: the language model-enhanced MSR-Video to Text Dataset
von: Yang, Yuchen, et al.
Veröffentlicht: (2024)
von: Yang, Yuchen, et al.
Veröffentlicht: (2024)
SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement
von: Lin, Yuqi, et al.
Veröffentlicht: (2025)
von: Lin, Yuqi, et al.
Veröffentlicht: (2025)
BESA: Pruning Large Language Models with Blockwise Parameter-Efficient Sparsity Allocation
von: Xu, Peng, et al.
Veröffentlicht: (2024)
von: Xu, Peng, et al.
Veröffentlicht: (2024)
Improving Autoregressive Image Generation through Coarse-to-Fine Token Prediction
von: Guo, Ziyao, et al.
Veröffentlicht: (2025)
von: Guo, Ziyao, et al.
Veröffentlicht: (2025)
Data Adaptive Traceback for Vision-Language Foundation Models in Image Classification
von: Peng, Wenshuo, et al.
Veröffentlicht: (2024)
von: Peng, Wenshuo, et al.
Veröffentlicht: (2024)
Rethinking the Evaluation Protocol of Domain Generalization
von: Yu, Han, et al.
Veröffentlicht: (2023)
von: Yu, Han, et al.
Veröffentlicht: (2023)
Prioritize Alignment in Dataset Distillation
von: Li, Zekai, et al.
Veröffentlicht: (2024)
von: Li, Zekai, et al.
Veröffentlicht: (2024)
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024)
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024)
Training-Free and Interpretable Hateful Video Detection via Multi-stage Adversarial Reasoning
von: Yang, Shuonan, et al.
Veröffentlicht: (2026)
von: Yang, Shuonan, et al.
Veröffentlicht: (2026)
T3M: Text Guided 3D Human Motion Synthesis from Speech
von: Peng, Wenshuo, et al.
Veröffentlicht: (2024)
von: Peng, Wenshuo, et al.
Veröffentlicht: (2024)
SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge
von: Li, Chuanhao, et al.
Veröffentlicht: (2024)
von: Li, Chuanhao, et al.
Veröffentlicht: (2024)
TinyLVLM-eHub: Towards Comprehensive and Efficient Evaluation for Large Vision-Language Models
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
Are AI-Generated Text Detectors Robust to Adversarial Perturbations?
von: Huang, Guanhua, et al.
Veröffentlicht: (2024)
von: Huang, Guanhua, et al.
Veröffentlicht: (2024)
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation
von: Zhou, Pengfei, et al.
Veröffentlicht: (2024)
von: Zhou, Pengfei, et al.
Veröffentlicht: (2024)
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
von: Feng, Yukang, et al.
Veröffentlicht: (2025)
von: Feng, Yukang, et al.
Veröffentlicht: (2025)
Enhanced Multimodal Hate Video Detection via Channel-wise and Modality-wise Fusion
von: Zhang, Yinghui, et al.
Veröffentlicht: (2025)
von: Zhang, Yinghui, et al.
Veröffentlicht: (2025)
Adapting LLaMA Decoder to Vision Transformer
von: Wang, Jiahao, et al.
Veröffentlicht: (2024)
von: Wang, Jiahao, et al.
Veröffentlicht: (2024)
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
SPOR: A Comprehensive and Practical Evaluation Method for Compositional Generalization in Data-to-Text Generation
von: Xu, Ziyao, et al.
Veröffentlicht: (2024)
von: Xu, Ziyao, et al.
Veröffentlicht: (2024)
Tarsier: Recipes for Training and Evaluating Large Video Description Models
von: Wang, Jiawei, et al.
Veröffentlicht: (2024)
von: Wang, Jiawei, et al.
Veröffentlicht: (2024)
Needle In A Multimodal Haystack
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Two Trades is not Baffled: Condensing Graph via Crafting Rational Gradient Matching
von: Zhang, Tianle, et al.
Veröffentlicht: (2024) -
Enhance-A-Video: Better Generated Video for Free
von: Luo, Yang, et al.
Veröffentlicht: (2025) -
DiffAgent: Fast and Accurate Text-to-Image API Selection with Large Language Model
von: Zhao, Lirui, et al.
Veröffentlicht: (2024) -
ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Capability for Large Vision-Language Models
von: Liu, Shuo, et al.
Veröffentlicht: (2024) -
PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)