FCMBench-Video: Benchmarking Document Video Intelligence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cui, Runze, Shang, Fangxin, Yang, Yehui, Yang, Qing, Xu, Yanwu, Chen, Tao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FCMBench: The First Large-scale Financial Credit Multimodal Benchmark for Real-world Applications
von: Yang, Yehui, et al.
Veröffentlicht: (2026)
von: Yang, Yehui, et al.
Veröffentlicht: (2026)
VMU-Diff: A Coarse-to-fine Multi-source Data Fusion Framework for Precipitation Nowcasting
von: Shi, Chunlei, et al.
Veröffentlicht: (2026)
von: Shi, Chunlei, et al.
Veröffentlicht: (2026)
BatSort: Enhanced Battery Classification with Transfer Learning for Battery Sorting and Recycling
von: Zhao, Yunyi, et al.
Veröffentlicht: (2024)
von: Zhao, Yunyi, et al.
Veröffentlicht: (2024)
Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
von: Liu, Fuxiao, et al.
Veröffentlicht: (2023)
von: Liu, Fuxiao, et al.
Veröffentlicht: (2023)
Decoding Scientific Experimental Images: The SPUR Benchmark for Perception, Understanding, and Reasoning
von: Ding, Junpeng, et al.
Veröffentlicht: (2026)
von: Ding, Junpeng, et al.
Veröffentlicht: (2026)
AUTV: Creating Underwater Video Datasets with Pixel-wise Annotations
von: Truong, Quang Trung, et al.
Veröffentlicht: (2025)
von: Truong, Quang Trung, et al.
Veröffentlicht: (2025)
FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step Computation
von: Tang, Zichen, et al.
Veröffentlicht: (2025)
von: Tang, Zichen, et al.
Veröffentlicht: (2025)
FinMTM: A Multi-Turn Multimodal Benchmark for Financial Reasoning and Agent Evaluation
von: Zhang, Chenxi, et al.
Veröffentlicht: (2026)
von: Zhang, Chenxi, et al.
Veröffentlicht: (2026)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
VC-Bench: Pioneering the Video Connecting Benchmark with a Dataset and Evaluation Metrics
von: Yin, Zhiyu, et al.
Veröffentlicht: (2026)
von: Yin, Zhiyu, et al.
Veröffentlicht: (2026)
VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models
von: Yang, Haibo, et al.
Veröffentlicht: (2024)
von: Yang, Haibo, et al.
Veröffentlicht: (2024)
VIoTGPT: Learning to Schedule Vision Tools in LLMs towards Intelligent Video Internet of Things
von: Zhong, Yaoyao, et al.
Veröffentlicht: (2023)
von: Zhong, Yaoyao, et al.
Veröffentlicht: (2023)
Face Consistency Benchmark for GenAI Video
von: Podstawski, Michal, et al.
Veröffentlicht: (2025)
von: Podstawski, Michal, et al.
Veröffentlicht: (2025)
TAVGBench: Benchmarking Text to Audible-Video Generation
von: Mao, Yuxin, et al.
Veröffentlicht: (2024)
von: Mao, Yuxin, et al.
Veröffentlicht: (2024)
Multimodal growth and development assessment model
von: Li, Ying, et al.
Veröffentlicht: (2024)
von: Li, Ying, et al.
Veröffentlicht: (2024)
MSMF: Multi-Scale Multi-Modal Fusion for Enhanced Stock Market Prediction
von: Qin, Jiahao
Veröffentlicht: (2024)
von: Qin, Jiahao
Veröffentlicht: (2024)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
neuralCAD-Edit: An Expert Benchmark for Multimodal-Instructed 3D CAD Model Editing
von: Perrett, Toby, et al.
Veröffentlicht: (2026)
von: Perrett, Toby, et al.
Veröffentlicht: (2026)
PRVR: Partially Relevant Video Retrieval
von: Chen, Xianke, et al.
Veröffentlicht: (2022)
von: Chen, Xianke, et al.
Veröffentlicht: (2022)
Video Watermarking: Safeguarding Your Video from (Unauthorized) Annotations by Video-based LLMs
von: Li, Jinmin, et al.
Veröffentlicht: (2024)
von: Li, Jinmin, et al.
Veröffentlicht: (2024)
FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging
von: Tang, Zichen, et al.
Veröffentlicht: (2025)
von: Tang, Zichen, et al.
Veröffentlicht: (2025)
FinAudio: A Benchmark for Audio Large Language Models in Financial Applications
von: Cao, Yupeng, et al.
Veröffentlicht: (2025)
von: Cao, Yupeng, et al.
Veröffentlicht: (2025)
CBVS: A Large-Scale Chinese Image-Text Benchmark for Real-World Short Video Search Scenarios
von: Qiao, Xiangshuo, et al.
Veröffentlicht: (2024)
von: Qiao, Xiangshuo, et al.
Veröffentlicht: (2024)
Looking Backward: Streaming Video-to-Video Translation with Feature Banks
von: Liang, Feng, et al.
Veröffentlicht: (2024)
von: Liang, Feng, et al.
Veröffentlicht: (2024)
PolySmart @ TRECVid 2024 Video Captioning (VTT)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
Hybrid Local-Global Context Learning for Neural Video Compression
von: Zhai, Yongqi, et al.
Veröffentlicht: (2024)
von: Zhai, Yongqi, et al.
Veröffentlicht: (2024)
GaussianCAD: Robust Self-Supervised CAD Reconstruction from Three Orthographic Views Using 3D Gaussian Splatting
von: Zhou, Zheng, et al.
Veröffentlicht: (2025)
von: Zhou, Zheng, et al.
Veröffentlicht: (2025)
TRIP: Temporal Residual Learning with Image Noise Prior for Image-to-Video Diffusion Models
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2024)
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2024)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
KeyVideoLLM: Towards Large-scale Video Keyframe Selection
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
PolySmart @ TRECVid 2024 Medical Video Question Answering
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing
von: Zhao, Pengcheng, et al.
Veröffentlicht: (2024)
von: Zhao, Pengcheng, et al.
Veröffentlicht: (2024)
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
GLObal Building heights for Urban Studies (UT-GLOBUS) for city- and street- scale urban simulations: Development and first applications
von: Kamath, Harsh G., et al.
Veröffentlicht: (2022)
von: Kamath, Harsh G., et al.
Veröffentlicht: (2022)
Human Motion Video Generation: A Survey
von: Xue, Haiwei, et al.
Veröffentlicht: (2025)
von: Xue, Haiwei, et al.
Veröffentlicht: (2025)
MFQE 2.0: A New Approach for Multi-frame Quality Enhancement on Compressed Video
von: Xing, Qunliang, et al.
Veröffentlicht: (2019)
von: Xing, Qunliang, et al.
Veröffentlicht: (2019)
Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
von: Gao, Jiayi, et al.
Veröffentlicht: (2025)
von: Gao, Jiayi, et al.
Veröffentlicht: (2025)
Generative Frame Sampler for Long Video Understanding
von: Yao, Linli, et al.
Veröffentlicht: (2025)
von: Yao, Linli, et al.
Veröffentlicht: (2025)
Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning
von: Zhang, Long, et al.
Veröffentlicht: (2025)
von: Zhang, Long, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FCMBench: The First Large-scale Financial Credit Multimodal Benchmark for Real-world Applications
von: Yang, Yehui, et al.
Veröffentlicht: (2026) -
VMU-Diff: A Coarse-to-fine Multi-source Data Fusion Framework for Precipitation Nowcasting
von: Shi, Chunlei, et al.
Veröffentlicht: (2026) -
BatSort: Enhanced Battery Classification with Transfer Learning for Battery Sorting and Recycling
von: Zhao, Yunyi, et al.
Veröffentlicht: (2024) -
Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
von: Liu, Fuxiao, et al.
Veröffentlicht: (2023) -
Decoding Scientific Experimental Images: The SPUR Benchmark for Perception, Understanding, and Reasoning
von: Ding, Junpeng, et al.
Veröffentlicht: (2026)