VisFinEval: A Scenario-Driven Chinese Multimodal Benchmark for Holistic Financial Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zhaowei, Guo, Xin, Xia, Haotian, Zeng, Lingfeng, Lou, Fangqi, Niu, Jinyi, Li, Mengping, Qi, Qi, Li, Jiahuan, Zhang, Wei, Wang, Yinglong, Cai, Weige, Shen, Weining, Zhang, Liwen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain
von: Zeng, Lingfeng, et al.
Veröffentlicht: (2025)
von: Zeng, Lingfeng, et al.
Veröffentlicht: (2025)
Fin-R1: A Large Language Model for Financial Reasoning through Reinforcement Learning
von: Liu, Zhaowei, et al.
Veröffentlicht: (2025)
von: Liu, Zhaowei, et al.
Veröffentlicht: (2025)
FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models
von: Guo, Xin, et al.
Veröffentlicht: (2023)
von: Guo, Xin, et al.
Veröffentlicht: (2023)
UniFinEval: Towards Unified Evaluation of Financial Multimodal Models across Text, Images and Videos
von: Yang, Zhi, et al.
Veröffentlicht: (2026)
von: Yang, Zhi, et al.
Veröffentlicht: (2026)
FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments
von: Yang, Zhi, et al.
Veröffentlicht: (2026)
von: Yang, Zhi, et al.
Veröffentlicht: (2026)
LightAgent: Production-level Open-source Agentic AI Framework
von: Cai, Weige, et al.
Veröffentlicht: (2025)
von: Cai, Weige, et al.
Veröffentlicht: (2025)
FinEval-KR: A Financial Domain Evaluation Framework for Large Language Models' Knowledge and Reasoning
von: Dou, Shaoyu, et al.
Veröffentlicht: (2025)
von: Dou, Shaoyu, et al.
Veröffentlicht: (2025)
FinTeam: A Multi-Agent Collaborative Intelligence System for Comprehensive Financial Scenarios
von: Wu, Yingqian, et al.
Veröffentlicht: (2025)
von: Wu, Yingqian, et al.
Veröffentlicht: (2025)
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios
von: Hou, Yutao, et al.
Veröffentlicht: (2026)
von: Hou, Yutao, et al.
Veröffentlicht: (2026)
UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation
von: Li, Yi, et al.
Veröffentlicht: (2025)
von: Li, Yi, et al.
Veröffentlicht: (2025)
FinBen: A Holistic Financial Benchmark for Large Language Models
von: Xie, Qianqian, et al.
Veröffentlicht: (2024)
von: Xie, Qianqian, et al.
Veröffentlicht: (2024)
FinMaster: A Holistic Benchmark for Mastering Full-Pipeline Financial Workflows with LLMs
von: Jiang, Junzhe, et al.
Veröffentlicht: (2025)
von: Jiang, Junzhe, et al.
Veröffentlicht: (2025)
FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step Computation
von: Tang, Zichen, et al.
Veröffentlicht: (2025)
von: Tang, Zichen, et al.
Veröffentlicht: (2025)
LLM-Driven Scenario-Aware Planning for Autonomous Driving
von: Li, He, et al.
Veröffentlicht: (2026)
von: Li, He, et al.
Veröffentlicht: (2026)
PhysCorr: Dual-Reward DPO for Physics-Constrained Text-to-Video Generation with Automated Preference Selection
von: Wang, Peiyao, et al.
Veröffentlicht: (2025)
von: Wang, Peiyao, et al.
Veröffentlicht: (2025)
FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks
von: Cao, Yupeng, et al.
Veröffentlicht: (2026)
von: Cao, Yupeng, et al.
Veröffentlicht: (2026)
FinReflectKG -- EvalBench: Benchmarking Financial KG with Multi-Dimensional Evaluation
von: Dimino, Fabrizio, et al.
Veröffentlicht: (2025)
von: Dimino, Fabrizio, et al.
Veröffentlicht: (2025)
FinDebate: Multi-Agent Collaborative Intelligence for Financial Analysis
von: Cai, Tianshi, et al.
Veröffentlicht: (2025)
von: Cai, Tianshi, et al.
Veröffentlicht: (2025)
Comparison of complications and indwelling time in midline catheters versus central venous catheters: A systematic review and meta‐analysis
von: Xin Li, et al.
Veröffentlicht: (2024)
von: Xin Li, et al.
Veröffentlicht: (2024)
BizFinBench.v2: A Unified Dual-Mode Bilingual Benchmark for Expert-Level Financial Capability Alignment
von: Guo, Xin, et al.
Veröffentlicht: (2026)
von: Guo, Xin, et al.
Veröffentlicht: (2026)
VisEval: A Benchmark for Data Visualization in the Era of Large Language Models
von: Chen, Nan, et al.
Veröffentlicht: (2024)
von: Chen, Nan, et al.
Veröffentlicht: (2024)
Sports Intelligence: Assessing the Sports Understanding Capabilities of Language Models through Question Answering from Text to Video
von: Yang, Zhengbang, et al.
Veröffentlicht: (2024)
von: Yang, Zhengbang, et al.
Veröffentlicht: (2024)
SceneJailEval: A Scenario-Adaptive Multi-Dimensional Framework for Jailbreak Evaluation
von: Jiang, Lai, et al.
Veröffentlicht: (2025)
von: Jiang, Lai, et al.
Veröffentlicht: (2025)
Agentar-Fin-R1: Enhancing Financial Intelligence through Domain Expertise, Training Efficiency, and Advanced Reasoning
von: Zheng, Yanjun, et al.
Veröffentlicht: (2025)
von: Zheng, Yanjun, et al.
Veröffentlicht: (2025)
FinSQL: Model-Agnostic LLMs-based Text-to-SQL Framework for Financial Analysis
von: Zhang, Chao, et al.
Veröffentlicht: (2024)
von: Zhang, Chao, et al.
Veröffentlicht: (2024)
SNFinLLM: Systematic and Nuanced Financial Domain Adaptation of Chinese Large Language Models
von: Zhao, Shujuan, et al.
Veröffentlicht: (2024)
von: Zhao, Shujuan, et al.
Veröffentlicht: (2024)
MTFinEval:A Multi-domain Chinese Financial Benchmark with Eurypalynous questions
von: Liu, Xinyu, et al.
Veröffentlicht: (2024)
von: Liu, Xinyu, et al.
Veröffentlicht: (2024)
VisMin: Visual Minimal-Change Understanding
von: Awal, Rabiul, et al.
Veröffentlicht: (2024)
von: Awal, Rabiul, et al.
Veröffentlicht: (2024)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
von: Jiang, Botian, et al.
Veröffentlicht: (2024)
von: Jiang, Botian, et al.
Veröffentlicht: (2024)
FinKario: Event-Enhanced Automated Construction of Financial Knowledge Graph
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
How Class Influences the Ethnic Identity of Chinese Immigrants in the UK: Citizenship, Work, and Solidarity
von: Zhaowei Yin
Veröffentlicht: (2026)
von: Zhaowei Yin
Veröffentlicht: (2026)
Holistic Adversarially Robust Pruning
von: Zhao, Qi, et al.
Veröffentlicht: (2024)
von: Zhao, Qi, et al.
Veröffentlicht: (2024)
CoVis: A Collaborative Framework for Fine-grained Graphic Visual Understanding
von: Deng, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Deng, Xiaoyu, et al.
Veröffentlicht: (2024)
FinCPRG: A Bidirectional Generation Pipeline for Hierarchical Queries and Rich Relevance in Financial Chinese Passage Retrieval
von: Xu, Xuan, et al.
Veröffentlicht: (2025)
von: Xu, Xuan, et al.
Veröffentlicht: (2025)
M$^3$FinMeeting: A Multilingual, Multi-Sector, and Multi-Task Financial Meeting Understanding Evaluation Dataset
von: Zhu, Jie, et al.
Veröffentlicht: (2025)
von: Zhu, Jie, et al.
Veröffentlicht: (2025)
TTP: Test-Time Padding for Adversarial Detection and Robust Adaptation on Vision-Language Models
von: Li, Zhiwei, et al.
Veröffentlicht: (2025)
von: Li, Zhiwei, et al.
Veröffentlicht: (2025)
TongGu: Mastering Classical Chinese Understanding with Knowledge-Grounded Large Language Models
von: Cao, Jiahuan, et al.
Veröffentlicht: (2024)
von: Cao, Jiahuan, et al.
Veröffentlicht: (2024)
FinGuard: Detecting Financial Regulatory Non-Compliance in LLM Interactions
von: Dou, Huaixia, et al.
Veröffentlicht: (2026)
von: Dou, Huaixia, et al.
Veröffentlicht: (2026)
PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
von: Zhou, Lingfeng, et al.
Veröffentlicht: (2025)
von: Zhou, Lingfeng, et al.
Veröffentlicht: (2025)
A Holistic Understanding of Chinese Children's Artificial Intelligence Literacy: An Exploratory Study
von: Ping Wang, et al.
Veröffentlicht: (2025)
von: Ping Wang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain
von: Zeng, Lingfeng, et al.
Veröffentlicht: (2025) -
Fin-R1: A Large Language Model for Financial Reasoning through Reinforcement Learning
von: Liu, Zhaowei, et al.
Veröffentlicht: (2025) -
FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models
von: Guo, Xin, et al.
Veröffentlicht: (2023) -
UniFinEval: Towards Unified Evaluation of Financial Multimodal Models across Text, Images and Videos
von: Yang, Zhi, et al.
Veröffentlicht: (2026) -
FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments
von: Yang, Zhi, et al.
Veröffentlicht: (2026)