FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain
Fuente:
arXiv
Saved in:
| Main Authors: | Zeng, Lingfeng, Lou, Fangqi, Wang, Zixuan, Xu, Jiajie, Niu, Jinyi, Li, Mengping, Dong, Yifan, Qi, Qi, Zhang, Wei, Yang, Ziwei, Han, Jun, Feng, Ruilun, Hu, Ruiqi, Zhang, Lejie, Feng, Zhengbo, Ren, Yicheng, Guo, Xin, Liu, Zhaowei, Cheng, Dongpo, Cai, Weige, Zhang, Liwen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VisFinEval: A Scenario-Driven Chinese Multimodal Benchmark for Holistic Financial Understanding
by: Liu, Zhaowei, et al.
Published: (2025)
by: Liu, Zhaowei, et al.
Published: (2025)
Fin-R1: A Large Language Model for Financial Reasoning through Reinforcement Learning
by: Liu, Zhaowei, et al.
Published: (2025)
by: Liu, Zhaowei, et al.
Published: (2025)
UniFinEval: Towards Unified Evaluation of Financial Multimodal Models across Text, Images and Videos
by: Yang, Zhi, et al.
Published: (2026)
by: Yang, Zhi, et al.
Published: (2026)
FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments
by: Yang, Zhi, et al.
Published: (2026)
by: Yang, Zhi, et al.
Published: (2026)
FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models
by: Guo, Xin, et al.
Published: (2023)
by: Guo, Xin, et al.
Published: (2023)
FinEval-KR: A Financial Domain Evaluation Framework for Large Language Models' Knowledge and Reasoning
by: Dou, Shaoyu, et al.
Published: (2025)
by: Dou, Shaoyu, et al.
Published: (2025)
LightAgent: Production-level Open-source Agentic AI Framework
by: Cai, Weige, et al.
Published: (2025)
by: Cai, Weige, et al.
Published: (2025)
FinSight: Towards Real-World Financial Deep Research
by: Jin, Jiajie, et al.
Published: (2025)
by: Jin, Jiajie, et al.
Published: (2025)
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios
by: Hou, Yutao, et al.
Published: (2026)
by: Hou, Yutao, et al.
Published: (2026)
FinTeam: A Multi-Agent Collaborative Intelligence System for Comprehensive Financial Scenarios
by: Wu, Yingqian, et al.
Published: (2025)
by: Wu, Yingqian, et al.
Published: (2025)
BizFinBench.v2: A Unified Dual-Mode Bilingual Benchmark for Expert-Level Financial Capability Alignment
by: Guo, Xin, et al.
Published: (2026)
by: Guo, Xin, et al.
Published: (2026)
FinGuard: Detecting Financial Regulatory Non-Compliance in LLM Interactions
by: Dou, Huaixia, et al.
Published: (2026)
by: Dou, Huaixia, et al.
Published: (2026)
FinDebate: Multi-Agent Collaborative Intelligence for Financial Analysis
by: Cai, Tianshi, et al.
Published: (2025)
by: Cai, Tianshi, et al.
Published: (2025)
SNFinLLM: Systematic and Nuanced Financial Domain Adaptation of Chinese Large Language Models
by: Zhao, Shujuan, et al.
Published: (2024)
by: Zhao, Shujuan, et al.
Published: (2024)
GAIA
by: Ernesto Cardenal
Published: (2007)
by: Ernesto Cardenal
Published: (2007)
FinSQL: Model-Agnostic LLMs-based Text-to-SQL Framework for Financial Analysis
by: Zhang, Chao, et al.
Published: (2024)
by: Zhang, Chao, et al.
Published: (2024)
FinS-Pilot: A Benchmark for Online Financial RAG System
by: Wang, Feng, et al.
Published: (2025)
by: Wang, Feng, et al.
Published: (2025)
GAIA: A General Agency Interaction Architecture for LLM-Human B2B Negotiation & Screening
by: Zhao, Siming, et al.
Published: (2025)
by: Zhao, Siming, et al.
Published: (2025)
Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models
by: Zhu, Jie, et al.
Published: (2025)
by: Zhu, Jie, et al.
Published: (2025)
Spectrally-large scale geometry via set-heaviness
by: Feng, Qi, et al.
Published: (2025)
by: Feng, Qi, et al.
Published: (2025)
Spectrally-large scale geometry in cotangent bundles
by: Feng, Qi, et al.
Published: (2024)
by: Feng, Qi, et al.
Published: (2024)
Notes on symplectic squeezing in $T^* \mathbb T^n$ and spectra of Finsler dynamics
by: Feng, Qi, et al.
Published: (2024)
by: Feng, Qi, et al.
Published: (2024)
Agentar-Fin-R1: Enhancing Financial Intelligence through Domain Expertise, Training Efficiency, and Advanced Reasoning
by: Zheng, Yanjun, et al.
Published: (2025)
by: Zheng, Yanjun, et al.
Published: (2025)
FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis
by: Zhu, Fengbin, et al.
Published: (2025)
by: Zhu, Fengbin, et al.
Published: (2025)
FinMTM: A Multi-Turn Multimodal Benchmark for Financial Reasoning and Agent Evaluation
by: Zhang, Chenxi, et al.
Published: (2026)
by: Zhang, Chenxi, et al.
Published: (2026)
Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset
by: Chen, Qian, et al.
Published: (2026)
by: Chen, Qian, et al.
Published: (2026)
Electrical Impedance Spectroscopy in Agricultural Food Quality Detection
by: Benhua Zhang, et al.
Published: (2025)
by: Benhua Zhang, et al.
Published: (2025)
FinToolSyn: A forward synthesis Framework for Financial Tool-Use Dialogue Data with Dynamic Tool Retrieval
by: Huang, Caishuang, et al.
Published: (2026)
by: Huang, Caishuang, et al.
Published: (2026)
OmniGAIA: Towards Native Omni-Modal AI Agents
by: Li, Xiaoxi, et al.
Published: (2026)
by: Li, Xiaoxi, et al.
Published: (2026)
FinCARDS: Card-Based Analyst Reranking for Financial Document Question Answering
by: Zhou, Yixi, et al.
Published: (2026)
by: Zhou, Yixi, et al.
Published: (2026)
FinAnchor: Aligned Multi-Model Representations for Financial Prediction
by: He, Zirui, et al.
Published: (2026)
by: He, Zirui, et al.
Published: (2026)
Comparison of complications and indwelling time in midline catheters versus central venous catheters: A systematic review and meta‐analysis
by: Xin Li, et al.
Published: (2024)
by: Xin Li, et al.
Published: (2024)
How Class Influences the Ethnic Identity of Chinese Immigrants in the UK: Citizenship, Work, and Solidarity
by: Zhaowei Yin
Published: (2026)
by: Zhaowei Yin
Published: (2026)
Azimuthal distribution of exponential format for particle collective motions in heavy-ion collisions under asynchronous assumption
by: Feng, Yicheng
Published: (2023)
by: Feng, Yicheng
Published: (2023)
FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol
by: Zhu, Jie, et al.
Published: (2026)
by: Zhu, Jie, et al.
Published: (2026)
Optimizing Sentence Embedding with Pseudo-Labeling and Model Ensembles: A Hierarchical Framework for Enhanced NLP Tasks
by: Liu, Ziwei, et al.
Published: (2025)
by: Liu, Ziwei, et al.
Published: (2025)
A Comprehensive Framework for Semantic Similarity Analysis of Human and AI-Generated Text Using Transformer Architectures and Ensemble Techniques
by: Gao, Lifu, et al.
Published: (2025)
by: Gao, Lifu, et al.
Published: (2025)
FinTagging: Benchmarking LLMs for Extracting and Structuring Financial Information
by: Wang, Yan, et al.
Published: (2025)
by: Wang, Yan, et al.
Published: (2025)
Spider-Sense: Intrinsic Risk Sensing for Efficient Agent Defense with Hierarchical Adaptive Screening
by: Yu, Zhenxiong, et al.
Published: (2026)
by: Yu, Zhenxiong, et al.
Published: (2026)
Financial Sentiment Analysis on News and Reports Using Large Language Models and FinBERT
by: Shen, Yanxin, et al.
Published: (2024)
by: Shen, Yanxin, et al.
Published: (2024)
Similar Items
-
VisFinEval: A Scenario-Driven Chinese Multimodal Benchmark for Holistic Financial Understanding
by: Liu, Zhaowei, et al.
Published: (2025) -
Fin-R1: A Large Language Model for Financial Reasoning through Reinforcement Learning
by: Liu, Zhaowei, et al.
Published: (2025) -
UniFinEval: Towards Unified Evaluation of Financial Multimodal Models across Text, Images and Videos
by: Yang, Zhi, et al.
Published: (2026) -
FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments
by: Yang, Zhi, et al.
Published: (2026) -
FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models
by: Guo, Xin, et al.
Published: (2023)