BizCompass: Benchmarking the Reasoning Capabilities of LLMs in Business Knowledge and Applications
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hao, Jianing, Wu, Yuhe, Xu, Yuanjian, Meng, Shichang, Yuan, Shuai, Zeng, Wei, Wang, Zixuan, Zhang, Guang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMs
von: Lu, Guilong, et al.
Veröffentlicht: (2025)
von: Lu, Guilong, et al.
Veröffentlicht: (2025)
Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks
von: Bigeard, Antoine, et al.
Veröffentlicht: (2025)
von: Bigeard, Antoine, et al.
Veröffentlicht: (2025)
Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs
von: Guo, Xingang, et al.
Veröffentlicht: (2025)
von: Guo, Xingang, et al.
Veröffentlicht: (2025)
Affective-ROPTester: Capability and Bias Analysis of LLMs in Predicting Retinopathy of Prematurity
von: Zhao, Shuai, et al.
Veröffentlicht: (2025)
von: Zhao, Shuai, et al.
Veröffentlicht: (2025)
Flow of Knowledge: Federated Fine-Tuning of LLMs in Healthcare under Non-IID Conditions
von: Chen, Zeyu, et al.
Veröffentlicht: (2025)
von: Chen, Zeyu, et al.
Veröffentlicht: (2025)
Fusing LLMs and KGs for Formal Causal Reasoning behind Financial Risk Contagion
von: Yu, Guanyuan, et al.
Veröffentlicht: (2024)
von: Yu, Guanyuan, et al.
Veröffentlicht: (2024)
FinBoardBench: Benchmarking Dynamic Wealth Management and Strategic Financial Reasoning of LLMs via Board Game Simulations
von: Hu, Xuesi, et al.
Veröffentlicht: (2026)
von: Hu, Xuesi, et al.
Veröffentlicht: (2026)
Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications
von: Huang, Jimin, et al.
Veröffentlicht: (2024)
von: Huang, Jimin, et al.
Veröffentlicht: (2024)
STRIDE: Post-Training LLMs to Reason and Refine Bio-Sequences via Edit Trajectories
von: Zhang, Daiheng, et al.
Veröffentlicht: (2026)
von: Zhang, Daiheng, et al.
Veröffentlicht: (2026)
FinDocMRE: A Benchmark for Document-Level Financial Multimodal Reasoning Evaluation
von: Zhu, Jiayong, et al.
Veröffentlicht: (2026)
von: Zhu, Jiayong, et al.
Veröffentlicht: (2026)
V2T-CoT: From Vision to Text Chain-of-Thought for Medical Reasoning and Diagnosis
von: Wang, Yuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuan, et al.
Veröffentlicht: (2025)
FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and Challenging
von: Tang, Zichen, et al.
Veröffentlicht: (2025)
von: Tang, Zichen, et al.
Veröffentlicht: (2025)
Chem-R: Learning to Reason as a Chemist
von: Wang, Weida, et al.
Veröffentlicht: (2025)
von: Wang, Weida, et al.
Veröffentlicht: (2025)
AFBench: A Large-scale Benchmark for Airfoil Design
von: Liu, Jian, et al.
Veröffentlicht: (2024)
von: Liu, Jian, et al.
Veröffentlicht: (2024)
Toward a Mapping of Capability and Skill Models using Asset Administration Shells and Ontologies
von: da Silva, Luis Miguel Vieira, et al.
Veröffentlicht: (2023)
von: da Silva, Luis Miguel Vieira, et al.
Veröffentlicht: (2023)
Update Strategy for Channel Knowledge Map in Complex Environments
von: Wang, Ting, et al.
Veröffentlicht: (2025)
von: Wang, Ting, et al.
Veröffentlicht: (2025)
QuantBench: Benchmarking AI Methods for Quantitative Investment
von: Wang, Saizhuo, et al.
Veröffentlicht: (2025)
von: Wang, Saizhuo, et al.
Veröffentlicht: (2025)
VisFinEval: A Scenario-Driven Chinese Multimodal Benchmark for Holistic Financial Understanding
von: Liu, Zhaowei, et al.
Veröffentlicht: (2025)
von: Liu, Zhaowei, et al.
Veröffentlicht: (2025)
RoadBench: A Vision-Language Foundation Model and Benchmark for Road Damage Understanding
von: Xiao, Xi, et al.
Veröffentlicht: (2025)
von: Xiao, Xi, et al.
Veröffentlicht: (2025)
AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets
von: Fan, Tianyu, et al.
Veröffentlicht: (2025)
von: Fan, Tianyu, et al.
Veröffentlicht: (2025)
Investigating the Surrogate Modeling Capabilities of Continuous Time Echo State Networks
von: Bhatnagar, Saakaar
Veröffentlicht: (2023)
von: Bhatnagar, Saakaar
Veröffentlicht: (2023)
MetaBench: A Multi-task Benchmark for Assessing LLMs in Metabolomics
von: Lu, Yuxing, et al.
Veröffentlicht: (2025)
von: Lu, Yuxing, et al.
Veröffentlicht: (2025)
ForTune: Running Offline Scenarios to Estimate Impact on Business Metrics
von: Dupret, Georges, et al.
Veröffentlicht: (2024)
von: Dupret, Georges, et al.
Veröffentlicht: (2024)
SECQUE: A Benchmark for Evaluating Real-World Financial Analysis Capabilities
von: Yoash, Noga Ben, et al.
Veröffentlicht: (2025)
von: Yoash, Noga Ben, et al.
Veröffentlicht: (2025)
PhysioFormer: Integrating Multimodal Physiological Signals and Symbolic Regression for Explainable Affective State Prediction
von: Wang, Zhifeng, et al.
Veröffentlicht: (2024)
von: Wang, Zhifeng, et al.
Veröffentlicht: (2024)
KARMA: Leveraging Multi-Agent LLMs for Automated Knowledge Graph Enrichment
von: Lu, Yuxing, et al.
Veröffentlicht: (2025)
von: Lu, Yuxing, et al.
Veröffentlicht: (2025)
KATO: Knowledge Alignment and Transfer for Transistor Sizing of Different Design and Technology
von: Xing, Wei W., et al.
Veröffentlicht: (2024)
von: Xing, Wei W., et al.
Veröffentlicht: (2024)
Investigate the efficiency of incompressible flow simulations on CPUs and GPUs with BSAMR
von: Liu, Dewen, et al.
Veröffentlicht: (2024)
von: Liu, Dewen, et al.
Veröffentlicht: (2024)
SGLDBench: A Benchmark Suite for Stress-Guided Lightweight 3D Designs
von: Wang, Junpeng, et al.
Veröffentlicht: (2025)
von: Wang, Junpeng, et al.
Veröffentlicht: (2025)
Identifying Evidence Subgraphs for Financial Risk Detection via Graph Counterfactual and Factual Reasoning
von: Du, Huaming, et al.
Veröffentlicht: (2025)
von: Du, Huaming, et al.
Veröffentlicht: (2025)
Benchmarking formalisms for dynamic structure system Modeling and Simulation
von: Attia, Aya, et al.
Veröffentlicht: (2024)
von: Attia, Aya, et al.
Veröffentlicht: (2024)
The Influence of Biomedical Research on Future Business Funding: Analyzing Scientific Impact and Content in Industrial Investments
von: Khanmohammadi, Reza, et al.
Veröffentlicht: (2024)
von: Khanmohammadi, Reza, et al.
Veröffentlicht: (2024)
Beyond Knowledge to Agency: Evaluating Expertise, Autonomy, and Integrity in Finance with CNFinBench
von: Ding, Jinru, et al.
Veröffentlicht: (2025)
von: Ding, Jinru, et al.
Veröffentlicht: (2025)
Seeing Beyond Words: MatVQA for Challenging Visual-Scientific Reasoning in Materials Science
von: Wu, Sifan, et al.
Veröffentlicht: (2025)
von: Wu, Sifan, et al.
Veröffentlicht: (2025)
FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets
von: Wang, Dannong, et al.
Veröffentlicht: (2025)
von: Wang, Dannong, et al.
Veröffentlicht: (2025)
ChemDFM-R: A Chemical Reasoning LLM Enhanced with Atomized Chemical Knowledge
von: Zhao, Zihan, et al.
Veröffentlicht: (2025)
von: Zhao, Zihan, et al.
Veröffentlicht: (2025)
Sci2Pol: Evaluating and Fine-tuning LLMs on Scientific-to-Policy Brief Generation
von: Wu, Weimin, et al.
Veröffentlicht: (2025)
von: Wu, Weimin, et al.
Veröffentlicht: (2025)
Design Editing for Offline Model-based Optimization
von: Yuan, Ye, et al.
Veröffentlicht: (2024)
von: Yuan, Ye, et al.
Veröffentlicht: (2024)
Location-Based Service (LBS) Data Quality Metrics and Effects on Mobility Inference
von: Wu, Xinhua, et al.
Veröffentlicht: (2024)
von: Wu, Xinhua, et al.
Veröffentlicht: (2024)
FinTagging: Benchmarking LLMs for Extracting and Structuring Financial Information
von: Wang, Yan, et al.
Veröffentlicht: (2025)
von: Wang, Yan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMs
von: Lu, Guilong, et al.
Veröffentlicht: (2025) -
Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks
von: Bigeard, Antoine, et al.
Veröffentlicht: (2025) -
Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs
von: Guo, Xingang, et al.
Veröffentlicht: (2025) -
Affective-ROPTester: Capability and Bias Analysis of LLMs in Predicting Retinopathy of Prematurity
von: Zhao, Shuai, et al.
Veröffentlicht: (2025) -
Flow of Knowledge: Federated Fine-Tuning of LLMs in Healthcare under Non-IID Conditions
von: Chen, Zeyu, et al.
Veröffentlicht: (2025)