FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain
Fuente:
arXiv
Salvato in:
| Autori principali: | Hu, Tiansheng, Hu, Tongyan, Bai, Liuyang, Zhao, Yilun, Cohan, Arman, Zhao, Chen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
di: Long, Yitao, et al.
Pubblicazione: (2025)
di: Long, Yitao, et al.
Pubblicazione: (2025)
SAGE: Benchmarking and Improving Retrieval for Deep Research Agents
di: Hu, Tiansheng, et al.
Pubblicazione: (2026)
di: Hu, Tiansheng, et al.
Pubblicazione: (2026)
FinDVer: Explainable Claim Verification over Long and Hybrid-Content Financial Documents
di: Zhao, Yilun, et al.
Pubblicazione: (2024)
di: Zhao, Yilun, et al.
Pubblicazione: (2024)
FinanceMath: Knowledge-Intensive Math Reasoning in Finance Domains
di: Zhao, Yilun, et al.
Pubblicazione: (2023)
di: Zhao, Yilun, et al.
Pubblicazione: (2023)
SciRAG: Adaptive, Citation-Aware, and Outline-Guided Retrieval and Synthesis for Scientific Literature
di: Ding, Hang, et al.
Pubblicazione: (2025)
di: Ding, Hang, et al.
Pubblicazione: (2025)
Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization
di: Liu, Yixin, et al.
Pubblicazione: (2023)
di: Liu, Yixin, et al.
Pubblicazione: (2023)
MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Search
di: Hu, Yunhai, et al.
Pubblicazione: (2025)
di: Hu, Yunhai, et al.
Pubblicazione: (2025)
Fin-Bias: Comprehensive Evaluation for LLM Decision-Making under human bias in Finance Domain
di: Hu, Xiaoyu, et al.
Pubblicazione: (2026)
di: Hu, Xiaoyu, et al.
Pubblicazione: (2026)
Observable Propagation: Uncovering Feature Vectors in Transformers
di: Dunefsky, Jacob, et al.
Pubblicazione: (2023)
di: Dunefsky, Jacob, et al.
Pubblicazione: (2023)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
di: Liu, Yixin, et al.
Pubblicazione: (2025)
di: Liu, Yixin, et al.
Pubblicazione: (2025)
MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models
di: Zhang, Yichi, et al.
Pubblicazione: (2024)
di: Zhang, Yichi, et al.
Pubblicazione: (2024)
Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models
di: Riddell, Martin, et al.
Pubblicazione: (2024)
di: Riddell, Martin, et al.
Pubblicazione: (2024)
Survey on Evaluation of LLM-based Agents
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
ReIFE: Re-evaluating Instruction-Following Evaluation
di: Liu, Yixin, et al.
Pubblicazione: (2024)
di: Liu, Yixin, et al.
Pubblicazione: (2024)
Understanding Reference Policies in Direct Preference Optimization
di: Liu, Yixin, et al.
Pubblicazione: (2024)
di: Liu, Yixin, et al.
Pubblicazione: (2024)
Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the Future
di: Wu, Sihong, et al.
Pubblicazione: (2026)
di: Wu, Sihong, et al.
Pubblicazione: (2026)
RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation
di: Wu, Sihong, et al.
Pubblicazione: (2026)
di: Wu, Sihong, et al.
Pubblicazione: (2026)
TrustLDM: Benchmarking Trustworthiness in Language Diffusion Models
di: Mo, Yichuan, et al.
Pubblicazione: (2026)
di: Mo, Yichuan, et al.
Pubblicazione: (2026)
Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack
di: Huang, Tiansheng, et al.
Pubblicazione: (2024)
di: Huang, Tiansheng, et al.
Pubblicazione: (2024)
SUCEA: Reasoning-Intensive Retrieval for Adversarial Fact-checking through Claim Decomposition and Editing
di: Liu, Hongjun, et al.
Pubblicazione: (2025)
di: Liu, Hongjun, et al.
Pubblicazione: (2025)
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
di: Song, Tingyu, et al.
Pubblicazione: (2025)
di: Song, Tingyu, et al.
Pubblicazione: (2025)
References Improve LLM Alignment in Non-Verifiable Domains
di: Shi, Kejian, et al.
Pubblicazione: (2026)
di: Shi, Kejian, et al.
Pubblicazione: (2026)
Table-R1: Inference-Time Scaling for Table Reasoning
di: Yang, Zheyuan, et al.
Pubblicazione: (2025)
di: Yang, Zheyuan, et al.
Pubblicazione: (2025)
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
di: Gao, Mingqi, et al.
Pubblicazione: (2024)
di: Gao, Mingqi, et al.
Pubblicazione: (2024)
Semantic Codebooks as Effective Priors for Neural Speech Compression
di: Bai, Liuyang, et al.
Pubblicazione: (2025)
di: Bai, Liuyang, et al.
Pubblicazione: (2025)
Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers
di: Xu, Zhijian, et al.
Pubblicazione: (2025)
di: Xu, Zhijian, et al.
Pubblicazione: (2025)
Investigating Data Contamination in Modern Benchmarks for Large Language Models
di: Deng, Chunyuan, et al.
Pubblicazione: (2023)
di: Deng, Chunyuan, et al.
Pubblicazione: (2023)
SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification
di: Wang, Chengye, et al.
Pubblicazione: (2025)
di: Wang, Chengye, et al.
Pubblicazione: (2025)
MDCure: A Scalable Pipeline for Multi-Document Instruction-Following
di: Liu, Gabrielle Kaili-May, et al.
Pubblicazione: (2024)
di: Liu, Gabrielle Kaili-May, et al.
Pubblicazione: (2024)
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
di: Li, Chuhan, et al.
Pubblicazione: (2024)
di: Li, Chuhan, et al.
Pubblicazione: (2024)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
di: Yu, Zhaojian, et al.
Pubblicazione: (2024)
di: Yu, Zhaojian, et al.
Pubblicazione: (2024)
CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models
di: Xia, Peng, et al.
Pubblicazione: (2024)
di: Xia, Peng, et al.
Pubblicazione: (2024)
Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
di: Zhao, Yilun, et al.
Pubblicazione: (2026)
di: Zhao, Yilun, et al.
Pubblicazione: (2026)
FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models
di: Mateega, Spencer, et al.
Pubblicazione: (2025)
di: Mateega, Spencer, et al.
Pubblicazione: (2025)
Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models
di: Zheng, Kening, et al.
Pubblicazione: (2024)
di: Zheng, Kening, et al.
Pubblicazione: (2024)
Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers
di: Zhao, Yilun, et al.
Pubblicazione: (2025)
di: Zhao, Yilun, et al.
Pubblicazione: (2025)
LimRank: Less is More for Reasoning-Intensive Information Reranking
di: Song, Tingyu, et al.
Pubblicazione: (2025)
di: Song, Tingyu, et al.
Pubblicazione: (2025)
Calibrating Long-form Generations from Large Language Models
di: Huang, Yukun, et al.
Pubblicazione: (2024)
di: Huang, Yukun, et al.
Pubblicazione: (2024)
LLM-TOPLA: Efficient LLM Ensemble by Maximising Diversity
di: Tekin, Selim Furkan, et al.
Pubblicazione: (2024)
di: Tekin, Selim Furkan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
di: Long, Yitao, et al.
Pubblicazione: (2025) -
SAGE: Benchmarking and Improving Retrieval for Deep Research Agents
di: Hu, Tiansheng, et al.
Pubblicazione: (2026) -
FinDVer: Explainable Claim Verification over Long and Hybrid-Content Financial Documents
di: Zhao, Yilun, et al.
Pubblicazione: (2024) -
FinanceMath: Knowledge-Intensive Math Reasoning in Finance Domains
di: Zhao, Yilun, et al.
Pubblicazione: (2023) -
SciRAG: Adaptive, Citation-Aware, and Outline-Guided Retrieval and Synthesis for Scientific Literature
di: Ding, Hang, et al.
Pubblicazione: (2025)