CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tang, Xiaqiang, Li, Jian, Hu, Keyu, Nan, Du, Li, Xiaolong, Zhang, Xi, Sun, Weigao, Xie, Sihong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SSFO: Self-Supervised Faithfulness Optimization for Retrieval-Augmented Generation
von: Tang, Xiaqiang, et al.
Veröffentlicht: (2025)
von: Tang, Xiaqiang, et al.
Veröffentlicht: (2025)
Adapting to Non-Stationary Environments: Multi-Armed Bandit Enhanced Retrieval-Augmented Generation on Knowledge Graphs
von: Tang, Xiaqiang, et al.
Veröffentlicht: (2024)
von: Tang, Xiaqiang, et al.
Veröffentlicht: (2024)
MBA-RAG: a Bandit Approach for Adaptive Retrieval-Augmented Generation through Question Complexity
von: Tang, Xiaqiang, et al.
Veröffentlicht: (2024)
von: Tang, Xiaqiang, et al.
Veröffentlicht: (2024)
MS-Net: A Multi-Path Sparse Model for Motion Prediction in Multi-Scenes
von: Tang, Xiaqiang, et al.
Veröffentlicht: (2024)
von: Tang, Xiaqiang, et al.
Veröffentlicht: (2024)
PSG-Nav: Probabilistic Scene Graph Navigation via Multiverse Decision Making
von: Chen, Rufeng, et al.
Veröffentlicht: (2026)
von: Chen, Rufeng, et al.
Veröffentlicht: (2026)
Comba: Improving Bilinear RNNs with Closed-loop Control
von: Hu, Jiaxi, et al.
Veröffentlicht: (2025)
von: Hu, Jiaxi, et al.
Veröffentlicht: (2025)
CogniDual Framework: Self-Training Large Language Models within a Dual-System Theoretical Framework for Improving Cognitive Tasks
von: Deng, Yongxin, et al.
Veröffentlicht: (2024)
von: Deng, Yongxin, et al.
Veröffentlicht: (2024)
Liger: Linearizing Large Language Models to Gated Recurrent Structures
von: Lan, Disen, et al.
Veröffentlicht: (2025)
von: Lan, Disen, et al.
Veröffentlicht: (2025)
FaithLM: Towards Faithful Explanations for Large Language Models
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
CogniDot: Vasoactivity-based Cognitive Load Monitoring with a Miniature On-skin Sensor
von: Lan, Hongbo, et al.
Veröffentlicht: (2024)
von: Lan, Hongbo, et al.
Veröffentlicht: (2024)
VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models
von: Dong, Nguyen Tien, et al.
Veröffentlicht: (2025)
von: Dong, Nguyen Tien, et al.
Veröffentlicht: (2025)
CogniTextes
Veröffentlicht: (2010)
Veröffentlicht: (2010)
CogniFold: Always-On Proactive Memory via Cognitive Folding
von: Wang, Suli, et al.
Veröffentlicht: (2026)
von: Wang, Suli, et al.
Veröffentlicht: (2026)
LegalFaith: A Faithfulness Benchmark for Legal QA
von: Singh, Rajshree
Veröffentlicht: (2026)
von: Singh, Rajshree
Veröffentlicht: (2026)
SPD-Faith Bench: Diagnosing and Improving Faithfulness in Chain-of-Thought for Multimodal Large Language Models
von: Lv, Weijiang, et al.
Veröffentlicht: (2026)
von: Lv, Weijiang, et al.
Veröffentlicht: (2026)
Context-Fidelity Boosting: Enhancing Faithful Generation through Watermark-Inspired Decoding
von: Zhang, Weixu, et al.
Veröffentlicht: (2026)
von: Zhang, Weixu, et al.
Veröffentlicht: (2026)
Attribution-Guided Continual Learning for Large Language Models
von: Liu, Yazheng, et al.
Veröffentlicht: (2026)
von: Liu, Yazheng, et al.
Veröffentlicht: (2026)
Causality‐inspired crop pest recognition based on Decoupled Feature Learning
von: Tao Hu, et al.
Veröffentlicht: (2024)
von: Tao Hu, et al.
Veröffentlicht: (2024)
NewsBench: A Systematic Evaluation Framework for Assessing Editorial Capabilities of Large Language Models in Chinese Journalism
von: Li, Miao, et al.
Veröffentlicht: (2024)
von: Li, Miao, et al.
Veröffentlicht: (2024)
CogniMap3D: Cognitive 3D Mapping and Rapid Retrieval
von: Wang, Feiran, et al.
Veröffentlicht: (2026)
von: Wang, Feiran, et al.
Veröffentlicht: (2026)
A Differential Geometric View and Explainability of GNN on Evolving Graphs
von: Liu, Yazheng, et al.
Veröffentlicht: (2024)
von: Liu, Yazheng, et al.
Veröffentlicht: (2024)
FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models
von: Jing, Liqiang, et al.
Veröffentlicht: (2023)
von: Jing, Liqiang, et al.
Veröffentlicht: (2023)
Mitigating Large Language Model Hallucination with Faithful Finetuning
von: Hu, Minda, et al.
Veröffentlicht: (2024)
von: Hu, Minda, et al.
Veröffentlicht: (2024)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
von: Ovcharov, Volodymyr
Veröffentlicht: (2026)
von: Ovcharov, Volodymyr
Veröffentlicht: (2026)
S-GRPO: Unified Post-Training for Large Vision-Language Models
von: Yan, Yuming, et al.
Veröffentlicht: (2026)
von: Yan, Yuming, et al.
Veröffentlicht: (2026)
SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
von: Xu, Peiran, et al.
Veröffentlicht: (2025)
von: Xu, Peiran, et al.
Veröffentlicht: (2025)
Proximity QA: Unleashing the Power of Multi-Modal Large Language Models for Spatial Proximity Analysis
von: Li, Jianing, et al.
Veröffentlicht: (2024)
von: Li, Jianing, et al.
Veröffentlicht: (2024)
CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
Are Large Language Models Good Statisticians?
von: Zhu, Yizhang, et al.
Veröffentlicht: (2024)
von: Zhu, Yizhang, et al.
Veröffentlicht: (2024)
LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification
von: Neto, Pedro Barbosa de Carvalho
Veröffentlicht: (2026)
von: Neto, Pedro Barbosa de Carvalho
Veröffentlicht: (2026)
LegalAgentBench: Evaluating LLM Agents in Legal Domain
von: Li, Haitao, et al.
Veröffentlicht: (2024)
von: Li, Haitao, et al.
Veröffentlicht: (2024)
KoALa-Bench: Evaluating Large Audio Language Models on Korean Speech Understanding and Faithfulness
von: Kim, Jinyoung, et al.
Veröffentlicht: (2026)
von: Kim, Jinyoung, et al.
Veröffentlicht: (2026)
LegalCiteBench: Evaluating Citation Reliability in Legal Language Models
von: Chen, Sijia, et al.
Veröffentlicht: (2026)
von: Chen, Sijia, et al.
Veröffentlicht: (2026)
CogniVoice: Multimodal and Multilingual Fusion Networks for Mild Cognitive Impairment Assessment from Spontaneous Speech
von: Cheng, Jiali, et al.
Veröffentlicht: (2024)
von: Cheng, Jiali, et al.
Veröffentlicht: (2024)
Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models
von: Zhang, Bang, et al.
Veröffentlicht: (2025)
von: Zhang, Bang, et al.
Veröffentlicht: (2025)
Are Large Language Models Chameleons? An Attempt to Simulate Social Surveys
von: Geng, Mingmeng, et al.
Veröffentlicht: (2024)
von: Geng, Mingmeng, et al.
Veröffentlicht: (2024)
OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice
von: Wang, Rongyang, et al.
Veröffentlicht: (2026)
von: Wang, Rongyang, et al.
Veröffentlicht: (2026)
A Short Review for Ontology Learning: Stride to Large Language Models Trend
von: Du, Rick, et al.
Veröffentlicht: (2024)
von: Du, Rick, et al.
Veröffentlicht: (2024)
PrivaCI-Bench: Evaluating Privacy with Contextual Integrity and Legal Compliance
von: Li, Haoran, et al.
Veröffentlicht: (2025)
von: Li, Haoran, et al.
Veröffentlicht: (2025)
VinaBench: Benchmark for Faithful and Consistent Visual Narratives
von: Gao, Silin, et al.
Veröffentlicht: (2025)
von: Gao, Silin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SSFO: Self-Supervised Faithfulness Optimization for Retrieval-Augmented Generation
von: Tang, Xiaqiang, et al.
Veröffentlicht: (2025) -
Adapting to Non-Stationary Environments: Multi-Armed Bandit Enhanced Retrieval-Augmented Generation on Knowledge Graphs
von: Tang, Xiaqiang, et al.
Veröffentlicht: (2024) -
MBA-RAG: a Bandit Approach for Adaptive Retrieval-Augmented Generation through Question Complexity
von: Tang, Xiaqiang, et al.
Veröffentlicht: (2024) -
MS-Net: A Multi-Path Sparse Model for Motion Prediction in Multi-Scenes
von: Tang, Xiaqiang, et al.
Veröffentlicht: (2024) -
PSG-Nav: Probabilistic Scene Graph Navigation via Multiverse Decision Making
von: Chen, Rufeng, et al.
Veröffentlicht: (2026)