Benchmarking Source-Sensitive Reasoning in Turkish: Humans and LLMs under Evidential Trust Manipulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Karakaş, Sercan, Şimşek, Yusuf |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Lemmas to Dependencies: What Signals Drive Light Verbs Classification?
von: Karakaş, Sercan, et al.
Veröffentlicht: (2026)
von: Karakaş, Sercan, et al.
Veröffentlicht: (2026)
Plausibility as Commonsense Reasoning: Humans Succeed, Large Language Models Do not
von: Karakaş, Sercan
Veröffentlicht: (2026)
von: Karakaş, Sercan
Veröffentlicht: (2026)
Clause-Internal or Clause-External? Testing Turkish Reflexive Binding in Adapted versus Chain of Thought Large Language Models
von: Karakaş, Sercan
Veröffentlicht: (2026)
von: Karakaş, Sercan
Veröffentlicht: (2026)
Transfer Learning for an Endangered Slavic Variety: Dependency Parsing in Pomak Across Contact-Shaped Dialects
von: Karakaş, Sercan
Veröffentlicht: (2026)
von: Karakaş, Sercan
Veröffentlicht: (2026)
HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning
von: Li, Chance Jiajie, et al.
Veröffentlicht: (2025)
von: Li, Chance Jiajie, et al.
Veröffentlicht: (2025)
TurkBench: A Benchmark for Evaluating Turkish Large Language Models
von: Toraman, Çağrı, et al.
Veröffentlicht: (2026)
von: Toraman, Çağrı, et al.
Veröffentlicht: (2026)
OCRTurk: A Comprehensive OCR Benchmark for Turkish
von: Yılmaz, Deniz, et al.
Veröffentlicht: (2026)
von: Yılmaz, Deniz, et al.
Veröffentlicht: (2026)
TurkColBERT: A Benchmark of Dense and Late-Interaction Models for Turkish Information Retrieval
von: Ezerceli, Özay, et al.
Veröffentlicht: (2025)
von: Ezerceli, Özay, et al.
Veröffentlicht: (2025)
Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data
von: Khatri, Mann, et al.
Veröffentlicht: (2025)
von: Khatri, Mann, et al.
Veröffentlicht: (2025)
Are You Human? An Adversarial Benchmark to Expose LLMs
von: Gressel, Gilad, et al.
Veröffentlicht: (2024)
von: Gressel, Gilad, et al.
Veröffentlicht: (2024)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
von: Maniparambil, Mayug, et al.
Veröffentlicht: (2026)
von: Maniparambil, Mayug, et al.
Veröffentlicht: (2026)
Organic Data-Driven Approach for Turkish Grammatical Error Correction and LLMs
von: Ersoy, Asım, et al.
Veröffentlicht: (2024)
von: Ersoy, Asım, et al.
Veröffentlicht: (2024)
A Large-Scale Dataset and Citation Intent Classification in Turkish with LLMs
von: Karaca, Kemal Sami, et al.
Veröffentlicht: (2025)
von: Karaca, Kemal Sami, et al.
Veröffentlicht: (2025)
RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Reasoning Intervention in Large Reasoning Models
von: Han, Yunseok, et al.
Veröffentlicht: (2026)
von: Han, Yunseok, et al.
Veröffentlicht: (2026)
Multilingual Mathematical Reasoning: Advancing Open-Source LLMs in Hindi and English
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions
von: Hong, Zijin, et al.
Veröffentlicht: (2025)
von: Hong, Zijin, et al.
Veröffentlicht: (2025)
From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark
von: Lei, Chao, et al.
Veröffentlicht: (2025)
von: Lei, Chao, et al.
Veröffentlicht: (2025)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
von: Jin, Bowen, et al.
Veröffentlicht: (2025)
von: Jin, Bowen, et al.
Veröffentlicht: (2025)
Unmasking Reasoning Processes: A Process-aware Benchmark for Evaluating Structural Mathematical Reasoning in LLMs
von: Zheng, Xiang, et al.
Veröffentlicht: (2026)
von: Zheng, Xiang, et al.
Veröffentlicht: (2026)
Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
von: Kolasani, Sai, et al.
Veröffentlicht: (2025)
von: Kolasani, Sai, et al.
Veröffentlicht: (2025)
AraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular Data
von: Alshaikh, Rana, et al.
Veröffentlicht: (2025)
von: Alshaikh, Rana, et al.
Veröffentlicht: (2025)
Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
von: Er, Yakup Abrek, et al.
Veröffentlicht: (2025)
von: Er, Yakup Abrek, et al.
Veröffentlicht: (2025)
Evaluating the Quality of Benchmark Datasets for Low-Resource Languages: A Case Study on Turkish
von: Cengiz, Ayşe Aysu, et al.
Veröffentlicht: (2025)
von: Cengiz, Ayşe Aysu, et al.
Veröffentlicht: (2025)
MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs
von: Fabbri, Alexander R., et al.
Veröffentlicht: (2025)
von: Fabbri, Alexander R., et al.
Veröffentlicht: (2025)
Nondeterministic Polynomial-time Problem Challenge: An Ever-Scaling Reasoning Benchmark for LLMs
von: Yang, Chang, et al.
Veröffentlicht: (2025)
von: Yang, Chang, et al.
Veröffentlicht: (2025)
Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG
von: Jin, Bowen, et al.
Veröffentlicht: (2024)
von: Jin, Bowen, et al.
Veröffentlicht: (2024)
ChatGPT Doesn't Trust Chargers Fans: Guardrail Sensitivity in Context
von: Li, Victoria R., et al.
Veröffentlicht: (2024)
von: Li, Victoria R., et al.
Veröffentlicht: (2024)
Do LLMs Exhibit Human-Like Reasoning? Evaluating Theory of Mind in LLMs for Open-Ended Responses
von: Amirizaniani, Maryam, et al.
Veröffentlicht: (2024)
von: Amirizaniani, Maryam, et al.
Veröffentlicht: (2024)
How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their Vulnerabilities
von: Mo, Lingbo, et al.
Veröffentlicht: (2023)
von: Mo, Lingbo, et al.
Veröffentlicht: (2023)
HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue
von: Iyer, Laya, et al.
Veröffentlicht: (2026)
von: Iyer, Laya, et al.
Veröffentlicht: (2026)
Shattering the Shortcut: A Topology-Regularized Benchmark for Multi-hop Medical Reasoning in LLMs
von: Zi, Xing, et al.
Veröffentlicht: (2026)
von: Zi, Xing, et al.
Veröffentlicht: (2026)
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
von: Zhou, Ruiwen, et al.
Veröffentlicht: (2024)
von: Zhou, Ruiwen, et al.
Veröffentlicht: (2024)
MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning
von: Li, Shuyue Stella, et al.
Veröffentlicht: (2024)
von: Li, Shuyue Stella, et al.
Veröffentlicht: (2024)
Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data
von: Wang, Qiongqiong, et al.
Veröffentlicht: (2025)
von: Wang, Qiongqiong, et al.
Veröffentlicht: (2025)
oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning
von: Xu, Ruiling, et al.
Veröffentlicht: (2025)
von: Xu, Ruiling, et al.
Veröffentlicht: (2025)
MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs
von: Zeng, Zhongshen, et al.
Veröffentlicht: (2024)
von: Zeng, Zhongshen, et al.
Veröffentlicht: (2024)
TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios
von: Wei, Shaohang, et al.
Veröffentlicht: (2025)
von: Wei, Shaohang, et al.
Veröffentlicht: (2025)
RAGTurk: Best Practices for Retrieval Augmented Generation in Turkish
von: Köse, Süha Kağan, et al.
Veröffentlicht: (2026)
von: Köse, Süha Kağan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
From Lemmas to Dependencies: What Signals Drive Light Verbs Classification?
von: Karakaş, Sercan, et al.
Veröffentlicht: (2026) -
Plausibility as Commonsense Reasoning: Humans Succeed, Large Language Models Do not
von: Karakaş, Sercan
Veröffentlicht: (2026) -
Clause-Internal or Clause-External? Testing Turkish Reflexive Binding in Adapted versus Chain of Thought Large Language Models
von: Karakaş, Sercan
Veröffentlicht: (2026) -
Transfer Learning for an Endangered Slavic Variety: Dependency Parsing in Pomak Across Contact-Shaped Dialects
von: Karakaş, Sercan
Veröffentlicht: (2026) -
HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning
von: Li, Chance Jiajie, et al.
Veröffentlicht: (2025)