Benchmarking Source-Sensitive Reasoning in Turkish: Humans and LLMs under Evidential Trust Manipulation
Fuente:
arXiv
Salvato in:
| Autori principali: | Karakaş, Sercan, Şimşek, Yusuf |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Lemmas to Dependencies: What Signals Drive Light Verbs Classification?
di: Karakaş, Sercan, et al.
Pubblicazione: (2026)
di: Karakaş, Sercan, et al.
Pubblicazione: (2026)
Plausibility as Commonsense Reasoning: Humans Succeed, Large Language Models Do not
di: Karakaş, Sercan
Pubblicazione: (2026)
di: Karakaş, Sercan
Pubblicazione: (2026)
Clause-Internal or Clause-External? Testing Turkish Reflexive Binding in Adapted versus Chain of Thought Large Language Models
di: Karakaş, Sercan
Pubblicazione: (2026)
di: Karakaş, Sercan
Pubblicazione: (2026)
Transfer Learning for an Endangered Slavic Variety: Dependency Parsing in Pomak Across Contact-Shaped Dialects
di: Karakaş, Sercan
Pubblicazione: (2026)
di: Karakaş, Sercan
Pubblicazione: (2026)
HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning
di: Li, Chance Jiajie, et al.
Pubblicazione: (2025)
di: Li, Chance Jiajie, et al.
Pubblicazione: (2025)
TurkBench: A Benchmark for Evaluating Turkish Large Language Models
di: Toraman, Çağrı, et al.
Pubblicazione: (2026)
di: Toraman, Çağrı, et al.
Pubblicazione: (2026)
OCRTurk: A Comprehensive OCR Benchmark for Turkish
di: Yılmaz, Deniz, et al.
Pubblicazione: (2026)
di: Yılmaz, Deniz, et al.
Pubblicazione: (2026)
TurkColBERT: A Benchmark of Dense and Late-Interaction Models for Turkish Information Retrieval
di: Ezerceli, Özay, et al.
Pubblicazione: (2025)
di: Ezerceli, Özay, et al.
Pubblicazione: (2025)
Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data
di: Khatri, Mann, et al.
Pubblicazione: (2025)
di: Khatri, Mann, et al.
Pubblicazione: (2025)
Are You Human? An Adversarial Benchmark to Expose LLMs
di: Gressel, Gilad, et al.
Pubblicazione: (2024)
di: Gressel, Gilad, et al.
Pubblicazione: (2024)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
di: Maniparambil, Mayug, et al.
Pubblicazione: (2026)
di: Maniparambil, Mayug, et al.
Pubblicazione: (2026)
Organic Data-Driven Approach for Turkish Grammatical Error Correction and LLMs
di: Ersoy, Asım, et al.
Pubblicazione: (2024)
di: Ersoy, Asım, et al.
Pubblicazione: (2024)
A Large-Scale Dataset and Citation Intent Classification in Turkish with LLMs
di: Karaca, Kemal Sami, et al.
Pubblicazione: (2025)
di: Karaca, Kemal Sami, et al.
Pubblicazione: (2025)
RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Reasoning Intervention in Large Reasoning Models
di: Han, Yunseok, et al.
Pubblicazione: (2026)
di: Han, Yunseok, et al.
Pubblicazione: (2026)
Multilingual Mathematical Reasoning: Advancing Open-Source LLMs in Hindi and English
di: Anand, Avinash, et al.
Pubblicazione: (2024)
di: Anand, Avinash, et al.
Pubblicazione: (2024)
Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions
di: Hong, Zijin, et al.
Pubblicazione: (2025)
di: Hong, Zijin, et al.
Pubblicazione: (2025)
From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark
di: Lei, Chao, et al.
Pubblicazione: (2025)
di: Lei, Chao, et al.
Pubblicazione: (2025)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
di: Jiang, Zhuohang, et al.
Pubblicazione: (2025)
di: Jiang, Zhuohang, et al.
Pubblicazione: (2025)
Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
di: Jin, Bowen, et al.
Pubblicazione: (2025)
di: Jin, Bowen, et al.
Pubblicazione: (2025)
Unmasking Reasoning Processes: A Process-aware Benchmark for Evaluating Structural Mathematical Reasoning in LLMs
di: Zheng, Xiang, et al.
Pubblicazione: (2026)
di: Zheng, Xiang, et al.
Pubblicazione: (2026)
Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
di: Liu, Xiao, et al.
Pubblicazione: (2024)
di: Liu, Xiao, et al.
Pubblicazione: (2024)
LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
di: Kolasani, Sai, et al.
Pubblicazione: (2025)
di: Kolasani, Sai, et al.
Pubblicazione: (2025)
AraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular Data
di: Alshaikh, Rana, et al.
Pubblicazione: (2025)
di: Alshaikh, Rana, et al.
Pubblicazione: (2025)
Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
di: Er, Yakup Abrek, et al.
Pubblicazione: (2025)
di: Er, Yakup Abrek, et al.
Pubblicazione: (2025)
Evaluating the Quality of Benchmark Datasets for Low-Resource Languages: A Case Study on Turkish
di: Cengiz, Ayşe Aysu, et al.
Pubblicazione: (2025)
di: Cengiz, Ayşe Aysu, et al.
Pubblicazione: (2025)
MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs
di: Fabbri, Alexander R., et al.
Pubblicazione: (2025)
di: Fabbri, Alexander R., et al.
Pubblicazione: (2025)
Nondeterministic Polynomial-time Problem Challenge: An Ever-Scaling Reasoning Benchmark for LLMs
di: Yang, Chang, et al.
Pubblicazione: (2025)
di: Yang, Chang, et al.
Pubblicazione: (2025)
Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG
di: Jin, Bowen, et al.
Pubblicazione: (2024)
di: Jin, Bowen, et al.
Pubblicazione: (2024)
ChatGPT Doesn't Trust Chargers Fans: Guardrail Sensitivity in Context
di: Li, Victoria R., et al.
Pubblicazione: (2024)
di: Li, Victoria R., et al.
Pubblicazione: (2024)
Do LLMs Exhibit Human-Like Reasoning? Evaluating Theory of Mind in LLMs for Open-Ended Responses
di: Amirizaniani, Maryam, et al.
Pubblicazione: (2024)
di: Amirizaniani, Maryam, et al.
Pubblicazione: (2024)
How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their Vulnerabilities
di: Mo, Lingbo, et al.
Pubblicazione: (2023)
di: Mo, Lingbo, et al.
Pubblicazione: (2023)
HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue
di: Iyer, Laya, et al.
Pubblicazione: (2026)
di: Iyer, Laya, et al.
Pubblicazione: (2026)
Shattering the Shortcut: A Topology-Regularized Benchmark for Multi-hop Medical Reasoning in LLMs
di: Zi, Xing, et al.
Pubblicazione: (2026)
di: Zi, Xing, et al.
Pubblicazione: (2026)
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
di: Zhou, Ruiwen, et al.
Pubblicazione: (2024)
di: Zhou, Ruiwen, et al.
Pubblicazione: (2024)
MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning
di: Li, Shuyue Stella, et al.
Pubblicazione: (2024)
di: Li, Shuyue Stella, et al.
Pubblicazione: (2024)
Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data
di: Wang, Qiongqiong, et al.
Pubblicazione: (2025)
di: Wang, Qiongqiong, et al.
Pubblicazione: (2025)
oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning
di: Xu, Ruiling, et al.
Pubblicazione: (2025)
di: Xu, Ruiling, et al.
Pubblicazione: (2025)
MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs
di: Zeng, Zhongshen, et al.
Pubblicazione: (2024)
di: Zeng, Zhongshen, et al.
Pubblicazione: (2024)
TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios
di: Wei, Shaohang, et al.
Pubblicazione: (2025)
di: Wei, Shaohang, et al.
Pubblicazione: (2025)
RAGTurk: Best Practices for Retrieval Augmented Generation in Turkish
di: Köse, Süha Kağan, et al.
Pubblicazione: (2026)
di: Köse, Süha Kağan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
From Lemmas to Dependencies: What Signals Drive Light Verbs Classification?
di: Karakaş, Sercan, et al.
Pubblicazione: (2026) -
Plausibility as Commonsense Reasoning: Humans Succeed, Large Language Models Do not
di: Karakaş, Sercan
Pubblicazione: (2026) -
Clause-Internal or Clause-External? Testing Turkish Reflexive Binding in Adapted versus Chain of Thought Large Language Models
di: Karakaş, Sercan
Pubblicazione: (2026) -
Transfer Learning for an Endangered Slavic Variety: Dependency Parsing in Pomak Across Contact-Shaped Dialects
di: Karakaş, Sercan
Pubblicazione: (2026) -
HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning
di: Li, Chance Jiajie, et al.
Pubblicazione: (2025)