Autonomous Evaluation of LLMs for Truth Maintenance and Reasoning Tasks
Fuente:
arXiv
Guardado en:
| Autores principales: | Karia, Rushang, Bramblett, Daniel, Dobhal, Daksh, Srivastava, Siddharth |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
$\forall$uto$\exists$val: Autonomous Assessment of LLMs in Formal Synthesis and Interpretation Tasks
por: Karia, Rushang, et al.
Publicado: (2024)
por: Karia, Rushang, et al.
Publicado: (2024)
Discovering and Learning Probabilistic Models of Black-Box AI Capabilities
por: Bramblett, Daniel, et al.
Publicado: (2025)
por: Bramblett, Daniel, et al.
Publicado: (2025)
Epistemic Exploration for Generalizable Planning and Learning in Non-Stationary Settings
por: Karia, Rushang, et al.
Publicado: (2024)
por: Karia, Rushang, et al.
Publicado: (2024)
Belief-State Query Policies for User-Aligned POMDPs
por: Bramblett, Daniel, et al.
Publicado: (2024)
por: Bramblett, Daniel, et al.
Publicado: (2024)
Using Explainable AI and Hierarchical Planning for Outreach with Robots
por: Karia, Rushang, et al.
Publicado: (2024)
por: Karia, Rushang, et al.
Publicado: (2024)
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
por: Barkett, Emilio, et al.
Publicado: (2025)
por: Barkett, Emilio, et al.
Publicado: (2025)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
por: Khatun, Aisha, et al.
Publicado: (2024)
por: Khatun, Aisha, et al.
Publicado: (2024)
Teaching and Evaluating LLMs to Reason About Polymer Design Related Tasks
por: Mohanty, Dikshya, et al.
Publicado: (2026)
por: Mohanty, Dikshya, et al.
Publicado: (2026)
Testing the Limits of Truth Directions in LLMs
por: Poulis, Angelos, et al.
Publicado: (2026)
por: Poulis, Angelos, et al.
Publicado: (2026)
How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs
por: Adarsh, Shivam, et al.
Publicado: (2026)
por: Adarsh, Shivam, et al.
Publicado: (2026)
Truth is Universal: Robust Detection of Lies in LLMs
por: Bürger, Lennart, et al.
Publicado: (2024)
por: Bürger, Lennart, et al.
Publicado: (2024)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
por: Kabir, Mohsinul, et al.
Publicado: (2026)
por: Kabir, Mohsinul, et al.
Publicado: (2026)
Collective Reasoning Among LLMs: A Framework for Answer Validation Without Ground Truth
por: Davoudi, Seyed Pouyan Mousavi, et al.
Publicado: (2025)
por: Davoudi, Seyed Pouyan Mousavi, et al.
Publicado: (2025)
Truth Knows No Language: Evaluating Truthfulness Beyond English
por: Figueras, Blanca Calvo, et al.
Publicado: (2025)
por: Figueras, Blanca Calvo, et al.
Publicado: (2025)
Truth, Trust, and Trouble: Medical AI on the Edge
por: Azeez, Mohammad Anas, et al.
Publicado: (2025)
por: Azeez, Mohammad Anas, et al.
Publicado: (2025)
Are LLMs good pragmatic speakers?
por: Jian, Mingyue, et al.
Publicado: (2024)
por: Jian, Mingyue, et al.
Publicado: (2024)
Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions
por: Murugadoss, Bhuvanashree, et al.
Publicado: (2024)
por: Murugadoss, Bhuvanashree, et al.
Publicado: (2024)
TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
por: Wei, Zhepei, et al.
Publicado: (2025)
por: Wei, Zhepei, et al.
Publicado: (2025)
Ground-Truth Subgraphs for Better Training and Evaluation of Knowledge Graph Augmented LLMs
por: Cattaneo, Alberto, et al.
Publicado: (2025)
por: Cattaneo, Alberto, et al.
Publicado: (2025)
MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine
por: Kong, Shufeng, et al.
Publicado: (2025)
por: Kong, Shufeng, et al.
Publicado: (2025)
Debating with More Persuasive LLMs Leads to More Truthful Answers
por: Khan, Akbir, et al.
Publicado: (2024)
por: Khan, Akbir, et al.
Publicado: (2024)
Fact or Fiction? Can LLMs be Reliable Annotators for Political Truths?
por: Chatrath, Veronica, et al.
Publicado: (2024)
por: Chatrath, Veronica, et al.
Publicado: (2024)
Graphing the Truth: Structured Visualizations for Automated Hallucination Detection in LLMs
por: Agrawal, Tanmay
Publicado: (2025)
por: Agrawal, Tanmay
Publicado: (2025)
Ground Truth Generation for Multilingual Historical NLP using LLMs
por: Gladstone, Clovis, et al.
Publicado: (2025)
por: Gladstone, Clovis, et al.
Publicado: (2025)
Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization
por: Jin, Keyan, et al.
Publicado: (2025)
por: Jin, Keyan, et al.
Publicado: (2025)
The Geometries of Truth Are Orthogonal Across Tasks
por: Azizian, Waiss, et al.
Publicado: (2025)
por: Azizian, Waiss, et al.
Publicado: (2025)
HALT-RAG: A Task-Adaptable Framework for Hallucination Detection with Calibrated NLI Ensembles and Abstention
por: Goswami, Saumya, et al.
Publicado: (2025)
por: Goswami, Saumya, et al.
Publicado: (2025)
Phonetic Perturbations Reveal Tokenizer-Rooted Safety Gaps in LLMs
por: Aswal, Darpan, et al.
Publicado: (2025)
por: Aswal, Darpan, et al.
Publicado: (2025)
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
por: Pires, Ramon, et al.
Publicado: (2026)
por: Pires, Ramon, et al.
Publicado: (2026)
SELT: Self-Evaluation Tree Search for LLMs with Task Decomposition
por: Wu, Mengsong, et al.
Publicado: (2025)
por: Wu, Mengsong, et al.
Publicado: (2025)
Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap
por: Srivastava, Saurabh, et al.
Publicado: (2024)
por: Srivastava, Saurabh, et al.
Publicado: (2024)
AggTruth: Contextual Hallucination Detection using Aggregated Attention Scores in LLMs
por: Matys, Piotr, et al.
Publicado: (2025)
por: Matys, Piotr, et al.
Publicado: (2025)
Evaluating LLMs for Hardware Design and Test
por: Blocklove, Jason, et al.
Publicado: (2024)
por: Blocklove, Jason, et al.
Publicado: (2024)
Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench
por: Narad, Reuben, et al.
Publicado: (2025)
por: Narad, Reuben, et al.
Publicado: (2025)
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
por: Zhang, Kongcheng, et al.
Publicado: (2025)
por: Zhang, Kongcheng, et al.
Publicado: (2025)
MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
por: Ghosh, Rajarshi, et al.
Publicado: (2025)
por: Ghosh, Rajarshi, et al.
Publicado: (2025)
What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
por: Do, Heejin, et al.
Publicado: (2025)
por: Do, Heejin, et al.
Publicado: (2025)
TruthStance: An Annotated Dataset of Conversations on Truth Social
por: Ameen, Fathima, et al.
Publicado: (2026)
por: Ameen, Fathima, et al.
Publicado: (2026)
LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
por: Choukrani, Omar, et al.
Publicado: (2025)
por: Choukrani, Omar, et al.
Publicado: (2025)
REAMS: Reasoning Enhanced Algorithm for Maths Solving
por: Singh, Eishkaran, et al.
Publicado: (2025)
por: Singh, Eishkaran, et al.
Publicado: (2025)
Ejemplares similares
-
$\forall$uto$\exists$val: Autonomous Assessment of LLMs in Formal Synthesis and Interpretation Tasks
por: Karia, Rushang, et al.
Publicado: (2024) -
Discovering and Learning Probabilistic Models of Black-Box AI Capabilities
por: Bramblett, Daniel, et al.
Publicado: (2025) -
Epistemic Exploration for Generalizable Planning and Learning in Non-Stationary Settings
por: Karia, Rushang, et al.
Publicado: (2024) -
Belief-State Query Policies for User-Aligned POMDPs
por: Bramblett, Daniel, et al.
Publicado: (2024) -
Using Explainable AI and Hierarchical Planning for Outreach with Robots
por: Karia, Rushang, et al.
Publicado: (2024)