Highlighting Case Studies in LLM Literature Review of Interdisciplinary System Science
Fuente:
arXiv
Saved in:
| Main Authors: | McGinness, Lachlan, Baumgartner, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Steamroller Problems: An Evaluation of LLM Reasoning Capability with Automated Theorem Prover Strategies
by: McGinness, Lachlan, et al.
Published: (2024)
by: McGinness, Lachlan, et al.
Published: (2024)
Automated Theorem Provers Help Improve Large Language Model Reasoning
by: McGinness, Lachlan, et al.
Published: (2024)
by: McGinness, Lachlan, et al.
Published: (2024)
Large Language Models Imitate Logical Reasoning, but at what Cost?
by: McGinness, Lachlan, et al.
Published: (2025)
by: McGinness, Lachlan, et al.
Published: (2025)
The AlphaPhysics Term Rewriting System for Marking Algebraic Expressions in Physics Exams
by: Baumgartner, Peter, et al.
Published: (2025)
by: Baumgartner, Peter, et al.
Published: (2025)
Large Language Models' Reasoning Stalls: An Investigation into the Capabilities of Frontier Models
by: McGinness, Lachlan, et al.
Published: (2025)
by: McGinness, Lachlan, et al.
Published: (2025)
CON-FOLD -- Explainable Machine Learning with Confidence
by: McGinness, Lachlan, et al.
Published: (2024)
by: McGinness, Lachlan, et al.
Published: (2024)
Overview of AI Grading of Physics Olympiad Exams
by: McGinness, Lachlan
Published: (2025)
by: McGinness, Lachlan
Published: (2025)
Can Large Language Models Correctly Interpret Equations with Errors?
by: McGinness, Lachlan, et al.
Published: (2025)
by: McGinness, Lachlan, et al.
Published: (2025)
The Benefits and Challenges of a Quantum Computing Concept Inventory
by: McGinness, Lachlan
Published: (2025)
by: McGinness, Lachlan
Published: (2025)
An Empirical Study of Conformal Prediction in LLM with ASP Scaffolds for Robust Reasoning
by: Kaur, Navdeep, et al.
Published: (2025)
by: Kaur, Navdeep, et al.
Published: (2025)
Sparking Scientific Creativity via LLM-Driven Interdisciplinary Inspiration
by: Kargupta, Priyanka, et al.
Published: (2026)
by: Kargupta, Priyanka, et al.
Published: (2026)
LitLLM: A Toolkit for Scientific Literature Review
by: Agarwal, Shubham, et al.
Published: (2024)
by: Agarwal, Shubham, et al.
Published: (2024)
The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
by: Sadallah, Abdelrahman, et al.
Published: (2025)
by: Sadallah, Abdelrahman, et al.
Published: (2025)
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
by: Baumgärtner, Tim, et al.
Published: (2026)
by: Baumgärtner, Tim, et al.
Published: (2026)
Modeling LLM Agent Reviewer Dynamics in Elo-Ranked Review System
by: Huang, Hsiang-Wei, et al.
Published: (2026)
by: Huang, Hsiang-Wei, et al.
Published: (2026)
Optimizing LLM-Based Multi-Agent System with Textual Feedback: A Case Study on Software Development
by: Shen, Ming, et al.
Published: (2025)
by: Shen, Ming, et al.
Published: (2025)
PeerQA: A Scientific Question Answering Dataset from Peer Reviews
by: Baumgärtner, Tim, et al.
Published: (2025)
by: Baumgärtner, Tim, et al.
Published: (2025)
Learning Evidence Highlighting for Frozen LLMs
by: Li, Shaoang, et al.
Published: (2026)
by: Li, Shaoang, et al.
Published: (2026)
Spectral Attention Steering for Prompt Highlighting
by: Li, Weixian Waylon, et al.
Published: (2026)
by: Li, Weixian Waylon, et al.
Published: (2026)
An Interdisciplinary Approach to Human-Centered Machine Translation
by: Carpuat, Marine, et al.
Published: (2025)
by: Carpuat, Marine, et al.
Published: (2025)
HoneyComb: A Flexible LLM-Based Agent System for Materials Science
by: Zhang, Huan, et al.
Published: (2024)
by: Zhang, Huan, et al.
Published: (2024)
Attention Flows: Tracing LLM Conceptual Engagement via Story Summaries
by: Hicke, Rebecca M. M., et al.
Published: (2026)
by: Hicke, Rebecca M. M., et al.
Published: (2026)
Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
by: Kim, Kon Woo, et al.
Published: (2025)
by: Kim, Kon Woo, et al.
Published: (2025)
Rethinking LLM Language Adaptation: A Case Study on Chinese Mixtral
by: Cui, Yiming, et al.
Published: (2024)
by: Cui, Yiming, et al.
Published: (2024)
Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic
by: Ingimundarson, Finnur Ágúst, et al.
Published: (2026)
by: Ingimundarson, Finnur Ágúst, et al.
Published: (2026)
DIAMOND: An LLM-Driven Agent for Context-Aware Baseball Highlight Summarization
by: Kang, Jeonghun, et al.
Published: (2025)
by: Kang, Jeonghun, et al.
Published: (2025)
Design Techniques for LLM-Powered Interactive Storytelling: A Case Study of the Dramamancer System
by: Wang, Tiffany, et al.
Published: (2026)
by: Wang, Tiffany, et al.
Published: (2026)
Case-Aware LLM-as-a-Judge Evaluation for Enterprise-Scale RAG Systems
by: Chhabra, Mukul, et al.
Published: (2026)
by: Chhabra, Mukul, et al.
Published: (2026)
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
by: Bai, Xiaoyan, et al.
Published: (2026)
by: Bai, Xiaoyan, et al.
Published: (2026)
Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction
by: Bailis, Suma, et al.
Published: (2024)
by: Bailis, Suma, et al.
Published: (2024)
Can Large Language Models Discern Evidence for Scientific Hypotheses? Case Studies in the Social Sciences
by: Koneru, Sai, et al.
Published: (2023)
by: Koneru, Sai, et al.
Published: (2023)
Matter-of-Fact: A Benchmark for Verifying the Feasibility of Literature-Supported Claims in Materials Science
by: Jansen, Peter, et al.
Published: (2025)
by: Jansen, Peter, et al.
Published: (2025)
From Natural Language to SQL: Review of LLM-based Text-to-SQL Systems
by: Mohammadjafari, Ali, et al.
Published: (2024)
by: Mohammadjafari, Ali, et al.
Published: (2024)
Generating Literature-Driven Scientific Theories at Scale
by: Jansen, Peter, et al.
Published: (2026)
by: Jansen, Peter, et al.
Published: (2026)
daVinci-LLM:Towards the Science of Pretraining
by: Qin, Yiwei, et al.
Published: (2026)
by: Qin, Yiwei, et al.
Published: (2026)
Disparities in LLM Reasoning Accuracy and Explanations: A Case Study on African American English
by: Zhou, Runtao, et al.
Published: (2025)
by: Zhou, Runtao, et al.
Published: (2025)
An Auditable Pipeline for Fuzzy Full-Text Screening in Systematic Reviews: Integrating Contrastive Semantic Highlighting and LLM Judgment
by: Mortezaagha, Pouria, et al.
Published: (2025)
by: Mortezaagha, Pouria, et al.
Published: (2025)
Text Production and Comprehension by Human and Artificial Intelligence: Interdisciplinary Workshop Report
by: Speltz, Emily Dux
Published: (2025)
by: Speltz, Emily Dux
Published: (2025)
Societal AI Research Has Become Less Interdisciplinary
by: Markus, Dror Kris, et al.
Published: (2025)
by: Markus, Dror Kris, et al.
Published: (2025)
Political-LLM: Large Language Models in Political Science
by: Li, Lincan, et al.
Published: (2024)
by: Li, Lincan, et al.
Published: (2024)
Similar Items
-
Steamroller Problems: An Evaluation of LLM Reasoning Capability with Automated Theorem Prover Strategies
by: McGinness, Lachlan, et al.
Published: (2024) -
Automated Theorem Provers Help Improve Large Language Model Reasoning
by: McGinness, Lachlan, et al.
Published: (2024) -
Large Language Models Imitate Logical Reasoning, but at what Cost?
by: McGinness, Lachlan, et al.
Published: (2025) -
The AlphaPhysics Term Rewriting System for Marking Algebraic Expressions in Physics Exams
by: Baumgartner, Peter, et al.
Published: (2025) -
Large Language Models' Reasoning Stalls: An Investigation into the Capabilities of Frontier Models
by: McGinness, Lachlan, et al.
Published: (2025)