Exploring the Limits of Fine-grained LLM-based Physics Inference via Premise Removal Interventions
Fuente:
arXiv
Saved in:
| Main Authors: | Meadows, Jordan, James, Tamsin, Freitas, Andre |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Controlling Equational Reasoning in Large Language Models with Prompt Interventions
by: Meadows, Jordan, et al.
Published: (2023)
by: Meadows, Jordan, et al.
Published: (2023)
A Survey in Mathematical Language Processing
by: Meadows, Jordan, et al.
Published: (2022)
by: Meadows, Jordan, et al.
Published: (2022)
FormalScience: Scalable Human-in-the-Loop Autoformalisation of Science with Agentic Code Generation in Lean
by: Meadows, Jordan, et al.
Published: (2026)
by: Meadows, Jordan, et al.
Published: (2026)
A Symbolic Framework for Evaluating Mathematical Reasoning and Generalisation with Transformers
by: Meadows, Jordan, et al.
Published: (2023)
by: Meadows, Jordan, et al.
Published: (2023)
FLRC: Fine-grained Low-Rank Compressor for Efficient LLM Inference
by: Lu, Yu-Chen, et al.
Published: (2025)
by: Lu, Yu-Chen, et al.
Published: (2025)
From Hypothesis to Premises: LLM-based Backward Logical Reasoning with Selective Symbolic Translation
by: Li, Qingchuan, et al.
Published: (2025)
by: Li, Qingchuan, et al.
Published: (2025)
Inferring Latent Intentions: Attributional Natural Language Inference in LLM Agents
by: Quan, Xin, et al.
Published: (2026)
by: Quan, Xin, et al.
Published: (2026)
Don't Take the Premise for Granted: Evaluating the Premise Critique Ability of Large Language Models
by: Li, Jinzhe, et al.
Published: (2025)
by: Li, Jinzhe, et al.
Published: (2025)
Abductive Inference in Retrieval-Augmented Language Models: Generating and Validating Missing Premises
by: Lin, Shiyin
Published: (2025)
by: Lin, Shiyin
Published: (2025)
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference
by: Jang, Wonsuk, et al.
Published: (2025)
by: Jang, Wonsuk, et al.
Published: (2025)
Redirected, Not Removed: Task-Dependent Stereotyping Reveals the Limits of LLM Alignments
by: Kumar, Divyanshu, et al.
Published: (2026)
by: Kumar, Divyanshu, et al.
Published: (2026)
Fine-grained Analysis of Brain-LLM Alignment through Input Attribution
by: Proietti, Michela, et al.
Published: (2025)
by: Proietti, Michela, et al.
Published: (2025)
Self-Explore: Enhancing Mathematical Reasoning in Language Models with Fine-grained Rewards
by: Hwang, Hyeonbin, et al.
Published: (2024)
by: Hwang, Hyeonbin, et al.
Published: (2024)
Adaptive LLM-Symbolic Reasoning via Dynamic Logical Solver Composition
by: Xu, Lei, et al.
Published: (2025)
by: Xu, Lei, et al.
Published: (2025)
TRACE: Training and Inference-Time Interpretability Analysis for Language Models
by: Aljaafari, Nura, et al.
Published: (2025)
by: Aljaafari, Nura, et al.
Published: (2025)
Towards Controllable Natural Language Inference through Lexical Inference Types
by: Zhang, Yingji, et al.
Published: (2023)
by: Zhang, Yingji, et al.
Published: (2023)
MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search
by: Yang, Zonglin, et al.
Published: (2025)
by: Yang, Zonglin, et al.
Published: (2025)
Improving Model Factuality with Fine-grained Critique-based Evaluator
by: Xie, Yiqing, et al.
Published: (2024)
by: Xie, Yiqing, et al.
Published: (2024)
Is Inference Mediated by Distinct Semantic Structures in LLMs? A Mechanistic Interpretation
by: Aljaafari, Nura, et al.
Published: (2026)
by: Aljaafari, Nura, et al.
Published: (2026)
Fine-grained Conversational Decoding via Isotropic and Proximal Search
by: Yao, Yuxuan, et al.
Published: (2023)
by: Yao, Yuxuan, et al.
Published: (2023)
CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data Synthesis
by: Feng, Ruixiang, et al.
Published: (2025)
by: Feng, Ruixiang, et al.
Published: (2025)
Controlled LLM-based Reasoning for Clinical Trial Retrieval
by: Jullien, Mael, et al.
Published: (2024)
by: Jullien, Mael, et al.
Published: (2024)
RuozhiBench: Evaluating LLMs with Logical Fallacies and Misleading Premises
by: Zhai, Zenan, et al.
Published: (2025)
by: Zhai, Zenan, et al.
Published: (2025)
Non-Linear Inference Time Intervention: Improving LLM Truthfulness
by: Hoscilowicz, Jakub, et al.
Published: (2024)
by: Hoscilowicz, Jakub, et al.
Published: (2024)
Don't Let It Hallucinate: Premise Verification via Retrieval-Augmented Logical Reasoning
by: Qin, Yuehan, et al.
Published: (2025)
by: Qin, Yuehan, et al.
Published: (2025)
Removing RLHF Protections in GPT-4 via Fine-Tuning
by: Zhan, Qiusi, et al.
Published: (2023)
by: Zhan, Qiusi, et al.
Published: (2023)
MASA: LLM-Driven Multi-Agent Systems for Autoformalization
by: Zhang, Lan, et al.
Published: (2025)
by: Zhang, Lan, et al.
Published: (2025)
FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
by: Ye, Seonghyeon, et al.
Published: (2023)
by: Ye, Seonghyeon, et al.
Published: (2023)
MultiHoax: A Dataset of Multi-hop False-Premise Questions
by: Shafiei, Mohammadamin, et al.
Published: (2025)
by: Shafiei, Mohammadamin, et al.
Published: (2025)
Premise Order Matters in Reasoning with Large Language Models
by: Chen, Xinyun, et al.
Published: (2024)
by: Chen, Xinyun, et al.
Published: (2024)
Making Implicit Premises Explicit in Logical Understanding of Enthymemes
by: Feng, Xuyao, et al.
Published: (2026)
by: Feng, Xuyao, et al.
Published: (2026)
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
by: Abdaljalil, Samir, et al.
Published: (2025)
by: Abdaljalil, Samir, et al.
Published: (2025)
XATU: A Fine-grained Instruction-based Benchmark for Explainable Text Updates
by: Zhang, Haopeng, et al.
Published: (2023)
by: Zhang, Haopeng, et al.
Published: (2023)
Decompose-and-Formalise: Recursively Verifiable Natural Language Inference
by: Quan, Xin, et al.
Published: (2026)
by: Quan, Xin, et al.
Published: (2026)
Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning
by: Zhang, Lan, et al.
Published: (2025)
by: Zhang, Lan, et al.
Published: (2025)
Training Language Models to Generate Text with Citations via Fine-grained Rewards
by: Huang, Chengyu, et al.
Published: (2024)
by: Huang, Chengyu, et al.
Published: (2024)
MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme Intervention
by: Jha, Prince, et al.
Published: (2024)
by: Jha, Prince, et al.
Published: (2024)
Reinforcement Retrieval Leveraging Fine-grained Feedback for Fact Checking News Claims with Black-Box LLM
by: Zhang, Xuan, et al.
Published: (2024)
by: Zhang, Xuan, et al.
Published: (2024)
FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain
by: Afzal, Anum, et al.
Published: (2025)
by: Afzal, Anum, et al.
Published: (2025)
FLAT-LLM: Fine-grained Low-rank Activation Space Transformation for Large Language Model Compression
by: Tian, Jiayi, et al.
Published: (2025)
by: Tian, Jiayi, et al.
Published: (2025)
Similar Items
-
Controlling Equational Reasoning in Large Language Models with Prompt Interventions
by: Meadows, Jordan, et al.
Published: (2023) -
A Survey in Mathematical Language Processing
by: Meadows, Jordan, et al.
Published: (2022) -
FormalScience: Scalable Human-in-the-Loop Autoformalisation of Science with Agentic Code Generation in Lean
by: Meadows, Jordan, et al.
Published: (2026) -
A Symbolic Framework for Evaluating Mathematical Reasoning and Generalisation with Transformers
by: Meadows, Jordan, et al.
Published: (2023) -
FLRC: Fine-grained Low-Rank Compressor for Efficient LLM Inference
by: Lu, Yu-Chen, et al.
Published: (2025)