Exploring the Limits of Fine-grained LLM-based Physics Inference via Premise Removal Interventions
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Meadows, Jordan, James, Tamsin, Freitas, Andre |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Controlling Equational Reasoning in Large Language Models with Prompt Interventions
par: Meadows, Jordan, et autres
Publié: (2023)
par: Meadows, Jordan, et autres
Publié: (2023)
A Survey in Mathematical Language Processing
par: Meadows, Jordan, et autres
Publié: (2022)
par: Meadows, Jordan, et autres
Publié: (2022)
FormalScience: Scalable Human-in-the-Loop Autoformalisation of Science with Agentic Code Generation in Lean
par: Meadows, Jordan, et autres
Publié: (2026)
par: Meadows, Jordan, et autres
Publié: (2026)
A Symbolic Framework for Evaluating Mathematical Reasoning and Generalisation with Transformers
par: Meadows, Jordan, et autres
Publié: (2023)
par: Meadows, Jordan, et autres
Publié: (2023)
FLRC: Fine-grained Low-Rank Compressor for Efficient LLM Inference
par: Lu, Yu-Chen, et autres
Publié: (2025)
par: Lu, Yu-Chen, et autres
Publié: (2025)
From Hypothesis to Premises: LLM-based Backward Logical Reasoning with Selective Symbolic Translation
par: Li, Qingchuan, et autres
Publié: (2025)
par: Li, Qingchuan, et autres
Publié: (2025)
Inferring Latent Intentions: Attributional Natural Language Inference in LLM Agents
par: Quan, Xin, et autres
Publié: (2026)
par: Quan, Xin, et autres
Publié: (2026)
Don't Take the Premise for Granted: Evaluating the Premise Critique Ability of Large Language Models
par: Li, Jinzhe, et autres
Publié: (2025)
par: Li, Jinzhe, et autres
Publié: (2025)
Abductive Inference in Retrieval-Augmented Language Models: Generating and Validating Missing Premises
par: Lin, Shiyin
Publié: (2025)
par: Lin, Shiyin
Publié: (2025)
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference
par: Jang, Wonsuk, et autres
Publié: (2025)
par: Jang, Wonsuk, et autres
Publié: (2025)
Redirected, Not Removed: Task-Dependent Stereotyping Reveals the Limits of LLM Alignments
par: Kumar, Divyanshu, et autres
Publié: (2026)
par: Kumar, Divyanshu, et autres
Publié: (2026)
Fine-grained Analysis of Brain-LLM Alignment through Input Attribution
par: Proietti, Michela, et autres
Publié: (2025)
par: Proietti, Michela, et autres
Publié: (2025)
Self-Explore: Enhancing Mathematical Reasoning in Language Models with Fine-grained Rewards
par: Hwang, Hyeonbin, et autres
Publié: (2024)
par: Hwang, Hyeonbin, et autres
Publié: (2024)
Adaptive LLM-Symbolic Reasoning via Dynamic Logical Solver Composition
par: Xu, Lei, et autres
Publié: (2025)
par: Xu, Lei, et autres
Publié: (2025)
TRACE: Training and Inference-Time Interpretability Analysis for Language Models
par: Aljaafari, Nura, et autres
Publié: (2025)
par: Aljaafari, Nura, et autres
Publié: (2025)
Towards Controllable Natural Language Inference through Lexical Inference Types
par: Zhang, Yingji, et autres
Publié: (2023)
par: Zhang, Yingji, et autres
Publié: (2023)
MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search
par: Yang, Zonglin, et autres
Publié: (2025)
par: Yang, Zonglin, et autres
Publié: (2025)
Improving Model Factuality with Fine-grained Critique-based Evaluator
par: Xie, Yiqing, et autres
Publié: (2024)
par: Xie, Yiqing, et autres
Publié: (2024)
Is Inference Mediated by Distinct Semantic Structures in LLMs? A Mechanistic Interpretation
par: Aljaafari, Nura, et autres
Publié: (2026)
par: Aljaafari, Nura, et autres
Publié: (2026)
Fine-grained Conversational Decoding via Isotropic and Proximal Search
par: Yao, Yuxuan, et autres
Publié: (2023)
par: Yao, Yuxuan, et autres
Publié: (2023)
CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data Synthesis
par: Feng, Ruixiang, et autres
Publié: (2025)
par: Feng, Ruixiang, et autres
Publié: (2025)
Controlled LLM-based Reasoning for Clinical Trial Retrieval
par: Jullien, Mael, et autres
Publié: (2024)
par: Jullien, Mael, et autres
Publié: (2024)
RuozhiBench: Evaluating LLMs with Logical Fallacies and Misleading Premises
par: Zhai, Zenan, et autres
Publié: (2025)
par: Zhai, Zenan, et autres
Publié: (2025)
Non-Linear Inference Time Intervention: Improving LLM Truthfulness
par: Hoscilowicz, Jakub, et autres
Publié: (2024)
par: Hoscilowicz, Jakub, et autres
Publié: (2024)
Don't Let It Hallucinate: Premise Verification via Retrieval-Augmented Logical Reasoning
par: Qin, Yuehan, et autres
Publié: (2025)
par: Qin, Yuehan, et autres
Publié: (2025)
Removing RLHF Protections in GPT-4 via Fine-Tuning
par: Zhan, Qiusi, et autres
Publié: (2023)
par: Zhan, Qiusi, et autres
Publié: (2023)
MASA: LLM-Driven Multi-Agent Systems for Autoformalization
par: Zhang, Lan, et autres
Publié: (2025)
par: Zhang, Lan, et autres
Publié: (2025)
FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
par: Ye, Seonghyeon, et autres
Publié: (2023)
par: Ye, Seonghyeon, et autres
Publié: (2023)
MultiHoax: A Dataset of Multi-hop False-Premise Questions
par: Shafiei, Mohammadamin, et autres
Publié: (2025)
par: Shafiei, Mohammadamin, et autres
Publié: (2025)
Premise Order Matters in Reasoning with Large Language Models
par: Chen, Xinyun, et autres
Publié: (2024)
par: Chen, Xinyun, et autres
Publié: (2024)
Making Implicit Premises Explicit in Logical Understanding of Enthymemes
par: Feng, Xuyao, et autres
Publié: (2026)
par: Feng, Xuyao, et autres
Publié: (2026)
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
par: Abdaljalil, Samir, et autres
Publié: (2025)
par: Abdaljalil, Samir, et autres
Publié: (2025)
XATU: A Fine-grained Instruction-based Benchmark for Explainable Text Updates
par: Zhang, Haopeng, et autres
Publié: (2023)
par: Zhang, Haopeng, et autres
Publié: (2023)
Decompose-and-Formalise: Recursively Verifiable Natural Language Inference
par: Quan, Xin, et autres
Publié: (2026)
par: Quan, Xin, et autres
Publié: (2026)
Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning
par: Zhang, Lan, et autres
Publié: (2025)
par: Zhang, Lan, et autres
Publié: (2025)
Training Language Models to Generate Text with Citations via Fine-grained Rewards
par: Huang, Chengyu, et autres
Publié: (2024)
par: Huang, Chengyu, et autres
Publié: (2024)
MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme Intervention
par: Jha, Prince, et autres
Publié: (2024)
par: Jha, Prince, et autres
Publié: (2024)
Reinforcement Retrieval Leveraging Fine-grained Feedback for Fact Checking News Claims with Black-Box LLM
par: Zhang, Xuan, et autres
Publié: (2024)
par: Zhang, Xuan, et autres
Publié: (2024)
FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain
par: Afzal, Anum, et autres
Publié: (2025)
par: Afzal, Anum, et autres
Publié: (2025)
FLAT-LLM: Fine-grained Low-rank Activation Space Transformation for Large Language Model Compression
par: Tian, Jiayi, et autres
Publié: (2025)
par: Tian, Jiayi, et autres
Publié: (2025)
Documents similaires
-
Controlling Equational Reasoning in Large Language Models with Prompt Interventions
par: Meadows, Jordan, et autres
Publié: (2023) -
A Survey in Mathematical Language Processing
par: Meadows, Jordan, et autres
Publié: (2022) -
FormalScience: Scalable Human-in-the-Loop Autoformalisation of Science with Agentic Code Generation in Lean
par: Meadows, Jordan, et autres
Publié: (2026) -
A Symbolic Framework for Evaluating Mathematical Reasoning and Generalisation with Transformers
par: Meadows, Jordan, et autres
Publié: (2023) -
FLRC: Fine-grained Low-Rank Compressor for Efficient LLM Inference
par: Lu, Yu-Chen, et autres
Publié: (2025)