Experiments or Outcomes? Probing Scientific Feasibility in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Mohammadi, Seyedali, Gaur, Manas, Ferraro, Francis |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WellDunn: On the Robustness and Explainability of Language Models and Large Language Models in Identifying Wellness Dimensions
by: Mohammadi, Seyedali, et al.
Published: (2024)
by: Mohammadi, Seyedali, et al.
Published: (2024)
Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions
by: Mohammadi, Seyedali, et al.
Published: (2025)
by: Mohammadi, Seyedali, et al.
Published: (2025)
IoT-Based Preventive Mental Health Using Knowledge Graphs and Standards for Better Well-Being
by: Gyrard, Amelie, et al.
Published: (2024)
by: Gyrard, Amelie, et al.
Published: (2024)
Can LLMs Obfuscate Code? A Systematic Analysis of Large Language Models into Assembly Code Obfuscation
by: Mohseni, Seyedreza, et al.
Published: (2024)
by: Mohseni, Seyedreza, et al.
Published: (2024)
Attribution in Scientific Literature: New Benchmark and Methods
by: Saxena, Yash, et al.
Published: (2024)
by: Saxena, Yash, et al.
Published: (2024)
SaGE: Evaluating Moral Consistency in Large Language Models
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
LingVarBench: Benchmarking LLMs on Entity Recognitions and Linguistic Verbalization Patterns in Phone-Call Transcripts
by: Mohammadi, Seyedali, et al.
Published: (2025)
by: Mohammadi, Seyedali, et al.
Published: (2025)
Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings
by: Das, Nilanjana, et al.
Published: (2026)
by: Das, Nilanjana, et al.
Published: (2026)
Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA
by: Lamba, Naveen, et al.
Published: (2025)
by: Lamba, Naveen, et al.
Published: (2025)
SymLoc: Symbolic Localization of Hallucination across HaluEval and TruthfulQA
by: Lamba, Naveen, et al.
Published: (2025)
by: Lamba, Naveen, et al.
Published: (2025)
Neurosymbolic Retrievers for Retrieval-augmented Generation
by: Saxena, Yash, et al.
Published: (2026)
by: Saxena, Yash, et al.
Published: (2026)
COBIAS: Assessing the Contextual Reliability of Bias Benchmarks for Language Models
by: Govil, Priyanshul, et al.
Published: (2024)
by: Govil, Priyanshul, et al.
Published: (2024)
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context
by: Das, Nilanjana, et al.
Published: (2024)
by: Das, Nilanjana, et al.
Published: (2024)
Bridging Reasoning Trajectories in On-Policy Distillation via Near-Future Guidance
by: Jiang, Yuxuan, et al.
Published: (2026)
by: Jiang, Yuxuan, et al.
Published: (2026)
Exploring Cultural Variations in Moral Judgments with Large Language Models
by: Mohammadi, Hadi, et al.
Published: (2025)
by: Mohammadi, Hadi, et al.
Published: (2025)
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
by: Haider, Batool, et al.
Published: (2025)
by: Haider, Batool, et al.
Published: (2025)
Explaining Large Language Models Decisions Using Shapley Values
by: Mohammadi, Behnam
Published: (2024)
by: Mohammadi, Behnam
Published: (2024)
Beyond Memorization: Testing LLM Reasoning on Unseen Theory of Computation Tasks
by: Shelat, Shlok, et al.
Published: (2026)
by: Shelat, Shlok, et al.
Published: (2026)
Probing Neural Topology of Large Language Models
by: Zheng, Yu, et al.
Published: (2025)
by: Zheng, Yu, et al.
Published: (2025)
Probing Causality Manipulation of Large Language Models
by: Zhang, Chenyang, et al.
Published: (2024)
by: Zhang, Chenyang, et al.
Published: (2024)
Large Language Models as Mirrors of Societal Moral Standards
by: Papadopoulou, Evi, et al.
Published: (2024)
by: Papadopoulou, Evi, et al.
Published: (2024)
Creativity Has Left the Chat: The Price of Debiasing Language Models
by: Mohammadi, Behnam
Published: (2024)
by: Mohammadi, Behnam
Published: (2024)
Probing the Robustness of Theory of Mind in Large Language Models
by: Nickel, Christian, et al.
Published: (2024)
by: Nickel, Christian, et al.
Published: (2024)
Probing the Difficulty Perception Mechanism of Large Language Models
by: Lee, Sunbowen, et al.
Published: (2025)
by: Lee, Sunbowen, et al.
Published: (2025)
Scientific Computing with Large Language Models
by: Culver, Christopher, et al.
Published: (2024)
by: Culver, Christopher, et al.
Published: (2024)
Set-Valued Prediction for Large Language Models with Feasibility-Aware Coverage Guarantees
by: Li, Ye, et al.
Published: (2026)
by: Li, Ye, et al.
Published: (2026)
ChatSR: Multimodal Large Language Models for Scientific Formula Discovery
by: Li, Yanjie, et al.
Published: (2024)
by: Li, Yanjie, et al.
Published: (2024)
Improving Scientific Hypothesis Generation with Knowledge Grounded Large Language Models
by: Xiong, Guangzhi, et al.
Published: (2024)
by: Xiong, Guangzhi, et al.
Published: (2024)
Towards Efficient Large Language Models for Scientific Text: A Review
by: To, Huy Quoc, et al.
Published: (2024)
by: To, Huy Quoc, et al.
Published: (2024)
Large Language Models for Automated Open-domain Scientific Hypotheses Discovery
by: Yang, Zonglin, et al.
Published: (2023)
by: Yang, Zonglin, et al.
Published: (2023)
EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models
by: Mohammadi, Hadi, et al.
Published: (2025)
by: Mohammadi, Hadi, et al.
Published: (2025)
Stuck in the Matrix: Probing Spatial Reasoning in Large Language Models
by: Bai, Maggie, et al.
Published: (2025)
by: Bai, Maggie, et al.
Published: (2025)
Neural Probe-Based Hallucination Detection for Large Language Models
by: Liang, Shize, et al.
Published: (2025)
by: Liang, Shize, et al.
Published: (2025)
Counterfactual Probing for Hallucination Detection and Mitigation in Large Language Models
by: Feng, Yijun
Published: (2025)
by: Feng, Yijun
Published: (2025)
PerMedCQA: Benchmarking Large Language Models on Medical Consumer Question Answering in Persian Language
by: Jamali, Naghmeh, et al.
Published: (2025)
by: Jamali, Naghmeh, et al.
Published: (2025)
Evaluating the Feasibility and Accuracy of Large Language Models for Medical History-Taking in Obstetrics and Gynecology
by: Liu, Dou, et al.
Published: (2025)
by: Liu, Dou, et al.
Published: (2025)
Large Language Models as Evaluators for Scientific Synthesis
by: Evans, Julia, et al.
Published: (2024)
by: Evans, Julia, et al.
Published: (2024)
Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning
by: Yan, Yibo, et al.
Published: (2025)
by: Yan, Yibo, et al.
Published: (2025)
Can Large Language Model Summarizers Adapt to Diverse Scientific Communication Goals?
by: Fonseca, Marcio, et al.
Published: (2024)
by: Fonseca, Marcio, et al.
Published: (2024)
MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding
by: Jiang, Mohan, et al.
Published: (2025)
by: Jiang, Mohan, et al.
Published: (2025)
Similar Items
-
WellDunn: On the Robustness and Explainability of Language Models and Large Language Models in Identifying Wellness Dimensions
by: Mohammadi, Seyedali, et al.
Published: (2024) -
Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions
by: Mohammadi, Seyedali, et al.
Published: (2025) -
IoT-Based Preventive Mental Health Using Knowledge Graphs and Standards for Better Well-Being
by: Gyrard, Amelie, et al.
Published: (2024) -
Can LLMs Obfuscate Code? A Systematic Analysis of Large Language Models into Assembly Code Obfuscation
by: Mohseni, Seyedreza, et al.
Published: (2024) -
Attribution in Scientific Literature: New Benchmark and Methods
by: Saxena, Yash, et al.
Published: (2024)