FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mohammadzadeh, Saeed, Hamdi, Erfan, Shor, Joel, Lejeune, Emma |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
JAM: Controllable and Responsible Text Generation via Causal Reasoning and Latent Vector Manipulation
von: Huang, Yingbing, et al.
Veröffentlicht: (2025)
von: Huang, Yingbing, et al.
Veröffentlicht: (2025)
QHackBench: Benchmarking Large Language Models for Quantum Code Generation Using PennyLane Hackathon Challenges
von: Basit, Abdul, et al.
Veröffentlicht: (2025)
von: Basit, Abdul, et al.
Veröffentlicht: (2025)
Towards Probabilistic Question Answering Over Tabular Data
von: Shen, Chen, et al.
Veröffentlicht: (2025)
von: Shen, Chen, et al.
Veröffentlicht: (2025)
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
von: Jia, Xiao
Veröffentlicht: (2026)
von: Jia, Xiao
Veröffentlicht: (2026)
Response Uncertainty and Probe Modeling: Two Sides of the Same Coin in LLM Interpretability?
von: Wang, Yongjie, et al.
Veröffentlicht: (2025)
von: Wang, Yongjie, et al.
Veröffentlicht: (2025)
Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent
von: Xia, Bowei, et al.
Veröffentlicht: (2026)
von: Xia, Bowei, et al.
Veröffentlicht: (2026)
An Iterative Optimizing Framework for Radiology Report Summarization with ChatGPT
von: Ma, Chong, et al.
Veröffentlicht: (2023)
von: Ma, Chong, et al.
Veröffentlicht: (2023)
Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering
von: Chen, Tiejin, et al.
Veröffentlicht: (2026)
von: Chen, Tiejin, et al.
Veröffentlicht: (2026)
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
von: Wang, Yihao, et al.
Veröffentlicht: (2026)
von: Wang, Yihao, et al.
Veröffentlicht: (2026)
Understanding the Uncertainty of LLM Explanations: A Perspective Based on Reasoning Topology
von: Da, Longchao, et al.
Veröffentlicht: (2025)
von: Da, Longchao, et al.
Veröffentlicht: (2025)
React-ing to Grace Hopper 200: Five Open-Weights Coding Models, One React Native App, One GH200, One Weekend
von: Potanin, Alex
Veröffentlicht: (2026)
von: Potanin, Alex
Veröffentlicht: (2026)
Kodezi Chronos: A Debugging-First Language Model for Repository-Scale Code Understanding
von: Khan, Ishraq, et al.
Veröffentlicht: (2025)
von: Khan, Ishraq, et al.
Veröffentlicht: (2025)
Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution
von: Liu, Xiaoou, et al.
Veröffentlicht: (2026)
von: Liu, Xiaoou, et al.
Veröffentlicht: (2026)
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
von: Zhang, Junwen, et al.
Veröffentlicht: (2025)
von: Zhang, Junwen, et al.
Veröffentlicht: (2025)
PennyCoder: Efficient Domain-Specific LLMs for PennyLane-Based Quantum Code Generation
von: Basit, Abdul, et al.
Veröffentlicht: (2025)
von: Basit, Abdul, et al.
Veröffentlicht: (2025)
AdvFusion: Adapter-based Knowledge Transfer for Code Summarization on Code Language Models
von: Saberi, Iman, et al.
Veröffentlicht: (2023)
von: Saberi, Iman, et al.
Veröffentlicht: (2023)
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
von: Badshah, Sher, et al.
Veröffentlicht: (2024)
von: Badshah, Sher, et al.
Veröffentlicht: (2024)
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics
von: Peyronnet, Antoine, et al.
Veröffentlicht: (2026)
von: Peyronnet, Antoine, et al.
Veröffentlicht: (2026)
Biomedical Visual Instruction Tuning with Clinician Preference Alignment
von: Cui, Hejie, et al.
Veröffentlicht: (2024)
von: Cui, Hejie, et al.
Veröffentlicht: (2024)
DUCTILE: Agentic LLM Orchestration of Engineering Analysis in Product Development Practice
von: Pradas-Gomez, Alejandro, et al.
Veröffentlicht: (2026)
von: Pradas-Gomez, Alejandro, et al.
Veröffentlicht: (2026)
Open Source Evolutionary Computation with Chips-n-Salsa
von: Cicirello, Vincent A.
Veröffentlicht: (2024)
von: Cicirello, Vincent A.
Veröffentlicht: (2024)
Towards Robust Surrogate Models: Benchmarking Machine Learning Approaches to Expediting Phase Field Simulations of Brittle Fracture
von: Hamdi, Erfan, et al.
Veröffentlicht: (2025)
von: Hamdi, Erfan, et al.
Veröffentlicht: (2025)
Large Language Models are Inconsistent and Biased Evaluators
von: Stureborg, Rickard, et al.
Veröffentlicht: (2024)
von: Stureborg, Rickard, et al.
Veröffentlicht: (2024)
D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree
von: Lei, Xiang, et al.
Veröffentlicht: (2025)
von: Lei, Xiang, et al.
Veröffentlicht: (2025)
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
von: Kaiser, Daniel, et al.
Veröffentlicht: (2025)
von: Kaiser, Daniel, et al.
Veröffentlicht: (2025)
On measuring grounding and generalizing grounding problems
von: Quigley, Daniel, et al.
Veröffentlicht: (2025)
von: Quigley, Daniel, et al.
Veröffentlicht: (2025)
See-Saw Generative Mechanism for Scalable Recursive Code Generation with Generative AI
von: Vsevolodovna, Ruslan Idelfonso Magaña
Veröffentlicht: (2024)
von: Vsevolodovna, Ruslan Idelfonso Magaña
Veröffentlicht: (2024)
MicroRemed: Benchmarking LLMs in Microservices Remediation
von: Zhang, Lingzhe, et al.
Veröffentlicht: (2025)
von: Zhang, Lingzhe, et al.
Veröffentlicht: (2025)
Superior Scoring Rules for Probabilistic Evaluation of Single-Label Multi-Class Classification Tasks
von: Ahmadian, Rouhollah, et al.
Veröffentlicht: (2024)
von: Ahmadian, Rouhollah, et al.
Veröffentlicht: (2024)
Low-Resource English-Tigrinya MT: Leveraging Multilingual Models, Custom Tokenizers, and Clean Evaluation Benchmarks
von: Teklehaymanot, Hailay Kidu, et al.
Veröffentlicht: (2025)
von: Teklehaymanot, Hailay Kidu, et al.
Veröffentlicht: (2025)
Harnessing non-adversarial robustness in large language models
von: Zhou, Qinghua, et al.
Veröffentlicht: (2026)
von: Zhou, Qinghua, et al.
Veröffentlicht: (2026)
Rethinking the Multilingual Reasoning Gap with Layer Swap
von: Lasbordes, Maxence, et al.
Veröffentlicht: (2026)
von: Lasbordes, Maxence, et al.
Veröffentlicht: (2026)
A Reproducible, Scalable Pipeline for Synthesizing Autoregressive Model Literature
von: Alpay, Faruk, et al.
Veröffentlicht: (2025)
von: Alpay, Faruk, et al.
Veröffentlicht: (2025)
How much do LLMs learn from negative examples?
von: Hamdan, Shadi, et al.
Veröffentlicht: (2025)
von: Hamdan, Shadi, et al.
Veröffentlicht: (2025)
Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
von: Fan, Jingxing, et al.
Veröffentlicht: (2025)
von: Fan, Jingxing, et al.
Veröffentlicht: (2025)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
von: Yang, Yibo
Veröffentlicht: (2025)
von: Yang, Yibo
Veröffentlicht: (2025)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
von: Kim, Heejun, et al.
Veröffentlicht: (2026)
von: Kim, Heejun, et al.
Veröffentlicht: (2026)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
von: Pather, Kaviraj, et al.
Veröffentlicht: (2025)
von: Pather, Kaviraj, et al.
Veröffentlicht: (2025)
CELI: Controller-Embedded Language Model Interactions
von: Wagner, Jan-Samuel, et al.
Veröffentlicht: (2024)
von: Wagner, Jan-Samuel, et al.
Veröffentlicht: (2024)
A transfer learning approach for automatic conflicts detection in software requirement sentence pairs based on dual encoders
von: Wang, Yizheng, et al.
Veröffentlicht: (2025)
von: Wang, Yizheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
JAM: Controllable and Responsible Text Generation via Causal Reasoning and Latent Vector Manipulation
von: Huang, Yingbing, et al.
Veröffentlicht: (2025) -
QHackBench: Benchmarking Large Language Models for Quantum Code Generation Using PennyLane Hackathon Challenges
von: Basit, Abdul, et al.
Veröffentlicht: (2025) -
Towards Probabilistic Question Answering Over Tabular Data
von: Shen, Chen, et al.
Veröffentlicht: (2025) -
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
von: Jia, Xiao
Veröffentlicht: (2026) -
Response Uncertainty and Probe Modeling: Two Sides of the Same Coin in LLM Interpretability?
von: Wang, Yongjie, et al.
Veröffentlicht: (2025)