EngTrace: A Symbolic Benchmark for Verifiable Process Supervision of Engineering Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Gull, Ayesha, Safder, Muhammad Usman, Elbadry, Rania, Zhang, Fan, Stoyanov, Veselin, Nakov, Preslav, Xie, Zhuohan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
by: Xie, Zhuohan, et al.
Published: (2025)
by: Xie, Zhuohan, et al.
Published: (2025)
SAHM: A Benchmark for Arabic Financial and Shari'ah-Compliant Reasoning
by: Elbadry, Rania, et al.
Published: (2026)
by: Elbadry, Rania, et al.
Published: (2026)
FinReporting: An Agentic Workflow for Localized Reporting of Cross-Jurisdiction Financial Disclosures
by: Zhang, Fan, et al.
Published: (2026)
by: Zhang, Fan, et al.
Published: (2026)
The CLEF-2026 FinMMEval Lab: Multilingual and Multimodal Evaluation of Financial AI Systems
by: Xie, Zhuohan, et al.
Published: (2026)
by: Xie, Zhuohan, et al.
Published: (2026)
The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations
by: Elbadry, Rania, et al.
Published: (2026)
by: Elbadry, Rania, et al.
Published: (2026)
DUAL-Bench: Measuring Over-Refusal and Robustness in Vision-Language Models
by: Ren, Kaixuan, et al.
Published: (2025)
by: Ren, Kaixuan, et al.
Published: (2025)
FinCARDS: Card-Based Analyst Reranking for Financial Document Question Answering
by: Zhou, Yixi, et al.
Published: (2026)
by: Zhou, Yixi, et al.
Published: (2026)
Instruction-Guided Poetry Generation in Arabic and Its Dialects
by: Sadallah, Abdelrahman, et al.
Published: (2026)
by: Sadallah, Abdelrahman, et al.
Published: (2026)
SQLStructEval: Structural Evaluation of LLM Text-to-SQL Generation
by: Zhou, Yixi, et al.
Published: (2026)
by: Zhou, Yixi, et al.
Published: (2026)
How Does Prefix Matter in Reasoning Model Tuning?
by: Tomar, Raj Vardhan, et al.
Published: (2026)
by: Tomar, Raj Vardhan, et al.
Published: (2026)
Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World
by: Almheiri, Saeed, et al.
Published: (2025)
by: Almheiri, Saeed, et al.
Published: (2025)
UnsafeChain: Enhancing Reasoning Model Safety via Hard Cases
by: Tomar, Raj Vardhan, et al.
Published: (2025)
by: Tomar, Raj Vardhan, et al.
Published: (2025)
From Generation to Detection: A Multimodal Multi-Task Dataset for Benchmarking Health Misinformation
by: Zhang, Zhihao, et al.
Published: (2025)
by: Zhang, Zhihao, et al.
Published: (2025)
MuDRiC: Multi-Dialect Reasoning for Arabic Commonsense Validation
by: Elozeiri, Kareem, et al.
Published: (2025)
by: Elozeiri, Kareem, et al.
Published: (2025)
Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
by: Su, Jinyan, et al.
Published: (2025)
by: Su, Jinyan, et al.
Published: (2025)
PRIME: A Process-Outcome Alignment Benchmark for Verifiable Reasoning in Mathematics and Engineering
by: Wang, Xiangfeng, et al.
Published: (2026)
by: Wang, Xiangfeng, et al.
Published: (2026)
Can Machines Resonate with Humans? Evaluating the Emotional and Empathic Comprehension of LMs
by: Manzoor, Muhammad Arslan, et al.
Published: (2024)
by: Manzoor, Muhammad Arslan, et al.
Published: (2024)
Large Language Models are Few-Shot Training Example Generators: A Case Study in Fallacy Recognition
by: Alhindi, Tariq, et al.
Published: (2023)
by: Alhindi, Tariq, et al.
Published: (2023)
Rethinking STS and NLI in Large Language Models
by: Wang, Yuxia, et al.
Published: (2023)
by: Wang, Yuxia, et al.
Published: (2023)
CoDet-M4: Detecting Machine-Generated Code in Multi-Lingual, Multi-Generator and Multi-Domain Settings
by: Orel, Daniil, et al.
Published: (2025)
by: Orel, Daniil, et al.
Published: (2025)
YaPO: Learnable Sparse Activation Steering Vectors for Domain Adaptation
by: Bounhar, Abdelaziz, et al.
Published: (2026)
by: Bounhar, Abdelaziz, et al.
Published: (2026)
Detecting Propaganda Techniques in Code-Switched Social Media Text
by: Salman, Muhammad Umar, et al.
Published: (2023)
by: Salman, Muhammad Umar, et al.
Published: (2023)
From Chaos to Clarity: Claim Normalization to Empower Fact-Checking
by: Sundriyal, Megha, et al.
Published: (2023)
by: Sundriyal, Megha, et al.
Published: (2023)
Adapting Fake News Detection to the Era of Large Language Models
by: Su, Jinyan, et al.
Published: (2023)
by: Su, Jinyan, et al.
Published: (2023)
FIRE: Fact-checking with Iterative Retrieval and Verification
by: Xie, Zhuohan, et al.
Published: (2024)
by: Xie, Zhuohan, et al.
Published: (2024)
Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models
by: Sahnan, Dhruv, et al.
Published: (2026)
by: Sahnan, Dhruv, et al.
Published: (2026)
Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages
by: Hu, Yujia, et al.
Published: (2025)
by: Hu, Yujia, et al.
Published: (2025)
TART: An Open-Source Tool-Augmented Framework for Explainable Table-based Reasoning
by: Lu, Xinyuan, et al.
Published: (2024)
by: Lu, Xinyuan, et al.
Published: (2024)
VISPA: Pluralistic Alignment via Automatic Value Selection and Activation
by: Zheng, Shenyan, et al.
Published: (2026)
by: Zheng, Shenyan, et al.
Published: (2026)
Correct Answers from Sound Reasoning: Verifiable Process Supervision for Language Models
by: Kim, Kyuyoung, et al.
Published: (2026)
by: Kim, Kyuyoung, et al.
Published: (2026)
Guided Verifier: Collaborative Multimodal Reasoning via Dynamic Process Supervision
by: Sun, Lingzhuang, et al.
Published: (2026)
by: Sun, Lingzhuang, et al.
Published: (2026)
ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety
by: Bates, Luke, et al.
Published: (2025)
by: Bates, Luke, et al.
Published: (2025)
UNCERTAINTY-LINE: Length-Invariant Estimation of Uncertainty for Large Language Models
by: Vashurin, Roman, et al.
Published: (2025)
by: Vashurin, Roman, et al.
Published: (2025)
Missci: Reconstructing Fallacies in Misrepresented Science
by: Glockner, Max, et al.
Published: (2024)
by: Glockner, Max, et al.
Published: (2024)
Grounding Fallacies Misrepresenting Scientific Publications in Evidence
by: Glockner, Max, et al.
Published: (2024)
by: Glockner, Max, et al.
Published: (2024)
Profiling News Media for Factuality and Bias Using LLMs and the Fact-Checking Methodology of Human Experts
by: Mujahid, Zain Muhammad, et al.
Published: (2025)
by: Mujahid, Zain Muhammad, et al.
Published: (2025)
Can LLMs Automate Fact-Checking Article Writing?
by: Sahnan, Dhruv, et al.
Published: (2025)
by: Sahnan, Dhruv, et al.
Published: (2025)
Multimodal Large Language Models to Support Real-World Fact-Checking
by: Geng, Jiahui, et al.
Published: (2024)
by: Geng, Jiahui, et al.
Published: (2024)
MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing
by: Agarwal, Siddhant, et al.
Published: (2024)
by: Agarwal, Siddhant, et al.
Published: (2024)
On a Novel Application of Wasserstein-Procrustes for Unsupervised Cross-Lingual Learning
by: Ramírez, Guillem, et al.
Published: (2020)
by: Ramírez, Guillem, et al.
Published: (2020)
Similar Items
-
FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
by: Xie, Zhuohan, et al.
Published: (2025) -
SAHM: A Benchmark for Arabic Financial and Shari'ah-Compliant Reasoning
by: Elbadry, Rania, et al.
Published: (2026) -
FinReporting: An Agentic Workflow for Localized Reporting of Cross-Jurisdiction Financial Disclosures
by: Zhang, Fan, et al.
Published: (2026) -
The CLEF-2026 FinMMEval Lab: Multilingual and Multimodal Evaluation of Financial AI Systems
by: Xie, Zhuohan, et al.
Published: (2026) -
The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations
by: Elbadry, Rania, et al.
Published: (2026)