LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Zenghui, Li, Man, Fang, Xiaoke, Zhou, Xinyi, Li, Weibin, Zheng, Zheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MASTEST: A LLM-Based Multi-Agent System For RESTful API Tests
di: Han, Xiaoke, et al.
Pubblicazione: (2025)
di: Han, Xiaoke, et al.
Pubblicazione: (2025)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
di: Gao, Shuzheng, et al.
Pubblicazione: (2025)
di: Gao, Shuzheng, et al.
Pubblicazione: (2025)
LiCoEval: Evaluating LLMs on License Compliance in Code Generation
di: Xu, Weiwei, et al.
Pubblicazione: (2024)
di: Xu, Weiwei, et al.
Pubblicazione: (2024)
Evaluating Human Trajectory Prediction with Metamorphic Testing
di: Spieker, Helge, et al.
Pubblicazione: (2024)
di: Spieker, Helge, et al.
Pubblicazione: (2024)
Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks
di: Zhou, Yongxi, et al.
Pubblicazione: (2026)
di: Zhou, Yongxi, et al.
Pubblicazione: (2026)
Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations
di: Palacio, David N., et al.
Pubblicazione: (2024)
di: Palacio, David N., et al.
Pubblicazione: (2024)
Can LLMs Reason Like Automated Theorem Provers for Rust Verification? VCoT-Bench: Evaluating via Verification Chain of Thought
di: Xie, Zichen, et al.
Pubblicazione: (2026)
di: Xie, Zichen, et al.
Pubblicazione: (2026)
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
di: Li, Yuangang, et al.
Pubblicazione: (2026)
di: Li, Yuangang, et al.
Pubblicazione: (2026)
Bidirectional Empowerment of Metamorphic Testing and Large Language Models: A Systematic Survey
di: Zheng, Zheng, et al.
Pubblicazione: (2026)
di: Zheng, Zheng, et al.
Pubblicazione: (2026)
Evaluating the Use of LLMs for Documentation to Code Traceability
di: Alor, Ebube, et al.
Pubblicazione: (2025)
di: Alor, Ebube, et al.
Pubblicazione: (2025)
EGSS: Entropy-guided Stepwise Scaling for Reliable Software Engineering
di: Mao, Chenhui, et al.
Pubblicazione: (2026)
di: Mao, Chenhui, et al.
Pubblicazione: (2026)
StackSight: Unveiling WebAssembly through Large Language Models and Neurosymbolic Chain-of-Thought Decompilation
di: Fang, Weike, et al.
Pubblicazione: (2024)
di: Fang, Weike, et al.
Pubblicazione: (2024)
CLOVER: A Test Case Generation Benchmark with Coverage, Long-Context, and Verification
di: Xu, Jiacheng, et al.
Pubblicazione: (2025)
di: Xu, Jiacheng, et al.
Pubblicazione: (2025)
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
di: Vulićević, Jelena Ilić
Pubblicazione: (2026)
di: Vulićević, Jelena Ilić
Pubblicazione: (2026)
SnipGen: A Mining Repository Framework for Evaluating LLMs for Code
di: Rodriguez-Cardenas, Daniel, et al.
Pubblicazione: (2025)
di: Rodriguez-Cardenas, Daniel, et al.
Pubblicazione: (2025)
Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Generation: An Empirical Study
di: Storhaug, André, et al.
Pubblicazione: (2024)
di: Storhaug, André, et al.
Pubblicazione: (2024)
MLAD: A Unified Model for Multi-system Log Anomaly Detection
di: Zang, Runqiang, et al.
Pubblicazione: (2024)
di: Zang, Runqiang, et al.
Pubblicazione: (2024)
Generative AI to Generate Test Data Generators
di: Baudry, Benoit, et al.
Pubblicazione: (2024)
di: Baudry, Benoit, et al.
Pubblicazione: (2024)
The Dual-State Architecture for Reliable LLM Agents
di: Thompson, Matthew
Pubblicazione: (2025)
di: Thompson, Matthew
Pubblicazione: (2025)
Exploring Pass-Rate Reward in Reinforcement Learning for Code Generation
di: Li, Xin-Ye, et al.
Pubblicazione: (2026)
di: Li, Xin-Ye, et al.
Pubblicazione: (2026)
MIST-RL: Mutation-based Incremental Suite Testing via Reinforcement Learning
di: Zhu, Sicheng, et al.
Pubblicazione: (2026)
di: Zhu, Sicheng, et al.
Pubblicazione: (2026)
PSALM: applying Proportional SAmpLing strategy in Metamorphic testing
di: Zhou, Zenghui, et al.
Pubblicazione: (2025)
di: Zhou, Zenghui, et al.
Pubblicazione: (2025)
LogReasoner: Empowering LLMs with Expert-like Coarse-to-Fine Reasoning for Automated Log Analysis
di: Ma, Lipeng, et al.
Pubblicazione: (2025)
di: Ma, Lipeng, et al.
Pubblicazione: (2025)
EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
di: Sharma, Aman, et al.
Pubblicazione: (2026)
di: Sharma, Aman, et al.
Pubblicazione: (2026)
Validating LLM-Generated Programs with Metamorphic Prompt Testing
di: Wang, Xiaoyin, et al.
Pubblicazione: (2024)
di: Wang, Xiaoyin, et al.
Pubblicazione: (2024)
ASSURE: Metamorphic Testing for AI-powered Browser Extensions
di: Gao, Xuanqi, et al.
Pubblicazione: (2025)
di: Gao, Xuanqi, et al.
Pubblicazione: (2025)
VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation
di: Bai, Yifan, et al.
Pubblicazione: (2026)
di: Bai, Yifan, et al.
Pubblicazione: (2026)
Semantic Voting: Execution-Grounded Consensus for LLM Code Generation
di: Jiang, Shan, et al.
Pubblicazione: (2026)
di: Jiang, Shan, et al.
Pubblicazione: (2026)
Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy
di: Dong, Yihong, et al.
Pubblicazione: (2026)
di: Dong, Yihong, et al.
Pubblicazione: (2026)
Advancing Software Security and Reliability in Cloud Platforms through AI-based Anomaly Detection
di: Saleh, Sabbir M., et al.
Pubblicazione: (2024)
di: Saleh, Sabbir M., et al.
Pubblicazione: (2024)
Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents
di: Huang, Jiawei, et al.
Pubblicazione: (2026)
di: Huang, Jiawei, et al.
Pubblicazione: (2026)
Can Search-Based Testing with Pareto Optimization Effectively Cover Failure-Revealing Test Inputs?
di: Sorokin, Lev, et al.
Pubblicazione: (2024)
di: Sorokin, Lev, et al.
Pubblicazione: (2024)
Operational Robustness of LLMs on Code Generation
di: Paul, Debalina Ghosh, et al.
Pubblicazione: (2026)
di: Paul, Debalina Ghosh, et al.
Pubblicazione: (2026)
On LLMs' Internal Representation of Code Correctness
di: Ribeiro, Francisco, et al.
Pubblicazione: (2025)
di: Ribeiro, Francisco, et al.
Pubblicazione: (2025)
Metamorphic Testing of Large Language Models for Natural Language Processing
di: Cho, Steven, et al.
Pubblicazione: (2025)
di: Cho, Steven, et al.
Pubblicazione: (2025)
FlakyFix: Using Large Language Models for Predicting Flaky Test Fix Categories and Test Code Repair
di: Fatima, Sakina, et al.
Pubblicazione: (2023)
di: Fatima, Sakina, et al.
Pubblicazione: (2023)
SPELL: Synthesis of Programmatic Edits using LLMs
di: Ramos, Daniel, et al.
Pubblicazione: (2026)
di: Ramos, Daniel, et al.
Pubblicazione: (2026)
Co-Located Tests, Better AI Code: How Test Syntax Structure Affects Foundation Model Code Generation
di: Jacopin, Éric
Pubblicazione: (2026)
di: Jacopin, Éric
Pubblicazione: (2026)
Understanding LLM-Driven Test Oracle Generation
di: Bodicoat, Adam, et al.
Pubblicazione: (2026)
di: Bodicoat, Adam, et al.
Pubblicazione: (2026)
Code Generation by Differential Test Time Scaling
di: He, Yifeng, et al.
Pubblicazione: (2026)
di: He, Yifeng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
MASTEST: A LLM-Based Multi-Agent System For RESTful API Tests
di: Han, Xiaoke, et al.
Pubblicazione: (2025) -
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
di: Gao, Shuzheng, et al.
Pubblicazione: (2025) -
LiCoEval: Evaluating LLMs on License Compliance in Code Generation
di: Xu, Weiwei, et al.
Pubblicazione: (2024) -
Evaluating Human Trajectory Prediction with Metamorphic Testing
di: Spieker, Helge, et al.
Pubblicazione: (2024) -
Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks
di: Zhou, Yongxi, et al.
Pubblicazione: (2026)