Bias Testing and Mitigation in Black Box LLMs using Metamorphic Relations
Fuente:
arXiv
Saved in:
| Main Authors: | Salimian, Sina, Uddin, Gias, Biswas, Sumon, Leung, Henry |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PCS: Perceived Confidence Scoring of Black Box LLMs with Metamorphic Relations
by: Salimian, Sina, et al.
Published: (2025)
by: Salimian, Sina, et al.
Published: (2025)
LLM Assisted Coding with Metamorphic Specification Mutation Agent
by: Akhond, Mostafijur Rahman, et al.
Published: (2025)
by: Akhond, Mostafijur Rahman, et al.
Published: (2025)
Efficient Fairness Testing in Large Language Models: Prioritizing Metamorphic Relations for Bias Detection
by: Giramata, Suavis, et al.
Published: (2025)
by: Giramata, Suavis, et al.
Published: (2025)
MORTAR: Multi-turn Metamorphic Testing for LLM-based Dialogue Systems
by: Guo, Guoxiang, et al.
Published: (2024)
by: Guo, Guoxiang, et al.
Published: (2024)
Effective Black Box Testing of Sentiment Analysis Classification Networks
by: Karbasizadeh, Parsa, et al.
Published: (2024)
by: Karbasizadeh, Parsa, et al.
Published: (2024)
Optimized Log Parsing with Syntactic Modifications
by: Enan, Nafid, et al.
Published: (2025)
by: Enan, Nafid, et al.
Published: (2025)
Metamorphic Testing for Audio Content Moderation Software
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
What Breaks When LLMs Code? Characterizing Operational Safety Failures of Agentic Code Assistants
by: Hasan, Alif Al, et al.
Published: (2026)
by: Hasan, Alif Al, et al.
Published: (2026)
From Output to Evaluation: Does Raw Instruction-Tuned Code LLMs Output Suffice for Fill-in-the-Middle Code Generation?
by: Ahmad, Wasi Uddin, et al.
Published: (2025)
by: Ahmad, Wasi Uddin, et al.
Published: (2025)
Metamorphic Debugging for Accountable Software
by: Tizpaz-Niari, Saeid, et al.
Published: (2024)
by: Tizpaz-Niari, Saeid, et al.
Published: (2024)
Secret Leak Detection in Software Issue Reports using LLMs: A Comprehensive Evaluation
by: Ahmed, Sadif, et al.
Published: (2024)
by: Ahmed, Sadif, et al.
Published: (2024)
FairCoder: Evaluating Social Bias of LLMs in Code Generation
by: Du, Yongkang, et al.
Published: (2025)
by: Du, Yongkang, et al.
Published: (2025)
On the Potential and Limitations of Few-Shot In-Context Learning to Generate Metamorphic Specifications for Tax Preparation Software
by: Srinivas, Dananjay, et al.
Published: (2023)
by: Srinivas, Dananjay, et al.
Published: (2023)
"How do people decide?": A Model for Software Library Selection
by: Tanzil, Minaoar Hossain, et al.
Published: (2024)
by: Tanzil, Minaoar Hossain, et al.
Published: (2024)
PAGENT: Learning to Patch Software Engineering Agents
by: Xue, Haoran, et al.
Published: (2025)
by: Xue, Haoran, et al.
Published: (2025)
Assessing the Influence of Toxic and Gender Discriminatory Communication on Perceptible Diversity in OSS Projects
by: Sultana, Sayma, et al.
Published: (2024)
by: Sultana, Sayma, et al.
Published: (2024)
ChatGPT Inaccuracy Mitigation during Technical Report Understanding: Are We There Yet?
by: Tamanna, Salma Begum, et al.
Published: (2024)
by: Tamanna, Salma Begum, et al.
Published: (2024)
OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs
by: Ahmad, Wasi Uddin, et al.
Published: (2025)
by: Ahmad, Wasi Uddin, et al.
Published: (2025)
TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generation
by: Liu, Steven, et al.
Published: (2026)
by: Liu, Steven, et al.
Published: (2026)
Benchmarking LLMs for Unit Test Generation from Real-World Functions
by: Huang, Dong, et al.
Published: (2025)
by: Huang, Dong, et al.
Published: (2025)
SimCT: A Simple Consistency Test Protocol in LLMs Development Lifecycle
by: Zhao, Fufangchen, et al.
Published: (2024)
by: Zhao, Fufangchen, et al.
Published: (2024)
A Mixed Method Study of DevOps Challenges
by: Tanzil, Minaoar Hossain, et al.
Published: (2024)
by: Tanzil, Minaoar Hossain, et al.
Published: (2024)
Evaluating the Environmental Impact of using SLMs and Prompt Engineering for Code Generation
by: Mamun, Md Afif Al, et al.
Published: (2026)
by: Mamun, Md Afif Al, et al.
Published: (2026)
Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure
by: Yang, Zheyuan, et al.
Published: (2025)
by: Yang, Zheyuan, et al.
Published: (2025)
Dealing with Data for RE: Mitigating Challenges while using NLP and Generative AI
by: Ghaisas, Smita, et al.
Published: (2024)
by: Ghaisas, Smita, et al.
Published: (2024)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
by: Aleithan, Reem, et al.
Published: (2024)
by: Aleithan, Reem, et al.
Published: (2024)
DiffuTester: Accelerating Unit Test Generation for Diffusion LLMs via Mining Structural Pattern
by: Yang, Lekang, et al.
Published: (2025)
by: Yang, Lekang, et al.
Published: (2025)
An Empirical Study on Bug Severity Estimation using Source Code Metrics and Static Analysis
by: Mashhadi, Ehsan, et al.
Published: (2022)
by: Mashhadi, Ehsan, et al.
Published: (2022)
Mitigating Gender Bias in Code Large Language Models via Model Editing
by: Qin, Zhanyue, et al.
Published: (2024)
by: Qin, Zhanyue, et al.
Published: (2024)
ABTest: Behavior-Driven Testing for AI Coding Agents
by: Dai, Wuyang, et al.
Published: (2026)
by: Dai, Wuyang, et al.
Published: (2026)
Reputation Gaming in Stack Overflow
by: Mazloomzadeh, Iren, et al.
Published: (2021)
by: Mazloomzadeh, Iren, et al.
Published: (2021)
BLAgent: Agentic RAG for File-Level Bug Localization
by: Mamun, Md Afif Al, et al.
Published: (2026)
by: Mamun, Md Afif Al, et al.
Published: (2026)
From SWE-ZERO to SWE-HERO: Execution-free to Execution-based Fine-tuning for Software Engineering Agents
by: Ludwig, Nikolai, et al.
Published: (2026)
by: Ludwig, Nikolai, et al.
Published: (2026)
ScaleBox: Enabling High-Fidelity and Scalable Code Verification for Large Language Models
by: Zheng, Jiasheng, et al.
Published: (2026)
by: Zheng, Jiasheng, et al.
Published: (2026)
Incremental Context-free Grammar Inference in Black Box Settings
by: Li, Feifei, et al.
Published: (2024)
by: Li, Feifei, et al.
Published: (2024)
From If-Statements to ML Pipelines: Revisiting Bias in Code-Generation
by: Bui, Minh Duc, et al.
Published: (2026)
by: Bui, Minh Duc, et al.
Published: (2026)
Repoformer: Selective Retrieval for Repository-Level Code Completion
by: Wu, Di, et al.
Published: (2024)
by: Wu, Di, et al.
Published: (2024)
Optimizing Metamorphic Testing: Prioritizing Relations Through Execution Profile Dissimilarity
by: Srinivasan, Madhusudan, et al.
Published: (2024)
by: Srinivasan, Madhusudan, et al.
Published: (2024)
LLM For Loop Invariant Generation and Fixing: How Far Are We?
by: Akhond, Mostafijur Rahman, et al.
Published: (2025)
by: Akhond, Mostafijur Rahman, et al.
Published: (2025)
MR-Scout: Automated Synthesis of Metamorphic Relations from Existing Test Cases
by: Xu, Congying, et al.
Published: (2023)
by: Xu, Congying, et al.
Published: (2023)
Similar Items
-
PCS: Perceived Confidence Scoring of Black Box LLMs with Metamorphic Relations
by: Salimian, Sina, et al.
Published: (2025) -
LLM Assisted Coding with Metamorphic Specification Mutation Agent
by: Akhond, Mostafijur Rahman, et al.
Published: (2025) -
Efficient Fairness Testing in Large Language Models: Prioritizing Metamorphic Relations for Bias Detection
by: Giramata, Suavis, et al.
Published: (2025) -
MORTAR: Multi-turn Metamorphic Testing for LLM-based Dialogue Systems
by: Guo, Guoxiang, et al.
Published: (2024) -
Effective Black Box Testing of Sentiment Analysis Classification Networks
by: Karbasizadeh, Parsa, et al.
Published: (2024)