Evaluating Human Trajectory Prediction with Metamorphic Testing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Spieker, Helge, Belmecheri, Nassim, Gotlieb, Arnaud, Lazaar, Nadjib |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Metamorphic Testing of Multimodal Human Trajectory Prediction
von: Spieker, Helge, et al.
Veröffentlicht: (2025)
von: Spieker, Helge, et al.
Veröffentlicht: (2025)
Towards Trustworthy Automated Driving through Qualitative Scene Understanding and Explanations
von: Belmecheri, Nassim, et al.
Veröffentlicht: (2024)
von: Belmecheri, Nassim, et al.
Veröffentlicht: (2024)
Trustworthy Automated Driving through Qualitative Scene Understanding and Explanations
von: Belmecheri, Nassim, et al.
Veröffentlicht: (2024)
von: Belmecheri, Nassim, et al.
Veröffentlicht: (2024)
Explainable Scene Understanding with Qualitative Representations and Graph Neural Networks
von: Belmecheri, Nassim, et al.
Veröffentlicht: (2025)
von: Belmecheri, Nassim, et al.
Veröffentlicht: (2025)
Rashomon in the Streets: Explanation Ambiguity in Scene Understanding
von: Spieker, Helge, et al.
Veröffentlicht: (2025)
von: Spieker, Helge, et al.
Veröffentlicht: (2025)
Prompting for Performance: Exploring LLMs for Configuring Software
von: Spieker, Helge, et al.
Veröffentlicht: (2025)
von: Spieker, Helge, et al.
Veröffentlicht: (2025)
Constraint-Guided Test Execution Scheduling: An Experience Report at ABB Robotics
von: Gotlieb, Arnaud, et al.
Veröffentlicht: (2023)
von: Gotlieb, Arnaud, et al.
Veröffentlicht: (2023)
Testing for Fault Diversity in Reinforcement Learning
von: Mazouni, Quentin, et al.
Veröffentlicht: (2024)
von: Mazouni, Quentin, et al.
Veröffentlicht: (2024)
Policy Testing with MDPFuzz (Replicability Study)
von: Mazouni, Quentin, et al.
Veröffentlicht: (2025)
von: Mazouni, Quentin, et al.
Veröffentlicht: (2025)
Efficiently Ranking Software Variants with Minimal Benchmarks
von: Matricon, Théo, et al.
Veröffentlicht: (2025)
von: Matricon, Théo, et al.
Veröffentlicht: (2025)
Mutation‐Guided Metamorphic Testing of Optimality in AI Planning
von: Quentin Mazouni, et al.
Veröffentlicht: (2024)
von: Quentin Mazouni, et al.
Veröffentlicht: (2024)
Validating LLM-Generated Programs with Metamorphic Prompt Testing
von: Wang, Xiaoyin, et al.
Veröffentlicht: (2024)
von: Wang, Xiaoyin, et al.
Veröffentlicht: (2024)
ASSURE: Metamorphic Testing for AI-powered Browser Extensions
von: Gao, Xuanqi, et al.
Veröffentlicht: (2025)
von: Gao, Xuanqi, et al.
Veröffentlicht: (2025)
Metamorphic Testing of Large Language Models for Natural Language Processing
von: Cho, Steven, et al.
Veröffentlicht: (2025)
von: Cho, Steven, et al.
Veröffentlicht: (2025)
Metamorphic Testing of Deep Code Models: A Systematic Literature Review
von: Asgari, Ali, et al.
Veröffentlicht: (2025)
von: Asgari, Ali, et al.
Veröffentlicht: (2025)
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs
von: Zhou, Zenghui, et al.
Veröffentlicht: (2026)
von: Zhou, Zenghui, et al.
Veröffentlicht: (2026)
MR-Coupler: Automated Metamorphic Test Generation via Functional Coupling Analysis
von: Xu, Congying, et al.
Veröffentlicht: (2026)
von: Xu, Congying, et al.
Veröffentlicht: (2026)
A Metamorphic Testing Approach to Diagnosing Memorization in LLM-Based Program Repair
von: De Koning, Milan, et al.
Veröffentlicht: (2026)
von: De Koning, Milan, et al.
Veröffentlicht: (2026)
Integrating Artificial Intelligence with Human Expertise: An In-depth Analysis of ChatGPT's Capabilities in Generating Metamorphic Relations
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
Metamorphic Testing for Audio Content Moderation Software
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025)
Optimizing Ethical Risk Reduction for Medical Intelligent Systems with Constraint Programming
von: Brayé, Clotilde, et al.
Veröffentlicht: (2025)
von: Brayé, Clotilde, et al.
Veröffentlicht: (2025)
Multi-Agent Specification-based Metamorphic Testing of FMU-Based Simulations
von: Kulshreshtha, Ashir, et al.
Veröffentlicht: (2026)
von: Kulshreshtha, Ashir, et al.
Veröffentlicht: (2026)
Metamorphic Testing for Pose Estimation Systems
von: Duran, Matias, et al.
Veröffentlicht: (2025)
von: Duran, Matias, et al.
Veröffentlicht: (2025)
Efficient Fairness Testing in Large Language Models: Prioritizing Metamorphic Relations for Bias Detection
von: Giramata, Suavis, et al.
Veröffentlicht: (2025)
von: Giramata, Suavis, et al.
Veröffentlicht: (2025)
Search-based Selection of Metamorphic Relations for Optimized Robustness Testing of Large Language Models
von: Hyun, Sangwon, et al.
Veröffentlicht: (2025)
von: Hyun, Sangwon, et al.
Veröffentlicht: (2025)
Verifying Memoryless Sequential Decision-making of Large Language Models
von: Gross, Dennis, et al.
Veröffentlicht: (2025)
von: Gross, Dennis, et al.
Veröffentlicht: (2025)
Bounded PCTL Model Checking of Large Language Model Outputs
von: Gross, Dennis, et al.
Veröffentlicht: (2025)
von: Gross, Dennis, et al.
Veröffentlicht: (2025)
DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents
von: Hong, Sirui, et al.
Veröffentlicht: (2026)
von: Hong, Sirui, et al.
Veröffentlicht: (2026)
From Untestable to Testable: Metamorphic Testing in the Age of LLMs
von: Terragni, Valerio
Veröffentlicht: (2026)
von: Terragni, Valerio
Veröffentlicht: (2026)
Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex
von: Yang, Zhonghao, et al.
Veröffentlicht: (2026)
von: Yang, Zhonghao, et al.
Veröffentlicht: (2026)
TAM-Eval: Evaluating LLMs for Automated Unit Test Maintenance
von: Bruches, Elena, et al.
Veröffentlicht: (2026)
von: Bruches, Elena, et al.
Veröffentlicht: (2026)
Evaluating LLM-Based Test Generation Under Software Evolution
von: Haroon, Sabaat, et al.
Veröffentlicht: (2026)
von: Haroon, Sabaat, et al.
Veröffentlicht: (2026)
Large Language Models as Test Case Generators: Performance Evaluation and Enhancement
von: Li, Kefan, et al.
Veröffentlicht: (2024)
von: Li, Kefan, et al.
Veröffentlicht: (2024)
Towards Reliable Evaluation of Neural Program Repair with Natural Robustness Testing
von: Le-Cong, Thanh, et al.
Veröffentlicht: (2024)
von: Le-Cong, Thanh, et al.
Veröffentlicht: (2024)
Mutation-based Consistency Testing for Evaluating the Code Understanding Capability of LLMs
von: Li, Ziyu, et al.
Veröffentlicht: (2024)
von: Li, Ziyu, et al.
Veröffentlicht: (2024)
GUITestScape: Towards Open-set Evaluation on Exploratory GUI Testing
von: Chen, Xiaoyi, et al.
Veröffentlicht: (2026)
von: Chen, Xiaoyi, et al.
Veröffentlicht: (2026)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
von: Gao, Shuzheng, et al.
Veröffentlicht: (2025)
von: Gao, Shuzheng, et al.
Veröffentlicht: (2025)
LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost
von: Huang, Donghao, et al.
Veröffentlicht: (2025)
von: Huang, Donghao, et al.
Veröffentlicht: (2025)
Evaluating Large Language Models for the Generation of Unit Tests with Equivalence Partitions and Boundary Values
von: Rodríguez, Martín, et al.
Veröffentlicht: (2025)
von: Rodríguez, Martín, et al.
Veröffentlicht: (2025)
Tricky$^2$: Towards a Benchmark for Evaluating Human and LLM Error Interactions
von: Granger, Cole, et al.
Veröffentlicht: (2026)
von: Granger, Cole, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Metamorphic Testing of Multimodal Human Trajectory Prediction
von: Spieker, Helge, et al.
Veröffentlicht: (2025) -
Towards Trustworthy Automated Driving through Qualitative Scene Understanding and Explanations
von: Belmecheri, Nassim, et al.
Veröffentlicht: (2024) -
Trustworthy Automated Driving through Qualitative Scene Understanding and Explanations
von: Belmecheri, Nassim, et al.
Veröffentlicht: (2024) -
Explainable Scene Understanding with Qualitative Representations and Graph Neural Networks
von: Belmecheri, Nassim, et al.
Veröffentlicht: (2025) -
Rashomon in the Streets: Explanation Ambiguity in Scene Understanding
von: Spieker, Helge, et al.
Veröffentlicht: (2025)