TAI3: Testing Agent Integrity in Interpreting User Intent
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Shiwei, Xu, Xiangzhe, Chen, Xuan, Zhang, Kaiyuan, Ahmed, Syed Yusuf, Su, Zian, Zheng, Mingwei, Zhang, Xiangyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Position: Intelligent Coding Systems Should Write Programs with Justifications
by: Xu, Xiangzhe, et al.
Published: (2025)
by: Xu, Xiangzhe, et al.
Published: (2025)
Source Code Foundation Models are Transferable Binary Analysis Knowledge Bases
by: Su, Zian, et al.
Published: (2024)
by: Su, Zian, et al.
Published: (2024)
ProSec: Fortifying Code LLMs with Proactive Security Alignment
by: Xu, Xiangzhe, et al.
Published: (2024)
by: Xu, Xiangzhe, et al.
Published: (2024)
SpecOps: A Fully Automated AI Agent Testing Framework in Real-World GUI Environments
by: Ahmed, Syed Yusuf, et al.
Published: (2026)
by: Ahmed, Syed Yusuf, et al.
Published: (2026)
Symbol Preference Aware Generative Models for Recovering Variable Names from Stripped Binary
by: Xu, Xiangzhe, et al.
Published: (2023)
by: Xu, Xiangzhe, et al.
Published: (2023)
Identifying Adversary Tactics and Techniques in Malware Binaries with an LLM Agent
by: Xuan, Zhou, et al.
Published: (2026)
by: Xuan, Zhou, et al.
Published: (2026)
RFCAudit: An LLM Agent for Functional Bug Detection in Network Protocols
by: Zheng, Mingwei, et al.
Published: (2025)
by: Zheng, Mingwei, et al.
Published: (2025)
CodeArt: Better Code Models by Attention Regularization When Symbols Are Lacking
by: Su, Zian, et al.
Published: (2024)
by: Su, Zian, et al.
Published: (2024)
RepoAudit: An Autonomous LLM-Agent for Repository-Level Code Auditing
by: Guo, Jinyao, et al.
Published: (2025)
by: Guo, Jinyao, et al.
Published: (2025)
Extracting Protocol Format as State Machine via Controlled Static Loop Analysis
by: Shi, Qingkai, et al.
Published: (2023)
by: Shi, Qingkai, et al.
Published: (2023)
ROCAS: Root Cause Analysis of Autonomous Driving Accidents via Cyber-Physical Co-mutation
by: Feng, Shiwei, et al.
Published: (2024)
by: Feng, Shiwei, et al.
Published: (2024)
CodeMirage: A Multi-Lingual Benchmark for Detecting AI-Generated and Paraphrased Source Code from Production-Level LLMs
by: Guo, Hanxi, et al.
Published: (2025)
by: Guo, Hanxi, et al.
Published: (2025)
Validating Network Protocol Parsers with Traceable RFC Document Interpretation
by: Zheng, Mingwei, et al.
Published: (2025)
by: Zheng, Mingwei, et al.
Published: (2025)
LLM Use, Cheating, and Academic Integrity in Software Engineering Education
by: Santos, Ronnie de Souza, et al.
Published: (2026)
by: Santos, Ronnie de Souza, et al.
Published: (2026)
Large Language Models for Validating Network Protocol Parsers
by: Zheng, Mingwei, et al.
Published: (2025)
by: Zheng, Mingwei, et al.
Published: (2025)
FastFixer: An Efficient and Effective Approach for Repairing Programming Assignments
by: Liu, Fang, et al.
Published: (2024)
by: Liu, Fang, et al.
Published: (2024)
Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
by: Chen, Xuan, et al.
Published: (2026)
by: Chen, Xuan, et al.
Published: (2026)
Agent for User: Testing Multi-User Interactive Features in TikTok
by: Feng, Sidong, et al.
Published: (2025)
by: Feng, Sidong, et al.
Published: (2025)
Automated Functional Testing for Malleable Mobile Application Driven from User Intent
by: Wang, Yuying, et al.
Published: (2026)
by: Wang, Yuying, et al.
Published: (2026)
Overwhelmed software developers: An Interpretative Phenomenological Analysis
by: Michels, Lisa-Marie, et al.
Published: (2024)
by: Michels, Lisa-Marie, et al.
Published: (2024)
Paths to Testing: Why Women Enter and Remain in Software Testing?
by: Silva, Kleice, et al.
Published: (2024)
by: Silva, Kleice, et al.
Published: (2024)
SWE-chat: Coding Agent Interactions From Real Users in the Wild
by: Baumann, Joachim, et al.
Published: (2026)
by: Baumann, Joachim, et al.
Published: (2026)
ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants
by: Xu, Xiangzhe, et al.
Published: (2025)
by: Xu, Xiangzhe, et al.
Published: (2025)
Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity
by: Jewitt, James, et al.
Published: (2026)
by: Jewitt, James, et al.
Published: (2026)
When the Specification Emerges: Benchmarking Faithfulness Loss in Long-Horizon Coding Agents
by: Yan, Lu, et al.
Published: (2026)
by: Yan, Lu, et al.
Published: (2026)
The First Issue Matters: Linking Task-Level Characteristics to Long-Term Newcomer Retention in OSS
by: Hao, Yichen, et al.
Published: (2026)
by: Hao, Yichen, et al.
Published: (2026)
LLMDFA: Analyzing Dataflow in Code with Large Language Models
by: Wang, Chengpeng, et al.
Published: (2024)
by: Wang, Chengpeng, et al.
Published: (2024)
Nova: Generative Language Models for Assembly Code with Hierarchical Attention and Contrastive Learning
by: Jiang, Nan, et al.
Published: (2023)
by: Jiang, Nan, et al.
Published: (2023)
IntentCoding: Amplifying User Intent in Code Generation
by: Fang, Zheng, et al.
Published: (2026)
by: Fang, Zheng, et al.
Published: (2026)
KTester: Leveraging Domain and Testing Knowledge for More Effective LLM-based Test Generation
by: Li, Anji, et al.
Published: (2025)
by: Li, Anji, et al.
Published: (2025)
Breaking Single-Tester Limits: Multi-Agent LLMs for Multi-User Feature Testing
by: Feng, Sidong, et al.
Published: (2025)
by: Feng, Sidong, et al.
Published: (2025)
Rapid Mobile App Development for Generative AI Agents on MIT App Inventor
by: Gao, Jaida, et al.
Published: (2024)
by: Gao, Jaida, et al.
Published: (2024)
Operationalizing Ethics for AI Agents: How Developers Encode Values into Repository Context Files
by: Treude, Christoph, et al.
Published: (2026)
by: Treude, Christoph, et al.
Published: (2026)
Bridging the Gap between User Intent and LLM: A Requirement Alignment Approach for Code Generation
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
Context-Specific Instruction: A Longitudinal Study on Debugging Skill Acquisition and Retention for Novice Programmers
by: Zhang, Ziyi, et al.
Published: (2025)
by: Zhang, Ziyi, et al.
Published: (2025)
MuMuTestUp: Mutation-based Multi-Agent Test Case Update
by: Tian, Dawei, et al.
Published: (2026)
by: Tian, Dawei, et al.
Published: (2026)
AUITestAgent: Automatic Requirements Oriented GUI Function Testing
by: Hu, Yongxiang, et al.
Published: (2024)
by: Hu, Yongxiang, et al.
Published: (2024)
Efficient Incremental Code Coverage Analysis for Regression Test Suites
by: Wang, Jiale Amber, et al.
Published: (2024)
by: Wang, Jiale Amber, et al.
Published: (2024)
Open, Small, Rigmarole -- Evaluating Llama 3.2 3B's Feedback for Programming Exercises
by: Azaiz, Imen, et al.
Published: (2025)
by: Azaiz, Imen, et al.
Published: (2025)
Who is using AI to code? Global diffusion and impact of generative AI
by: Daniotti, Simone, et al.
Published: (2025)
by: Daniotti, Simone, et al.
Published: (2025)
Similar Items
-
Position: Intelligent Coding Systems Should Write Programs with Justifications
by: Xu, Xiangzhe, et al.
Published: (2025) -
Source Code Foundation Models are Transferable Binary Analysis Knowledge Bases
by: Su, Zian, et al.
Published: (2024) -
ProSec: Fortifying Code LLMs with Proactive Security Alignment
by: Xu, Xiangzhe, et al.
Published: (2024) -
SpecOps: A Fully Automated AI Agent Testing Framework in Real-World GUI Environments
by: Ahmed, Syed Yusuf, et al.
Published: (2026) -
Symbol Preference Aware Generative Models for Recovering Variable Names from Stripped Binary
by: Xu, Xiangzhe, et al.
Published: (2023)