SoK: Measuring What Matters for Closed-Loop Security Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Khurana, Mudita, Jain, Raunak |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Collaborative Causal Sensemaking: Closing the Complementarity Gap in Human-AI Decision Support
by: Jain, Raunak
Published: (2025)
by: Jain, Raunak
Published: (2025)
From Sycophancy to Sensemaking: Premise Governance for Human-AI Decision Making
by: Jain, Raunak
Published: (2026)
by: Jain, Raunak
Published: (2026)
SoK: Prompt Hacking of Large Language Models
by: Rababah, Baha, et al.
Published: (2024)
by: Rababah, Baha, et al.
Published: (2024)
SoK: Large Language Model Copyright Auditing via Fingerprinting
by: Shao, Shuo, et al.
Published: (2025)
by: Shao, Shuo, et al.
Published: (2025)
SwissNYF: Tool Grounded LLM Agents for Black Box Setting
by: Kumar, Somnath Sendhil, et al.
Published: (2024)
by: Kumar, Somnath Sendhil, et al.
Published: (2024)
SoK: Security and Privacy of AI Agents for Blockchain
by: Romandini, Nicolò, et al.
Published: (2025)
by: Romandini, Nicolò, et al.
Published: (2025)
SoK: Agentic Retrieval-Augmented Generation (RAG): Taxonomy, Architectures, Evaluation, and Research Directions
by: Mishra, Saroj, et al.
Published: (2026)
by: Mishra, Saroj, et al.
Published: (2026)
SoK: Towards Security and Safety of Edge AI
by: Wingarz, Tatjana, et al.
Published: (2024)
by: Wingarz, Tatjana, et al.
Published: (2024)
SoK: On the Offensive Potential of AI
by: Schröer, Saskia Laura, et al.
Published: (2024)
by: Schröer, Saskia Laura, et al.
Published: (2024)
SoK: On the Semantic AI Security in Autonomous Driving
by: Shen, Junjie, et al.
Published: (2022)
by: Shen, Junjie, et al.
Published: (2022)
SLIDE: Reference-free Evaluation for Machine Translation using a Sliding Document Window
by: Raunak, Vikas, et al.
Published: (2023)
by: Raunak, Vikas, et al.
Published: (2023)
SoK: Security and Privacy Risks of Healthcare AI
by: Chang, Yuanhaur, et al.
Published: (2024)
by: Chang, Yuanhaur, et al.
Published: (2024)
Security in LLM-as-a-Judge: A Comprehensive SoK
by: Masoud, Aiman Al, et al.
Published: (2026)
by: Masoud, Aiman Al, et al.
Published: (2026)
On Instruction-Finetuning Neural Machine Translation Models
by: Raunak, Vikas, et al.
Published: (2024)
by: Raunak, Vikas, et al.
Published: (2024)
SoK: Agentic Skills -- Beyond Tool Use in LLM Agents
by: Jiang, Yanna, et al.
Published: (2026)
by: Jiang, Yanna, et al.
Published: (2026)
Measuring What Matters -- or What's Convenient?: Robustness of LLM-Based Scoring Systems to Construct-Irrelevant Factors
by: Walsh, Cole, et al.
Published: (2026)
by: Walsh, Cole, et al.
Published: (2026)
SoK: Taxonomy and Evaluation of Prompt Security in Large Language Models
by: Hong, Hanbin, et al.
Published: (2025)
by: Hong, Hanbin, et al.
Published: (2025)
SoK: Trust-Authorization Mismatch in LLM Agent Interactions
by: Shi, Guanquan, et al.
Published: (2025)
by: Shi, Guanquan, et al.
Published: (2025)
Pay Attention to What Matters
by: Silva, Pedro Luiz, et al.
Published: (2024)
by: Silva, Pedro Luiz, et al.
Published: (2024)
LLM for SoC Security: A Paradigm Shift
by: Saha, Dipayan, et al.
Published: (2023)
by: Saha, Dipayan, et al.
Published: (2023)
InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification
by: InternAgent Team, et al.
Published: (2025)
by: InternAgent Team, et al.
Published: (2025)
LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
by: Zhang, Kangning, et al.
Published: (2025)
by: Zhang, Kangning, et al.
Published: (2025)
SoK: Pitfalls in Evaluating Black-Box Attacks
by: Suya, Fnu, et al.
Published: (2023)
by: Suya, Fnu, et al.
Published: (2023)
EndoCogniAgent: Closed-Loop Agentic Reasoning with Self-Consistency Validation for Endoscopic Diagnosis
by: Tang, Yi, et al.
Published: (2025)
by: Tang, Yi, et al.
Published: (2025)
Closing the Data Loop: Using OpenDataArena to Engineer Superior Training Datasets
by: Gao, Xin, et al.
Published: (2025)
by: Gao, Xin, et al.
Published: (2025)
Way to Specialist: Closing Loop Between Specialized LLM and Evolving Domain Knowledge Graph
by: Zhang, Yutong, et al.
Published: (2024)
by: Zhang, Yutong, et al.
Published: (2024)
Topology Matters: Measuring Memory Leakage in Multi-Agent LLMs
by: Liu, Jinbo, et al.
Published: (2025)
by: Liu, Jinbo, et al.
Published: (2025)
CyberCorrect: A Cybernetic Framework for Closed-Loop Self-Correction in Large Language Models
by: Wu, Yuning, et al.
Published: (2026)
by: Wu, Yuning, et al.
Published: (2026)
SoK: On Gradient Leakage in Federated Learning
by: Du, Jiacheng, et al.
Published: (2024)
by: Du, Jiacheng, et al.
Published: (2024)
SoK: a Comprehensive Causality Analysis Framework for Large Language Model Security
by: Zhao, Wei, et al.
Published: (2025)
by: Zhao, Wei, et al.
Published: (2025)
SoK: Machine Learning for Misinformation Detection
by: Xiao, Madelyne, et al.
Published: (2023)
by: Xiao, Madelyne, et al.
Published: (2023)
SoK: Are Watermarks in LLMs Ready for Deployment?
by: Dang, Kieu, et al.
Published: (2025)
by: Dang, Kieu, et al.
Published: (2025)
Middo: Model-Informed Dynamic Data Optimization for Enhanced LLM Fine-Tuning via Closed-Loop Learning
by: Tang, Zinan, et al.
Published: (2025)
by: Tang, Zinan, et al.
Published: (2025)
What Matters For Safety Alignment?
by: Li, Xing, et al.
Published: (2026)
by: Li, Xing, et al.
Published: (2026)
Every Answer Matters: Evaluating Commonsense with Probabilistic Measures
by: Cheng, Qi, et al.
Published: (2024)
by: Cheng, Qi, et al.
Published: (2024)
SIPDO: Closed-Loop Prompt Optimization via Synthetic Data Feedback
by: Yu, Yaoning, et al.
Published: (2025)
by: Yu, Yaoning, et al.
Published: (2025)
ChatCLIDS: Simulating Persuasive AI Dialogues to Promote Closed-Loop Insulin Adoption in Type 1 Diabetes Care
by: Yao, Zonghai, et al.
Published: (2025)
by: Yao, Zonghai, et al.
Published: (2025)
Evolving and Executing Research Plans via Double-Loop Multi-Agent Collaboration
by: Zhang, Zhi, et al.
Published: (2025)
by: Zhang, Zhi, et al.
Published: (2025)
SoK: Verifiable Cross-Silo FL
by: Korneev, Aleksei, et al.
Published: (2024)
by: Korneev, Aleksei, et al.
Published: (2024)
SoK: Watermarking for AI-Generated Content
by: Zhao, Xuandong, et al.
Published: (2024)
by: Zhao, Xuandong, et al.
Published: (2024)
Similar Items
-
Collaborative Causal Sensemaking: Closing the Complementarity Gap in Human-AI Decision Support
by: Jain, Raunak
Published: (2025) -
From Sycophancy to Sensemaking: Premise Governance for Human-AI Decision Making
by: Jain, Raunak
Published: (2026) -
SoK: Prompt Hacking of Large Language Models
by: Rababah, Baha, et al.
Published: (2024) -
SoK: Large Language Model Copyright Auditing via Fingerprinting
by: Shao, Shuo, et al.
Published: (2025) -
SwissNYF: Tool Grounded LLM Agents for Black Box Setting
by: Kumar, Somnath Sendhil, et al.
Published: (2024)