Mind the Gap: Evaluating Model- and Agentic-Level Vulnerabilities in LLMs with Action Graphs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wicaksono, Ilham, Wu, Zekun, Patel, Rahul, King, Theo, Koshiyama, Adriano, Treleaven, Philip |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Mind the Gap: Comparing Model- vs Agentic-Level Red Teaming with Action-Graph Observability on GPT-OSS-20B
par: Wicaksono, Ilham, et autres
Publié: (2025)
par: Wicaksono, Ilham, et autres
Publié: (2025)
HEARTS: A Holistic Framework for Explainable, Sustainable and Robust Text Stereotype Detection
par: King, Theo, et autres
Publié: (2024)
par: King, Theo, et autres
Publié: (2024)
Personality as a Probe for LLM Evaluation: Method Trade-offs and Downstream Effects
par: Handa, Gunmay, et autres
Publié: (2025)
par: Handa, Gunmay, et autres
Publié: (2025)
Knowledge Collapse in LLMs: When Fluency Survives but Facts Fail under Recursive Synthetic Training
par: Keisha, Figarri, et autres
Publié: (2025)
par: Keisha, Figarri, et autres
Publié: (2025)
THaMES: An End-to-End Tool for Hallucination Mitigation and Evaluation in Large Language Models
par: Liang, Mengfei, et autres
Publié: (2024)
par: Liang, Mengfei, et autres
Publié: (2024)
From Text to Emoji: How PEFT-Driven Personality Manipulation Unleashes the Emoji Potential in LLMs
par: Jain, Navya, et autres
Publié: (2024)
par: Jain, Navya, et autres
Publié: (2024)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
par: Cho, Seonglae, et autres
Publié: (2026)
par: Cho, Seonglae, et autres
Publié: (2026)
Stereotype Detection in LLMs: A Multiclass, Explainable, and Benchmark-Driven Approach
par: Wu, Zekun, et autres
Publié: (2024)
par: Wu, Zekun, et autres
Publié: (2024)
LibVulnWatch: A Deep Assessment Agent System and Leaderboard for Uncovering Hidden Vulnerabilities in Open-Source AI Libraries
par: Wu, Zekun, et autres
Publié: (2025)
par: Wu, Zekun, et autres
Publié: (2025)
Eliciting Personality Traits in Large Language Models
par: Hilliard, Airlie, et autres
Publié: (2024)
par: Hilliard, Airlie, et autres
Publié: (2024)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
par: Cho, Seonglae, et autres
Publié: (2025)
par: Cho, Seonglae, et autres
Publié: (2025)
The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models
par: Cho, Seonglae, et autres
Publié: (2026)
par: Cho, Seonglae, et autres
Publié: (2026)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
par: Demchak, Nathaniel, et autres
Publié: (2024)
par: Demchak, Nathaniel, et autres
Publié: (2024)
MPF: Aligning and Debiasing Language Models post Deployment via Multi Perspective Fusion
par: Guan, Xin, et autres
Publié: (2025)
par: Guan, Xin, et autres
Publié: (2025)
Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values?
par: Shen, Hua, et autres
Publié: (2025)
par: Shen, Hua, et autres
Publié: (2025)
Tool Calling is Linearly Readable and Steerable in Language Models
par: Wu, Zekun, et autres
Publié: (2026)
par: Wu, Zekun, et autres
Publié: (2026)
BERT vs GPT for financial engineering
par: Sharkey, Edward, et autres
Publié: (2024)
par: Sharkey, Edward, et autres
Publié: (2024)
Agentic Medical Knowledge Graphs Enhance Medical Question Answering: Bridging the Gap Between LLMs and Evolving Medical Knowledge
par: Rezaei, Mohammad Reza, et autres
Publié: (2025)
par: Rezaei, Mohammad Reza, et autres
Publié: (2025)
JobFair: A Framework for Benchmarking Gender Hiring Bias in Large Language Models
par: Wang, Ze, et autres
Publié: (2024)
par: Wang, Ze, et autres
Publié: (2024)
Cultural Alignment in Large Language Models Using Soft Prompt Tuning
par: Masoud, Reem I., et autres
Publié: (2025)
par: Masoud, Reem I., et autres
Publié: (2025)
Bridging the Knowledge-Action Gap by Evaluating LLMs in Dynamic Dental Clinical Scenarios
par: Ma, Hongyang, et autres
Publié: (2026)
par: Ma, Hongyang, et autres
Publié: (2026)
HyPA-RAG: A Hybrid Parameter Adaptive Retrieval-Augmented Generation System for AI Legal and Policy Applications
par: Kalra, Rishi, et autres
Publié: (2024)
par: Kalra, Rishi, et autres
Publié: (2024)
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
par: Pathade, Chetan
Publié: (2025)
par: Pathade, Chetan
Publié: (2025)
Agentic LLMs for Question Answering over Tabular Data
par: Tyagi, Rishit, et autres
Publié: (2025)
par: Tyagi, Rishit, et autres
Publié: (2025)
Mind the (Belief) Gap: Group Identity in the World of LLMs
par: Borah, Angana, et autres
Publié: (2025)
par: Borah, Angana, et autres
Publié: (2025)
SAGED: A Holistic Bias-Benchmarking Pipeline for Language Models with Customisable Fairness Calibration
par: Guan, Xin, et autres
Publié: (2024)
par: Guan, Xin, et autres
Publié: (2024)
Mind the Language Gap: Automated and Augmented Evaluation of Bias in LLMs for High- and Low-Resource Languages
par: Buscemi, Alessio, et autres
Publié: (2025)
par: Buscemi, Alessio, et autres
Publié: (2025)
Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection
par: Hakim, Muhammad Alif Al, et autres
Publié: (2026)
par: Hakim, Muhammad Alif Al, et autres
Publié: (2026)
Mind the Gap! Static and Interactive Evaluations of Large Audio Models
par: Li, Minzhi, et autres
Publié: (2025)
par: Li, Minzhi, et autres
Publié: (2025)
Mind the Gap: A Closer Look at Tokenization for Multiple-Choice Question Answering with LLMs
par: Sanz-Guerrero, Mario, et autres
Publié: (2025)
par: Sanz-Guerrero, Mario, et autres
Publié: (2025)
Mind the Style Gap: Meta-Evaluation of Style and Attribute Transfer Metrics
par: Pauli, Amalie Brogaard, et autres
Publié: (2025)
par: Pauli, Amalie Brogaard, et autres
Publié: (2025)
Mind the Gap Between Spatial Reasoning and Acting! Step-by-Step Evaluation of Agents With Spatial-Gym
par: Kaesberg, Lars Benedikt, et autres
Publié: (2026)
par: Kaesberg, Lars Benedikt, et autres
Publié: (2026)
Cultural Alignment in Large Language Models: An Explanatory Analysis Based on Hofstede's Cultural Dimensions
par: Masoud, Reem I., et autres
Publié: (2023)
par: Masoud, Reem I., et autres
Publié: (2023)
Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results
par: Peter, Jan-Thorsten, et autres
Publié: (2025)
par: Peter, Jan-Thorsten, et autres
Publié: (2025)
IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages
par: Azmi, Muhammad Falensi, et autres
Publié: (2025)
par: Azmi, Muhammad Falensi, et autres
Publié: (2025)
GraphPlanner: Graph Memory-Augmented Agentic Routing for Multi-Agent LLMs
par: Feng, Tao, et autres
Publié: (2026)
par: Feng, Tao, et autres
Publié: (2026)
MemRepair: Hierarchical Memory for Agentic Repository-Level Vulnerability Repair
par: Liu, Simiao, et autres
Publié: (2026)
par: Liu, Simiao, et autres
Publié: (2026)
Mapping the Minds of LLMs: A Graph-Based Analysis of Reasoning LLM
par: Xiong, Zhen, et autres
Publié: (2025)
par: Xiong, Zhen, et autres
Publié: (2025)
Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge
par: Gou, Boyu, et autres
Publié: (2025)
par: Gou, Boyu, et autres
Publié: (2025)
Your Agent is More Brittle Than You Think: Uncovering Indirect Injection Vulnerabilities in Agentic LLMs
par: Zhu, Wenhui, et autres
Publié: (2026)
par: Zhu, Wenhui, et autres
Publié: (2026)
Documents similaires
-
Mind the Gap: Comparing Model- vs Agentic-Level Red Teaming with Action-Graph Observability on GPT-OSS-20B
par: Wicaksono, Ilham, et autres
Publié: (2025) -
HEARTS: A Holistic Framework for Explainable, Sustainable and Robust Text Stereotype Detection
par: King, Theo, et autres
Publié: (2024) -
Personality as a Probe for LLM Evaluation: Method Trade-offs and Downstream Effects
par: Handa, Gunmay, et autres
Publié: (2025) -
Knowledge Collapse in LLMs: When Fluency Survives but Facts Fail under Recursive Synthetic Training
par: Keisha, Figarri, et autres
Publié: (2025) -
THaMES: An End-to-End Tool for Hallucination Mitigation and Evaluation in Large Language Models
par: Liang, Mengfei, et autres
Publié: (2024)