Mind the Gap: Comparing Model- vs Agentic-Level Red Teaming with Action-Graph Observability on GPT-OSS-20B
Fuente:
arXiv
Saved in:
| Main Authors: | Wicaksono, Ilham, Wu, Zekun, Patel, Rahul, King, Theo, Koshiyama, Adriano, Treleaven, Philip |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mind the Gap: Evaluating Model- and Agentic-Level Vulnerabilities in LLMs with Action Graphs
by: Wicaksono, Ilham, et al.
Published: (2025)
by: Wicaksono, Ilham, et al.
Published: (2025)
HEARTS: A Holistic Framework for Explainable, Sustainable and Robust Text Stereotype Detection
by: King, Theo, et al.
Published: (2024)
by: King, Theo, et al.
Published: (2024)
Personality as a Probe for LLM Evaluation: Method Trade-offs and Downstream Effects
by: Handa, Gunmay, et al.
Published: (2025)
by: Handa, Gunmay, et al.
Published: (2025)
Knowledge Collapse in LLMs: When Fluency Survives but Facts Fail under Recursive Synthetic Training
by: Keisha, Figarri, et al.
Published: (2025)
by: Keisha, Figarri, et al.
Published: (2025)
BERT vs GPT for financial engineering
by: Sharkey, Edward, et al.
Published: (2024)
by: Sharkey, Edward, et al.
Published: (2024)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2026)
by: Cho, Seonglae, et al.
Published: (2026)
HyPA-RAG: A Hybrid Parameter Adaptive Retrieval-Augmented Generation System for AI Legal and Policy Applications
by: Kalra, Rishi, et al.
Published: (2024)
by: Kalra, Rishi, et al.
Published: (2024)
THaMES: An End-to-End Tool for Hallucination Mitigation and Evaluation in Large Language Models
by: Liang, Mengfei, et al.
Published: (2024)
by: Liang, Mengfei, et al.
Published: (2024)
From Text to Emoji: How PEFT-Driven Personality Manipulation Unleashes the Emoji Potential in LLMs
by: Jain, Navya, et al.
Published: (2024)
by: Jain, Navya, et al.
Published: (2024)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2025)
by: Cho, Seonglae, et al.
Published: (2025)
The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models
by: Cho, Seonglae, et al.
Published: (2026)
by: Cho, Seonglae, et al.
Published: (2026)
Eliciting Personality Traits in Large Language Models
by: Hilliard, Airlie, et al.
Published: (2024)
by: Hilliard, Airlie, et al.
Published: (2024)
AgenticRed: Evolving Agentic Systems for Red-Teaming
by: Yuan, Jiayi, et al.
Published: (2026)
by: Yuan, Jiayi, et al.
Published: (2026)
Stereotype Detection in LLMs: A Multiclass, Explainable, and Benchmark-Driven Approach
by: Wu, Zekun, et al.
Published: (2024)
by: Wu, Zekun, et al.
Published: (2024)
GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
by: Kumar, Deepak, et al.
Published: (2025)
by: Kumar, Deepak, et al.
Published: (2025)
LibVulnWatch: A Deep Assessment Agent System and Leaderboard for Uncovering Hidden Vulnerabilities in Open-Source AI Libraries
by: Wu, Zekun, et al.
Published: (2025)
by: Wu, Zekun, et al.
Published: (2025)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
by: Demchak, Nathaniel, et al.
Published: (2024)
by: Demchak, Nathaniel, et al.
Published: (2024)
An Empirical Analysis of Community and Coding Patterns in OSS4SG vs. Conventional OSS
by: Ouf, Mohamed, et al.
Published: (2026)
by: Ouf, Mohamed, et al.
Published: (2026)
Federated Computing as Code (FCaC): Sovereignty-aware Systems by Design
by: Fenoglio, Enzo, et al.
Published: (2026)
by: Fenoglio, Enzo, et al.
Published: (2026)
Quant Fever, Reasoning Blackholes, Schrodinger's Compliance, and More: Probing GPT-OSS-20B
by: Lin, Shuyi, et al.
Published: (2025)
by: Lin, Shuyi, et al.
Published: (2025)
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?
by: Chen, Shuo, et al.
Published: (2024)
by: Chen, Shuo, et al.
Published: (2024)
OpenAI's GPT-OSS-20B Model and Safety Alignment Issues in a Low-Resource Language
by: Inuwa-Dutse, Isa
Published: (2025)
by: Inuwa-Dutse, Isa
Published: (2025)
Do LLMs Know They Are Being Tested? Evaluation Awareness and Incentive-Sensitive Failures in GPT-OSS-20B
by: Ahmed, Nisar, et al.
Published: (2025)
by: Ahmed, Nisar, et al.
Published: (2025)
RedVLA: Physical Red Teaming for Vision-Language-Action Models
by: Zhang, Yuhao, et al.
Published: (2026)
by: Zhang, Yuhao, et al.
Published: (2026)
Heterogeneous Team Coordination on Partially Observable Graphs with Realistic Communication
by: Zhou, Yanlin, et al.
Published: (2024)
by: Zhou, Yanlin, et al.
Published: (2024)
RedTeamLLM: an Agentic AI framework for offensive security
by: Challita, Brian, et al.
Published: (2025)
by: Challita, Brian, et al.
Published: (2025)
MPF: Aligning and Debiasing Language Models post Deployment via Multi Perspective Fusion
by: Guan, Xin, et al.
Published: (2025)
by: Guan, Xin, et al.
Published: (2025)
Tool Calling is Linearly Readable and Steerable in Language Models
by: Wu, Zekun, et al.
Published: (2026)
by: Wu, Zekun, et al.
Published: (2026)
Strategize Globally, Adapt Locally: A Multi-Turn Red Teaming Agent with Dual-Level Learning
by: Chen, Si, et al.
Published: (2025)
by: Chen, Si, et al.
Published: (2025)
DeTEcT: Dynamic and Probabilistic Parameters Extension
by: Sadykhov, Rem, et al.
Published: (2024)
by: Sadykhov, Rem, et al.
Published: (2024)
Economic Policy Taxonomy
by: Sadykhov, Rem, et al.
Published: (2025)
by: Sadykhov, Rem, et al.
Published: (2025)
Impacts of Economic Policies on Wealth Distribution in Token Economies
by: Sadykhov, Rem, et al.
Published: (2026)
by: Sadykhov, Rem, et al.
Published: (2026)
Lagrangian Relaxation for Multi-Action Partially Observable Restless Bandits: Heuristic Policies and Indexability
by: Meshram, Rahul, et al.
Published: (2025)
by: Meshram, Rahul, et al.
Published: (2025)
Towards Bridging Language Gaps in OSS with LLM-Driven Documentation Translation
by: Adejumo, Elijah Kayode, et al.
Published: (2025)
by: Adejumo, Elijah Kayode, et al.
Published: (2025)
Red Teaming AI Red Teaming
by: Majumdar, Subhabrata, et al.
Published: (2025)
by: Majumdar, Subhabrata, et al.
Published: (2025)
Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours
by: Dheekonda, Raja Sekhar Rao, et al.
Published: (2026)
by: Dheekonda, Raja Sekhar Rao, et al.
Published: (2026)
Mind the Gap: Evaluating LLMs for High-Level Malicious Package Detection vs. Fine-Grained Indicator Identification
by: Ryan, Ahmed, et al.
Published: (2026)
by: Ryan, Ahmed, et al.
Published: (2026)
Beyond Code: Empirical Insights into How Team Dynamics Influence OSS Project Selection
by: Nirmani, Shashiwadana, et al.
Published: (2026)
by: Nirmani, Shashiwadana, et al.
Published: (2026)
Bias Amplification: Large Language Models as Increasingly Biased Media
by: Wang, Ze, et al.
Published: (2024)
by: Wang, Ze, et al.
Published: (2024)
Extending Puzzle for Mixture-of-Experts Reasoning Models with Application to GPT-OSS Acceleration
by: Bercovich, Akhiad, et al.
Published: (2026)
by: Bercovich, Akhiad, et al.
Published: (2026)
Similar Items
-
Mind the Gap: Evaluating Model- and Agentic-Level Vulnerabilities in LLMs with Action Graphs
by: Wicaksono, Ilham, et al.
Published: (2025) -
HEARTS: A Holistic Framework for Explainable, Sustainable and Robust Text Stereotype Detection
by: King, Theo, et al.
Published: (2024) -
Personality as a Probe for LLM Evaluation: Method Trade-offs and Downstream Effects
by: Handa, Gunmay, et al.
Published: (2025) -
Knowledge Collapse in LLMs: When Fluency Survives but Facts Fail under Recursive Synthetic Training
by: Keisha, Figarri, et al.
Published: (2025) -
BERT vs GPT for financial engineering
by: Sharkey, Edward, et al.
Published: (2024)