LibVulnWatch: A Deep Assessment Agent System and Leaderboard for Uncovering Hidden Vulnerabilities in Open-Source AI Libraries
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Zekun, Cho, Seonglae, Mohammed, Umar, Munoz, Cristian, Costa, Kleyton, Guan, Xin, King, Theo, Wang, Ze, Kazim, Emre, Koshiyama, Adriano |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models
by: Cho, Seonglae, et al.
Published: (2026)
by: Cho, Seonglae, et al.
Published: (2026)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2026)
by: Cho, Seonglae, et al.
Published: (2026)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2025)
by: Cho, Seonglae, et al.
Published: (2025)
HEARTS: A Holistic Framework for Explainable, Sustainable and Robust Text Stereotype Detection
by: King, Theo, et al.
Published: (2024)
by: King, Theo, et al.
Published: (2024)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
by: Demchak, Nathaniel, et al.
Published: (2024)
by: Demchak, Nathaniel, et al.
Published: (2024)
MPF: Aligning and Debiasing Language Models post Deployment via Multi Perspective Fusion
by: Guan, Xin, et al.
Published: (2025)
by: Guan, Xin, et al.
Published: (2025)
Tool Calling is Linearly Readable and Steerable in Language Models
by: Wu, Zekun, et al.
Published: (2026)
by: Wu, Zekun, et al.
Published: (2026)
Mind the Gap: Evaluating Model- and Agentic-Level Vulnerabilities in LLMs with Action Graphs
by: Wicaksono, Ilham, et al.
Published: (2025)
by: Wicaksono, Ilham, et al.
Published: (2025)
From Text to Emoji: How PEFT-Driven Personality Manipulation Unleashes the Emoji Potential in LLMs
by: Jain, Navya, et al.
Published: (2024)
by: Jain, Navya, et al.
Published: (2024)
SAGED: A Holistic Bias-Benchmarking Pipeline for Language Models with Customisable Fairness Calibration
by: Guan, Xin, et al.
Published: (2024)
by: Guan, Xin, et al.
Published: (2024)
Evaluating Explainability in Machine Learning Predictions through Explainer-Agnostic Metrics
by: Munoz, Cristian, et al.
Published: (2023)
by: Munoz, Cristian, et al.
Published: (2023)
THaMES: An End-to-End Tool for Hallucination Mitigation and Evaluation in Large Language Models
by: Liang, Mengfei, et al.
Published: (2024)
by: Liang, Mengfei, et al.
Published: (2024)
Mind the Gap: Comparing Model- vs Agentic-Level Red Teaming with Action-Graph Observability on GPT-OSS-20B
by: Wicaksono, Ilham, et al.
Published: (2025)
by: Wicaksono, Ilham, et al.
Published: (2025)
Eliciting Personality Traits in Large Language Models
by: Hilliard, Airlie, et al.
Published: (2024)
by: Hilliard, Airlie, et al.
Published: (2024)
Knowledge Collapse in LLMs: When Fluency Survives but Facts Fail under Recursive Synthetic Training
by: Keisha, Figarri, et al.
Published: (2025)
by: Keisha, Figarri, et al.
Published: (2025)
Personality as a Probe for LLM Evaluation: Method Trade-offs and Downstream Effects
by: Handa, Gunmay, et al.
Published: (2025)
by: Handa, Gunmay, et al.
Published: (2025)
Bias Amplification: Large Language Models as Increasingly Biased Media
by: Wang, Ze, et al.
Published: (2024)
by: Wang, Ze, et al.
Published: (2024)
Stereotype Detection in LLMs: A Multiclass, Explainable, and Benchmark-Driven Approach
by: Wu, Zekun, et al.
Published: (2024)
by: Wu, Zekun, et al.
Published: (2024)
HyPA-RAG: A Hybrid Parameter Adaptive Retrieval-Augmented Generation System for AI Legal and Policy Applications
by: Kalra, Rishi, et al.
Published: (2024)
by: Kalra, Rishi, et al.
Published: (2024)
VulnScout-C: A Lightweight Transformer for C Code Vulnerability Detection
by: Lassoued, Aymen, et al.
Published: (2026)
by: Lassoued, Aymen, et al.
Published: (2026)
VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection
by: Nie, Yuzhou, et al.
Published: (2025)
by: Nie, Yuzhou, et al.
Published: (2025)
JobFair: A Framework for Benchmarking Gender Hiring Bias in Large Language Models
by: Wang, Ze, et al.
Published: (2024)
by: Wang, Ze, et al.
Published: (2024)
Patch2Vuln: Agentic Reconstruction of Vulnerabilities from Linux Distribution Binary Patches
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
VulnAgent-X: A Layered Agentic Framework for Repository-Level Vulnerability Detection
by: Meng, Renwei, et al.
Published: (2026)
by: Meng, Renwei, et al.
Published: (2026)
SpeechColab Leaderboard: An Open-Source Platform for Automatic Speech Recognition Evaluation
by: Du, Jiayu, et al.
Published: (2024)
by: Du, Jiayu, et al.
Published: (2024)
VulnResolver: A Hybrid Agent Framework for LLM-Based Automated Vulnerability Issue Resolution
by: Zhang, Mingming, et al.
Published: (2026)
by: Zhang, Mingming, et al.
Published: (2026)
VulnLLMEval: A Framework for Evaluating Large Language Models in Software Vulnerability Detection and Patching
by: Zibaeirad, Arastoo, et al.
Published: (2024)
by: Zibaeirad, Arastoo, et al.
Published: (2024)
LLM4Vuln: A Unified Evaluation Framework for Decoupling and Enhancing LLMs' Vulnerability Reasoning
by: Sun, Yuqiang, et al.
Published: (2024)
by: Sun, Yuqiang, et al.
Published: (2024)
AutoVulnPHP: LLM-Powered Two-Stage PHP Vulnerability Detection and Automated Localization
by: Wang, Zhiqiang, et al.
Published: (2026)
by: Wang, Zhiqiang, et al.
Published: (2026)
Uncovering Hidden Inclusions of Vulnerable Dependencies in Real-World Java Projects
by: Schott, Stefan, et al.
Published: (2026)
by: Schott, Stefan, et al.
Published: (2026)
Open Universal Arabic ASR Leaderboard
by: Wang, Yingzhi, et al.
Published: (2024)
by: Wang, Yingzhi, et al.
Published: (2024)
Depressive symptoms and motor performance in the elderly: a population based study
by: Kleyton T. Santos
Published: (2012)
by: Kleyton T. Santos
Published: (2012)
MulVuln: Enhancing Pre-trained LMs with Shared and Language-Specific Knowledge for Multilingual Vulnerability Detection
by: Nguyen, Van, et al.
Published: (2025)
by: Nguyen, Van, et al.
Published: (2025)
Proxion: Uncovering Hidden Proxy Smart Contracts for Finding Collision Vulnerabilities in Ethereum
by: Chen, Cheng-Kang, et al.
Published: (2024)
by: Chen, Cheng-Kang, et al.
Published: (2024)
VulnRepairEval: An Exploit-Based Evaluation Framework for Assessing Large Language Model Vulnerability Repair Capabilities
by: Wang, Weizhe, et al.
Published: (2025)
by: Wang, Weizhe, et al.
Published: (2025)
FuzzySQL: Uncovering Hidden Vulnerabilities in DBMS Special Features with LLM-Driven Fuzzing
by: Chen, Yongxin, et al.
Published: (2026)
by: Chen, Yongxin, et al.
Published: (2026)
CrossCommitVuln-Bench: A Dataset of Multi-Commit Python Vulnerabilities Invisible to Per-Commit Static Analysis
by: Majumdar, Arunabh
Published: (2026)
by: Majumdar, Arunabh
Published: (2026)
Projeto Tuning europeu para a educação superior: reflexões sobre o seu delineamento
by: Kleyton Carlos Ferreira
Published: (2017)
by: Kleyton Carlos Ferreira
Published: (2017)
The first occurrence of a freshwater percomorph fish (Actinopterygii: Teleostei) in the Ixtapa Formation (Miocene), Chiapas, southeastern Mexico
by: Kleyton M. Cantalice
Published: (2019)
by: Kleyton M. Cantalice
Published: (2019)
Biological aspects and mating behavior of Leucothyreus albopilosus (Coleoptera: Scarabaeidae)
by: Kleyton Rezende Ferreira
Published: (2016)
by: Kleyton Rezende Ferreira
Published: (2016)
Similar Items
-
The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models
by: Cho, Seonglae, et al.
Published: (2026) -
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2026) -
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2025) -
HEARTS: A Holistic Framework for Explainable, Sustainable and Robust Text Stereotype Detection
by: King, Theo, et al.
Published: (2024) -
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
by: Demchak, Nathaniel, et al.
Published: (2024)