Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems
Fuente:
arXiv
Guardado en:
| Autores principales: | Cui, Tianyu, Wang, Yanling, Fu, Chuanpu, Xiao, Yong, Li, Sijia, Deng, Xinhao, Liu, Yunpeng, Zhang, Qinglin, Qiu, Ziyi, Li, Peiyang, Tan, Zhixing, Xiong, Junwu, Kong, Xinyu, Wen, Zujie, Xu, Ke, Li, Qi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
How Social is It? A Benchmark for LLMs' Capabilities in Multi-user Multi-turn Social Agent Tasks
por: Wu, Yusen, et al.
Publicado: (2025)
por: Wu, Yusen, et al.
Publicado: (2025)
Robust and Reliable Early-Stage Website Fingerprinting Attacks via Spatial-Temporal Distribution Analysis
por: Deng, Xinhao, et al.
Publicado: (2024)
por: Deng, Xinhao, et al.
Publicado: (2024)
Hummer: Towards Limited Competitive Preference Dataset
por: Jiang, Li, et al.
Publicado: (2024)
por: Jiang, Li, et al.
Publicado: (2024)
Towards Fine-Grained Webpage Fingerprinting at Scale
por: Zhao, Xiyuan, et al.
Publicado: (2024)
por: Zhao, Xiyuan, et al.
Publicado: (2024)
TrafficLLM: Enhancing Large Language Models for Network Traffic Analysis with Generic Traffic Representation
por: Cui, Tianyu, et al.
Publicado: (2025)
por: Cui, Tianyu, et al.
Publicado: (2025)
GRAB: A Risk Taxonomy--Grounded Benchmark for Unsupervised Topic Discovery in Financial Disclosures
por: Li, Ying, et al.
Publicado: (2025)
por: Li, Ying, et al.
Publicado: (2025)
Flexible ZSM‐5 Zeolite Membrane for High‐Performance Helium Separation
por: Yongyan Deng, et al.
Publicado: (2024)
por: Yongyan Deng, et al.
Publicado: (2024)
Flexible ZSM‐5 Zeolite Membrane for High‐Performance Helium Separation
por: Yongyan Deng, et al.
Publicado: (2024)
por: Yongyan Deng, et al.
Publicado: (2024)
Mapping AI Risk Mitigations: Evidence Scan and Preliminary AI Risk Mitigation Taxonomy
por: Saeri, Alexander K., et al.
Publicado: (2025)
por: Saeri, Alexander K., et al.
Publicado: (2025)
Causal Directed Acyclic Graphs to Mitigate Confounding Bias in Exposure‐Response Analyses
por: Sebastiaan C. Goulooze, et al.
Publicado: (2026)
por: Sebastiaan C. Goulooze, et al.
Publicado: (2026)
Enhancing Risk Assessment in Transformers with Loss-at-Risk Functions
por: Zhang, Jinghan, et al.
Publicado: (2024)
por: Zhang, Jinghan, et al.
Publicado: (2024)
Threat, Risk and Mitigation Taxonomy for Digital Identity Systems
por: SHEIK, AL TARIQ, et al.
Publicado: (2024)
por: SHEIK, AL TARIQ, et al.
Publicado: (2024)
Management and Assessment of Displacement Risks in the Construction of High‐Risk Structural Engineering Projects
por: Tianyu Wang, et al.
Publicado: (2026)
por: Tianyu Wang, et al.
Publicado: (2026)
Exposing LLM User Privacy via Traffic Fingerprint Analysis: A Study of Privacy Risks in LLM Agent Interactions
por: Zhang, Yixiang, et al.
Publicado: (2025)
por: Zhang, Yixiang, et al.
Publicado: (2025)
GenTS: A Comprehensive Benchmark Library for Generative Time Series Models
por: Wang, Chenxi, et al.
Publicado: (2026)
por: Wang, Chenxi, et al.
Publicado: (2026)
Mitigating Financial Frictions in Agriculture: A Framework for Stablecoin Adoption
por: Li, Xinyu
Publicado: (2025)
por: Li, Xinyu
Publicado: (2025)
A Visualization Verification Method for Startup Risk of Substation Equipment Based on Multidimensional Graph Models
por: Lifan Mao, et al.
Publicado: (2025)
por: Lifan Mao, et al.
Publicado: (2025)
Hallucination by Code Generation LLMs: Taxonomy, Benchmarks, Mitigation, and Challenges
por: Lee, Yunseo, et al.
Publicado: (2025)
por: Lee, Yunseo, et al.
Publicado: (2025)
A Hard-Label Black-Box Evasion Attack against ML-based Malicious Traffic Detection Systems
por: Liu, Zixuan, et al.
Publicado: (2025)
por: Liu, Zixuan, et al.
Publicado: (2025)
Boosting Medical Image Synthesis via Registration-guided Consistency and Disentanglement Learning
por: Li, Chuanpu, et al.
Publicado: (2024)
por: Li, Chuanpu, et al.
Publicado: (2024)
Financial Openness, Bank Systematic Risk, and Macroprudential Supervision
por: Yanling Chen, et al.
Publicado: (2024)
por: Yanling Chen, et al.
Publicado: (2024)
When MCP Servers Attack: Taxonomy, Feasibility, and Mitigation
por: Zhao, Weibo, et al.
Publicado: (2025)
por: Zhao, Weibo, et al.
Publicado: (2025)
Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling
por: Xiong, Shengwu., et al.
Publicado: (2025)
por: Xiong, Shengwu., et al.
Publicado: (2025)
Quantitative Risk Assessment for Autonomous Vehicles: Integrating Functional Resonance Analysis Method and Bayesian Network
por: Chengwen Deng, et al.
Publicado: (2024)
por: Chengwen Deng, et al.
Publicado: (2024)
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
por: Li, Miles Q., et al.
Publicado: (2026)
por: Li, Miles Q., et al.
Publicado: (2026)
Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings
por: Yang, Shujian, et al.
Publicado: (2025)
por: Yang, Shujian, et al.
Publicado: (2025)
Incidence, Risk Factors, and Modified Risk Assessment Model of Venous Thromboembolism in Non‐Hodgkin Lymphoma Patients
por: Wen Li, et al.
Publicado: (2024)
por: Wen Li, et al.
Publicado: (2024)
Energy spectrum of two-dimensional isotropic rapidly rotating turbulence
por: Li, Peiyang, et al.
Publicado: (2024)
por: Li, Peiyang, et al.
Publicado: (2024)
Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN
por: Xia, Ziyi, et al.
Publicado: (2026)
por: Xia, Ziyi, et al.
Publicado: (2026)
OminiAdapt: Learning Cross-Task Invariance for Robust and Environment-Aware Robotic Manipulation
por: Wang, Yongxu, et al.
Publicado: (2025)
por: Wang, Yongxu, et al.
Publicado: (2025)
VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
por: Li, Xinhao, et al.
Publicado: (2024)
por: Li, Xinhao, et al.
Publicado: (2024)
Cyclic Strain and Macrophage‐Mediated Transport Govern Micron‐Sized PM 2 . 5 Translocation across the Air–Blood Barrier
por: Yongjian Li, et al.
Publicado: (2025)
por: Yongjian Li, et al.
Publicado: (2025)
Guardrails Beat Guidance: A Large-Scale Study of Rules, Skills, and Persistent Configuration for Coding Agents
por: Zhang, Xing, et al.
Publicado: (2026)
por: Zhang, Xing, et al.
Publicado: (2026)
Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents
por: Zhang, Xing, et al.
Publicado: (2026)
por: Zhang, Xing, et al.
Publicado: (2026)
Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries
por: Zhang, Xing, et al.
Publicado: (2026)
por: Zhang, Xing, et al.
Publicado: (2026)
Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems
por: Zhang, Xing, et al.
Publicado: (2026)
por: Zhang, Xing, et al.
Publicado: (2026)
Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents
por: Zhang, Xing, et al.
Publicado: (2026)
por: Zhang, Xing, et al.
Publicado: (2026)
The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs
por: Zhang, Xing, et al.
Publicado: (2026)
por: Zhang, Xing, et al.
Publicado: (2026)
Revisiting the shadow of Johannsen-Psaltis black holes
por: Wang, Xinyu, et al.
Publicado: (2025)
por: Wang, Xinyu, et al.
Publicado: (2025)
Automating Agent Hijacking via Structural Template Injection
por: Deng, Xinhao, et al.
Publicado: (2026)
por: Deng, Xinhao, et al.
Publicado: (2026)
Ejemplares similares
-
How Social is It? A Benchmark for LLMs' Capabilities in Multi-user Multi-turn Social Agent Tasks
por: Wu, Yusen, et al.
Publicado: (2025) -
Robust and Reliable Early-Stage Website Fingerprinting Attacks via Spatial-Temporal Distribution Analysis
por: Deng, Xinhao, et al.
Publicado: (2024) -
Hummer: Towards Limited Competitive Preference Dataset
por: Jiang, Li, et al.
Publicado: (2024) -
Towards Fine-Grained Webpage Fingerprinting at Scale
por: Zhao, Xiyuan, et al.
Publicado: (2024) -
TrafficLLM: Enhancing Large Language Models for Network Traffic Analysis with Generic Traffic Representation
por: Cui, Tianyu, et al.
Publicado: (2025)