FlexGuard: Continuous Risk Scoring for Strictness-Adaptive LLM Content Moderation
Fuente:
arXiv
Salvato in:
| Autori principali: | Ding, Zhihao, Li, Jinming, Lu, Ze, Shi, Jieming |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Effective Illicit Account Detection on Large Cryptocurrency MultiGraphs
di: Ding, Zhihao, et al.
Pubblicazione: (2023)
di: Ding, Zhihao, et al.
Pubblicazione: (2023)
RingFormer: A Ring-Enhanced Graph Transformer for Organic Solar Cell Property Prediction
di: Ding, Zhihao, et al.
Pubblicazione: (2024)
di: Ding, Zhihao, et al.
Pubblicazione: (2024)
LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models
di: Elesedy, Hayder, et al.
Pubblicazione: (2024)
di: Elesedy, Hayder, et al.
Pubblicazione: (2024)
SGOOD: Substructure-enhanced Graph-Level Out-of-Distribution Detection
di: Ding, Zhihao, et al.
Pubblicazione: (2023)
di: Ding, Zhihao, et al.
Pubblicazione: (2023)
Standardized Interpretable Fairness Measures for Continuous Risk Scores
di: Becker, Ann-Kristin, et al.
Pubblicazione: (2023)
di: Becker, Ann-Kristin, et al.
Pubblicazione: (2023)
Reliability Auditing for Downstream LLM tasks in Psychiatry: LLM-Generated Hospitalization Risk Scores
di: Panda, Shevya, et al.
Pubblicazione: (2026)
di: Panda, Shevya, et al.
Pubblicazione: (2026)
From Parameter Dynamics to Risk Scoring : Quantifying Sample-Level Safety Degradation in LLM Fine-tuning
di: Wang, Xiao, et al.
Pubblicazione: (2026)
di: Wang, Xiao, et al.
Pubblicazione: (2026)
Distillation Traps and Guards: A Calibration Knob for LLM Distillability
di: Zhan, Weixiao, et al.
Pubblicazione: (2026)
di: Zhan, Weixiao, et al.
Pubblicazione: (2026)
FlexLLM: Composable HLS Library for Flexible Hybrid LLM Accelerator Design
di: Zhang, Jiahao, et al.
Pubblicazione: (2026)
di: Zhang, Jiahao, et al.
Pubblicazione: (2026)
Choosing How to Remember: Adaptive Memory Structures for LLM Agents
di: Lu, Mingfei, et al.
Pubblicazione: (2026)
di: Lu, Mingfei, et al.
Pubblicazione: (2026)
Is One Score Enough? Rethinking the Evaluation of Sequentially Evolving LLM Memory
di: Dong, Songwei, et al.
Pubblicazione: (2026)
di: Dong, Songwei, et al.
Pubblicazione: (2026)
Accelerating Asynchronous Federated Learning Convergence via Opportunistic Mobile Relaying
di: Bian, Jieming, et al.
Pubblicazione: (2022)
di: Bian, Jieming, et al.
Pubblicazione: (2022)
GuardReasoner: Towards Reasoning-based LLM Safeguards
di: Liu, Yue, et al.
Pubblicazione: (2025)
di: Liu, Yue, et al.
Pubblicazione: (2025)
MSSR: Memory-Aware Adaptive Replay for Continual LLM Fine-Tuning
di: Lu, Yiyang, et al.
Pubblicazione: (2026)
di: Lu, Yiyang, et al.
Pubblicazione: (2026)
COIN: Uncertainty-Guarding Selective Question Answering for Foundation Models with Provable Risk Guarantees
di: Wang, Zhiyuan, et al.
Pubblicazione: (2025)
di: Wang, Zhiyuan, et al.
Pubblicazione: (2025)
Think Inside the JSON: Reinforcement Strategy for Strict LLM Schema Adherence
di: Agarwal, Bhavik, et al.
Pubblicazione: (2025)
di: Agarwal, Bhavik, et al.
Pubblicazione: (2025)
Content Moderation by LLM: From Accuracy to Legitimacy
di: Huang, Tao
Pubblicazione: (2024)
di: Huang, Tao
Pubblicazione: (2024)
Guarding Graph Neural Networks for Unsupervised Graph Anomaly Detection
di: Bei, Yuanchen, et al.
Pubblicazione: (2024)
di: Bei, Yuanchen, et al.
Pubblicazione: (2024)
FlexMol: A Flexible Toolkit for Benchmarking Molecular Relational Learning
di: Liu, Sizhe, et al.
Pubblicazione: (2024)
di: Liu, Sizhe, et al.
Pubblicazione: (2024)
FlexCare: Leveraging Cross-Task Synergy for Flexible Multimodal Healthcare Prediction
di: Xu, Muhao, et al.
Pubblicazione: (2024)
di: Xu, Muhao, et al.
Pubblicazione: (2024)
CourtGuard: A Model-Agnostic Framework for Zero-Shot Policy Adaptation in LLM Safety
di: Suleymanov, Umid, et al.
Pubblicazione: (2026)
di: Suleymanov, Umid, et al.
Pubblicazione: (2026)
On the Implicit Adversariality of Catastrophic Forgetting in Deep Continual Learning
di: Peng, Ze, et al.
Pubblicazione: (2025)
di: Peng, Ze, et al.
Pubblicazione: (2025)
Online Domain-aware LLM Decoding for Continual Domain Evolution
di: Abu-Shaira, Mohammad, et al.
Pubblicazione: (2026)
di: Abu-Shaira, Mohammad, et al.
Pubblicazione: (2026)
GuardFed: A Trustworthy Federated Learning Framework Against Dual-Facet Attacks
di: Li, Yanli, et al.
Pubblicazione: (2025)
di: Li, Yanli, et al.
Pubblicazione: (2025)
The WidthWall: A Strict Expressivity Hierarchy for Hypergraph Neural Networks
di: Jiang, Fengqing, et al.
Pubblicazione: (2026)
di: Jiang, Fengqing, et al.
Pubblicazione: (2026)
HalluGuard: Demystifying Data-Driven and Reasoning-Driven Hallucinations in LLMs
di: Zeng, Xinyue, et al.
Pubblicazione: (2026)
di: Zeng, Xinyue, et al.
Pubblicazione: (2026)
FedALT: Federated Fine-Tuning through Adaptive Local Training with Rest-of-World LoRA
di: Bian, Jieming, et al.
Pubblicazione: (2025)
di: Bian, Jieming, et al.
Pubblicazione: (2025)
PropGuard: Safeguarding LLM-MAS via Propagation-Aware Exploration and Remediation
di: Yan, Bingyu, et al.
Pubblicazione: (2026)
di: Yan, Bingyu, et al.
Pubblicazione: (2026)
Reliable Self-Harm Risk Screening via Adaptive Multi-Agent LLM Systems
di: Karnam, Meghana, et al.
Pubblicazione: (2026)
di: Karnam, Meghana, et al.
Pubblicazione: (2026)
Shared LoRA Subspaces for almost Strict Continual Learning
di: Kaushik, Prakhar, et al.
Pubblicazione: (2026)
di: Kaushik, Prakhar, et al.
Pubblicazione: (2026)
RAP: Runtime Adaptive Pruning for LLM Inference
di: Liu, Huanrong, et al.
Pubblicazione: (2025)
di: Liu, Huanrong, et al.
Pubblicazione: (2025)
X-Guard: Multilingual Guard Agent for Content Moderation
di: Upadhayay, Bibek, et al.
Pubblicazione: (2025)
di: Upadhayay, Bibek, et al.
Pubblicazione: (2025)
DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
di: Wang, Yaxuan, et al.
Pubblicazione: (2025)
di: Wang, Yaxuan, et al.
Pubblicazione: (2025)
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
di: Yang, Lijie, et al.
Pubblicazione: (2024)
di: Yang, Lijie, et al.
Pubblicazione: (2024)
FlexTSF: A Flexible Forecasting Model for Time Series with Variable Regularities
di: Xiao, Jingge, et al.
Pubblicazione: (2024)
di: Xiao, Jingge, et al.
Pubblicazione: (2024)
CSAttention: Centroid-Scoring Attention for Accelerating LLM Inference
di: Song, Chuxu, et al.
Pubblicazione: (2026)
di: Song, Chuxu, et al.
Pubblicazione: (2026)
Semi-Supervised Learning for Large Language Models Safety and Content Moderation
di: Dinuta, Eduard Stefan, et al.
Pubblicazione: (2025)
di: Dinuta, Eduard Stefan, et al.
Pubblicazione: (2025)
FlexMS is a flexible framework for benchmarking deep learning-based mass spectrum prediction tools in metabolomics
di: Zhong, Yunhua, et al.
Pubblicazione: (2026)
di: Zhong, Yunhua, et al.
Pubblicazione: (2026)
Dimensional Characterization and Pathway Modeling for Catastrophic AI Risks
di: Chin, Ze Shen
Pubblicazione: (2025)
di: Chin, Ze Shen
Pubblicazione: (2025)
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
di: Lee, Jung Hyun, et al.
Pubblicazione: (2023)
di: Lee, Jung Hyun, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Effective Illicit Account Detection on Large Cryptocurrency MultiGraphs
di: Ding, Zhihao, et al.
Pubblicazione: (2023) -
RingFormer: A Ring-Enhanced Graph Transformer for Organic Solar Cell Property Prediction
di: Ding, Zhihao, et al.
Pubblicazione: (2024) -
LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models
di: Elesedy, Hayder, et al.
Pubblicazione: (2024) -
SGOOD: Substructure-enhanced Graph-Level Out-of-Distribution Detection
di: Ding, Zhihao, et al.
Pubblicazione: (2023) -
Standardized Interpretable Fairness Measures for Continuous Risk Scores
di: Becker, Ann-Kristin, et al.
Pubblicazione: (2023)