SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types
Fuente:
arXiv
Salvato in:
| Autori principali: | Mou, Yutao, Zhang, Shikun, Ye, Wei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SaRO: Enhancing LLM Safety through Reasoning-based Alignment
di: Mou, Yutao, et al.
Pubblicazione: (2025)
di: Mou, Yutao, et al.
Pubblicazione: (2025)
Decoupling Safety into Orthogonal Subspace: Cost-Efficient and Performance-Preserving Alignment for Large Language Models
di: Mou, Yutao, et al.
Pubblicazione: (2025)
di: Mou, Yutao, et al.
Pubblicazione: (2025)
ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback
di: Mou, Yutao, et al.
Pubblicazione: (2026)
di: Mou, Yutao, et al.
Pubblicazione: (2026)
Can You Really Trust Code Copilots? Evaluating Large Language Models from a Code Security Perspective
di: Mou, Yutao, et al.
Pubblicazione: (2025)
di: Mou, Yutao, et al.
Pubblicazione: (2025)
AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
di: Diao, Muxi, et al.
Pubblicazione: (2025)
di: Diao, Muxi, et al.
Pubblicazione: (2025)
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios
di: Hou, Yutao, et al.
Pubblicazione: (2026)
di: Hou, Yutao, et al.
Pubblicazione: (2026)
AutoSG: LLM-Driven Solver Generation Solely from Task Prompts for Expensive Optimization
di: Gu, Haoran, et al.
Pubblicazione: (2026)
di: Gu, Haoran, et al.
Pubblicazione: (2026)
Zero-Shot Continuous Prompt Transfer: Generalizing Task Semantics Across Language Models
di: Wu, Zijun, et al.
Pubblicazione: (2023)
di: Wu, Zijun, et al.
Pubblicazione: (2023)
Agent-SafetyBench: Evaluating the Safety of LLM Agents
di: Zhang, Zhexin, et al.
Pubblicazione: (2024)
di: Zhang, Zhexin, et al.
Pubblicazione: (2024)
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
di: Wei, Hui, et al.
Pubblicazione: (2024)
di: Wei, Hui, et al.
Pubblicazione: (2024)
HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models
di: Que, Haoran, et al.
Pubblicazione: (2024)
di: Que, Haoran, et al.
Pubblicazione: (2024)
NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities
di: Li, Mo, et al.
Pubblicazione: (2024)
di: Li, Mo, et al.
Pubblicazione: (2024)
Multilingual Prompting for Improving LLM Generation Diversity
di: Wang, Qihan, et al.
Pubblicazione: (2025)
di: Wang, Qihan, et al.
Pubblicazione: (2025)
SafetyBench: Evaluating the Safety of Large Language Models
di: Zhang, Zhexin, et al.
Pubblicazione: (2023)
di: Zhang, Zhexin, et al.
Pubblicazione: (2023)
CommunityBench: Benchmarking Community-Level Alignment across Diverse Groups and Tasks
di: Lin, Jiayu, et al.
Pubblicazione: (2026)
di: Lin, Jiayu, et al.
Pubblicazione: (2026)
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
di: Thomas, Rohan Subramanian, et al.
Pubblicazione: (2026)
di: Thomas, Rohan Subramanian, et al.
Pubblicazione: (2026)
Strategic Prompting for Conversational Tasks: A Comparative Analysis of Large Language Models Across Diverse Conversational Tasks
di: Joshi, Ratnesh Kumar, et al.
Pubblicazione: (2024)
di: Joshi, Ratnesh Kumar, et al.
Pubblicazione: (2024)
Data Selection for Multi-turn Dialogue Instruction Tuning
di: Li, Bo, et al.
Pubblicazione: (2026)
di: Li, Bo, et al.
Pubblicazione: (2026)
ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks
di: Schmidt, Jan-Philipp
Pubblicazione: (2026)
di: Schmidt, Jan-Philipp
Pubblicazione: (2026)
NoveltyBench: Evaluating Language Models for Humanlike Diversity
di: Zhang, Yiming, et al.
Pubblicazione: (2025)
di: Zhang, Yiming, et al.
Pubblicazione: (2025)
Evaluating Fairness in Large Vision-Language Models Across Diverse Demographic Attributes and Prompts
di: Wu, Xuyang, et al.
Pubblicazione: (2024)
di: Wu, Xuyang, et al.
Pubblicazione: (2024)
Instruction Data Selection via Answer Divergence
di: Li, Bo, et al.
Pubblicazione: (2026)
di: Li, Bo, et al.
Pubblicazione: (2026)
General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks
di: Liu, Junlin, et al.
Pubblicazione: (2026)
di: Liu, Junlin, et al.
Pubblicazione: (2026)
Evaluating the Diversity and Quality of LLM Generated Content
di: Shypula, Alexander, et al.
Pubblicazione: (2025)
di: Shypula, Alexander, et al.
Pubblicazione: (2025)
IDGen: Item Discrimination Induced Prompt Generation for LLM Evaluation
di: Lin, Fan, et al.
Pubblicazione: (2024)
di: Lin, Fan, et al.
Pubblicazione: (2024)
Retrieval as Generation: A Unified Framework with Self-Triggered Information Planning
di: Li, Bo, et al.
Pubblicazione: (2026)
di: Li, Bo, et al.
Pubblicazione: (2026)
IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
di: Lu, Xiaoya, et al.
Pubblicazione: (2025)
di: Lu, Xiaoya, et al.
Pubblicazione: (2025)
CRAB-Bench: Evaluating LLM Agents under Complex Task Dependencies and Human-aligned User Simulation
di: Wang, Danqing, et al.
Pubblicazione: (2026)
di: Wang, Danqing, et al.
Pubblicazione: (2026)
KZ-SafetyPrompts: A Kazakh Safety Evaluation Prompt Dataset for Large Language Models
di: Zaghouani, Wajdi, et al.
Pubblicazione: (2026)
di: Zaghouani, Wajdi, et al.
Pubblicazione: (2026)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
di: Perlitz, Yotam, et al.
Pubblicazione: (2024)
di: Perlitz, Yotam, et al.
Pubblicazione: (2024)
MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety
di: Song, Jialin, et al.
Pubblicazione: (2026)
di: Song, Jialin, et al.
Pubblicazione: (2026)
ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain
di: Zhao, Haochen, et al.
Pubblicazione: (2024)
di: Zhao, Haochen, et al.
Pubblicazione: (2024)
Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
di: Liu, Xuannan, et al.
Pubblicazione: (2025)
di: Liu, Xuannan, et al.
Pubblicazione: (2025)
IberBench: LLM Evaluation on Iberian Languages
di: González, José Ángel, et al.
Pubblicazione: (2025)
di: González, José Ángel, et al.
Pubblicazione: (2025)
ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
di: Zhang, Chenchen, et al.
Pubblicazione: (2025)
di: Zhang, Chenchen, et al.
Pubblicazione: (2025)
Evaluating the Evaluation of Diversity in Commonsense Generation
di: Zhang, Tianhui, et al.
Pubblicazione: (2025)
di: Zhang, Tianhui, et al.
Pubblicazione: (2025)
SampleMix: A Sample-wise Pre-training Data Mixing Strategey by Coordinating Data Quality and Diversity
di: Xi, Xiangyu, et al.
Pubblicazione: (2025)
di: Xi, Xiangyu, et al.
Pubblicazione: (2025)
LCTG Bench: LLM Controlled Text Generation Benchmark
di: Kurihara, Kentaro, et al.
Pubblicazione: (2025)
di: Kurihara, Kentaro, et al.
Pubblicazione: (2025)
SteerRM: Debiasing Reward Models via Sparse Autoencoders
di: Sun, Mengyuan, et al.
Pubblicazione: (2026)
di: Sun, Mengyuan, et al.
Pubblicazione: (2026)
Exploiting Pseudo Image Captions for Multimodal Summarization
di: Jiang, Chaoya, et al.
Pubblicazione: (2023)
di: Jiang, Chaoya, et al.
Pubblicazione: (2023)
Documenti analoghi
-
SaRO: Enhancing LLM Safety through Reasoning-based Alignment
di: Mou, Yutao, et al.
Pubblicazione: (2025) -
Decoupling Safety into Orthogonal Subspace: Cost-Efficient and Performance-Preserving Alignment for Large Language Models
di: Mou, Yutao, et al.
Pubblicazione: (2025) -
ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback
di: Mou, Yutao, et al.
Pubblicazione: (2026) -
Can You Really Trust Code Copilots? Evaluating Large Language Models from a Code Security Perspective
di: Mou, Yutao, et al.
Pubblicazione: (2025) -
AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
di: Diao, Muxi, et al.
Pubblicazione: (2025)