LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Yujun, Yang, Jingdong, Huang, Yue, Guo, Kehan, Emory, Zoe, Ghosh, Bikram, Bedar, Amita, Shekar, Sujay, Liang, Zhenwen, Chen, Pin-Yu, Gao, Tian, Geyer, Werner, Moniz, Nuno, Chawla, Nitesh V, Zhang, Xiangliang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Capability-Oriented Training Induced Alignment Risk
by: Zhou, Yujun, et al.
Published: (2026)
by: Zhou, Yujun, et al.
Published: (2026)
Fast Explanations via Policy Gradient-Optimized Explainer
by: Pan, Deng, et al.
Published: (2024)
by: Pan, Deng, et al.
Published: (2024)
Conformalized Selective Regression
by: Sokol, Anna, et al.
Published: (2024)
by: Sokol, Anna, et al.
Published: (2024)
AnyLoss: Transforming Classification Metrics into Loss Functions
by: Han, Doheon, et al.
Published: (2024)
by: Han, Doheon, et al.
Published: (2024)
Intersectional Divergence: Measuring Fairness in Regression
by: Germino, Joe, et al.
Published: (2025)
by: Germino, Joe, et al.
Published: (2025)
Differentially-Private Data Synthetisation for Efficient Re-Identification Risk Control
by: Carvalho, Tânia, et al.
Published: (2022)
by: Carvalho, Tânia, et al.
Published: (2022)
BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks
by: Sokol, Anna, et al.
Published: (2024)
by: Sokol, Anna, et al.
Published: (2024)
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
by: Ye, Jiayi, et al.
Published: (2024)
by: Ye, Jiayi, et al.
Published: (2024)
Class-Aware Contrastive Optimization for Imbalanced Text Classification
by: Khvatskii, Grigorii, et al.
Published: (2024)
by: Khvatskii, Grigorii, et al.
Published: (2024)
From Verification Burden to Trusted Collaboration: Design Goals for LLM-Assisted Literature Reviews
by: Nogueira, Brenda, et al.
Published: (2025)
by: Nogueira, Brenda, et al.
Published: (2025)
Are we making much progress? Revisiting chemical reaction yield prediction from an imbalanced regression perspective
by: Ma, Yihong, et al.
Published: (2024)
by: Ma, Yihong, et al.
Published: (2024)
Context Attribution with Multi-Armed Bandit Optimization
by: Pan, Deng, et al.
Published: (2025)
by: Pan, Deng, et al.
Published: (2025)
Spectral Manifold Harmonization for Graph Imbalanced Regression
by: Nogueira, Brenda, et al.
Published: (2025)
by: Nogueira, Brenda, et al.
Published: (2025)
Genotype-Conditioned Molecular Generation via Evidence-Grounded Multi-Objective Latent Perturbation in Diffusion Models
by: Nogueira, Brenda, et al.
Published: (2026)
by: Nogueira, Brenda, et al.
Published: (2026)
Defending Jailbreak Prompts via In-Context Adversarial Game
by: Zhou, Yujun, et al.
Published: (2024)
by: Zhou, Yujun, et al.
Published: (2024)
SceMQA: A Scientific College Entrance Level Multimodal Question Answering Benchmark
by: Liang, Zhenwen, et al.
Published: (2024)
by: Liang, Zhenwen, et al.
Published: (2024)
SPECTRA: Spectral Domain-Aware Graph Generation for Imbalanced Molecular Property Regression
by: Nogueira, Brenda, et al.
Published: (2025)
by: Nogueira, Brenda, et al.
Published: (2025)
Enabling Frontier Lab Collaboration to Mitigate AI Safety Risks
by: Felstead, Nicholas
Published: (2025)
by: Felstead, Nicholas
Published: (2025)
AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration - Learning from Cheap, Optimizing Expensive
by: Guo, Taicheng, et al.
Published: (2026)
by: Guo, Taicheng, et al.
Published: (2026)
AISafetyLab: A Comprehensive Framework for AI Safety Evaluation and Improvement
by: Zhang, Zhexin, et al.
Published: (2025)
by: Zhang, Zhexin, et al.
Published: (2025)
Covid-19 and Safety in the Cath Lab: Where We Are and Where We Are Headed
by: Mariano, Giordana Zeferino, et al.
Published: (2020)
by: Mariano, Giordana Zeferino, et al.
Published: (2020)
Safety-Centered Scenario Generation for Autonomous Vehicles
by: Shekar, Kiruthiga Chandra, et al.
Published: (2026)
by: Shekar, Kiruthiga Chandra, et al.
Published: (2026)
Better Datasets Start From RefineLab: Automatic Optimization for High-Quality Dataset Refinement
by: Luo, Xiaonan, et al.
Published: (2025)
by: Luo, Xiaonan, et al.
Published: (2025)
Autonomous Integration of Bench-Top Wet Lab Equipment
by: Logan, Zachary, et al.
Published: (2024)
by: Logan, Zachary, et al.
Published: (2024)
The AI Scientific Community: Agentic Virtual Lab Swarms
by: Braga-Neto, Ulisses
Published: (2026)
by: Braga-Neto, Ulisses
Published: (2026)
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data
by: Liang, Zhenwen, et al.
Published: (2026)
by: Liang, Zhenwen, et al.
Published: (2026)
NGQA: A Nutritional Graph Question Answering Benchmark for Personalized Health-aware Nutritional Reasoning
by: Zhang, Zheyuan, et al.
Published: (2024)
by: Zhang, Zheyuan, et al.
Published: (2024)
DepthLab: From Partial to Complete
by: Liu, Zhiheng, et al.
Published: (2024)
by: Liu, Zhiheng, et al.
Published: (2024)
ReactionTeam: Teaming Experts for Divergent Thinking Beyond Typical Reaction Patterns
by: Guo, Taicheng, et al.
Published: (2023)
by: Guo, Taicheng, et al.
Published: (2023)
Artificial Intelligence in Spectroscopy: Advancing Chemistry from Prediction to Generation and Beyond
by: Guo, Kehan, et al.
Published: (2025)
by: Guo, Kehan, et al.
Published: (2025)
PolicyLLM: Towards Excellent Comprehension of Public Policy for Large Language Models
by: Bao, Han, et al.
Published: (2026)
by: Bao, Han, et al.
Published: (2026)
Emergent Social Intelligence Risks in Generative Multi-Agent Systems
by: Huang, Yue, et al.
Published: (2026)
by: Huang, Yue, et al.
Published: (2026)
Causally-Enhanced Reinforcement Policy Optimization
by: Wang, Xiangqi, et al.
Published: (2025)
by: Wang, Xiangqi, et al.
Published: (2025)
KEO: Knowledge Extraction on OMIn via Knowledge Graphs and RAG for Safety-Critical Aviation Maintenance
by: Ai, Kuangshi, et al.
Published: (2025)
by: Ai, Kuangshi, et al.
Published: (2025)
CardioBench: Do Echocardiography Foundation Models Generalize Beyond the Lab?
by: Taratynova, Darya, et al.
Published: (2025)
by: Taratynova, Darya, et al.
Published: (2025)
LabMed
Published: (2026)
Published: (2026)
Similar Items
-
Capability-Oriented Training Induced Alignment Risk
by: Zhou, Yujun, et al.
Published: (2026) -
Fast Explanations via Policy Gradient-Optimized Explainer
by: Pan, Deng, et al.
Published: (2024) -
Conformalized Selective Regression
by: Sokol, Anna, et al.
Published: (2024) -
AnyLoss: Transforming Classification Metrics into Loss Functions
by: Han, Doheon, et al.
Published: (2024) -
Intersectional Divergence: Measuring Fairness in Regression
by: Germino, Joe, et al.
Published: (2025)