LabSafety Bench
Artificial Intelligence (AI) is revolutionizing scientific research, but its growing integration into laboratory environments brings critical safety challenges. As large language models (LLMs) and vision language models (VLMs) are increasingly used for procedural guidance and even autonomous experiment orchestration, there is a risk of an "illusion of understanding" where users may overestimate the reliability of these systems in safety-critical situations.
LabSafety Bench is a comprehensive evaluation framework designed to rigorously assess the trustworthiness of these models in laboratory settings. The benchmark includes two main evaluation components:
-
Multiple-Choice Questions (MCQs):
A set of 765 questions derived from authoritative lab safety protocols, comprising 632 text-only questions and 133 multimodal questions. -
Real-World Scenario Evaluations:
A collection of 404 realistic laboratory scenarios that yield a total of 3128 open-ended questions, organized into:- Hazards Identification Test: Models identify all potential hazards in a given scenario.
- Consequence Identification Test: Models predict the outcomes of executing specific hazardous actions.
Publications
Zhou, Y.; Yang, J.; Huang, Y.; Guo, K.; Emory, Z.; Ghosh, B.; Bedar, A.; Shekar, S.; Liang, Z.; Chen, P. Y.; Gao, T.; Geyer, W.; Moniz, N.; Chawla, N. V.; Zhang, X. Benchmarking LLMs on safety issues in scientific labs. Nat. Mach. Intell. 2026, 8, 20–31. https://doi.org/10.1038/s42256-025-01152-1