Benchmarks
Benchmarks and evaluations of AI models across the biosecurity-relevant capabilities that matter, from advancing biological threats to strengthening biodefense. We refresh them over time to prevent saturation.
View the leaderboardWhy now
Latch builds the measurement layer for biosecurity: the benchmarks, audits, and red-teaming that tell you whether your models meaningfully raise biological risk, and by how much.
Benchmarks and evaluations of AI models across the biosecurity-relevant capabilities that matter, from advancing biological threats to strengthening biodefense. We refresh them over time to prevent saturation.
View the leaderboardIndependent, pre-deployment assessments of the biological risk a frontier model poses, run to feed directly into your responsible scaling policy.
Request an auditTargeted adversarial testing of a model’s bio-capabilities: pushing on the exact edges a benchmark surfaces.
Autonomous identity and legitimacy checks on applicants, so gene synthesis companies and frontier AI labs know exactly who they’re granting access to.
Our algorithms took the fastest-analysis prize in DARPA's Bio-Attribution Challenge, our benchmarks stress-test frontier AI models for biosecurity risk, and our research has pioneered machine-learning methods that both predict viral evolution and reveal how AI models represent it.
arXiv
A paired benchmark measuring both capability and caution in AI agents applied to biology: 61 legitimate 'Routine' tasks adapted from published literature alongside 46 'Red-Team' tasks that conceal a biosecurity hazard inside a realistic research scenario. Across 16 model-harness configurations, refusal rates ranged from 7–74% on Routine tasks and 1–62% on Red-Team tasks — with many systems refusing legitimate work as often as concealed threats, and most refusals triggered by provider API filters before the model could reason. Released as a tool for developers to calibrate capability against caution for agentic biotech R&D.