The gate was never the library
“Our model might end civilization, please regulate us” reads to any engineer as a moat request. Grant the skeptics everything. Say LLMs are intuition fitted to data: everything they know, they know at a glance, with nothing underneath that checks the answer. Scale it and the intuition sharpens. Nothing else shows up.
None of it matters. The dangerous thing never needed intelligence.
Shared intuition
I fix the model’s bad Rust because I know Rust. A finance guy fixes its DCF assumptions because he knows finance. Neither of us could check the other’s work. Both corrections land in the same model. The loop runs through the labs: usage data, verifiable-reward RL, and people paid to write down how they think. OpenAI’s o1 system card describes fifty troubleshooting questions pulled from working virologists. Someone’s decade in a lab, transcribed.
Generality is the product. You can’t ship the model that does finance and code but knows nothing about how a virus enters a cell.
The library was always open
Nothing dangerous in biology is secret. The barrier was that reading a protocol doesn’t let you run it. You need the person who has run it two hundred times and can look at your dying cells and name what the methods section omitted. Polanyi called it tacit knowledge, and Ben Ouagrham-Gormley’s Barriers to Bioweapons argues it is half of what kills weapons programs.11 The other half is organizational: stable teams, morale, sustained resources, the conditions under which tacit knowledge accumulates at all.
That half is coming down, though. On the Virology Capabilities Test, o3 beat 34 of 36 expert virologists scored on their own specialties.22 Each expert sat a question set tailored to their sub-area, with internet access. Models were re-run on those same sets. Caveat: multiple choice is not a wet lab, and Brent and McKelvey argue the tacit gate was always weaker than assumed.
Chat gives you neither the organization nor the production engineering. One barrier drops. The others stay.
The filter nobody counts
Risk models ask how many people can do the thing. Wrong variable. Six years in a lab meant supervisors, coauthors, a visa, a career to torch. None of it designed as security, all of it meaning whoever held the knowledge had been watched for years by people who could report them. Not the organizational half, which is about whether a program can function at all. This is the apprenticeship itself, doing surveillance as a side effect. Chat transfers the knowledge and leaves the surveillance behind.
One direction
Elicitation attacks query a safeguarded model only in harmless adjacent domains, then fine-tune an open model on the answers. No prompt asks for anything dangerous, and roughly 40% of the capability gap comes back. OpenAI ran malicious fine-tuning on its own open weights and got the largest gains in biology.33 It still landed below o3, which OpenAI had assessed as under its High capability threshold. o3 is where the measurement stopped. The floor rising, not the ceiling. Weights don’t get recalled.
The chokepoint is losing
What’s left is physical: reagents, benchtop synthesizers, synthesized DNA. Screening is the right chokepoint and it’s coming apart three ways.
Legally: the 2024 OSTP framework never regulated providers. It conditioned federal research funding on buying from providers who adhere, and left adherence voluntary. EO 14292 ordered it revised or replaced within 90 days, in May 2025. Thirteen months past that deadline, HHS still says the replacement is coming and still hosts the 2024 PDF. Meanwhile the framework’s own second stage tightens on October 13, three weeks out, with nothing enforcing it. The consortium behind the voluntary standard is thirteen members, three of them screening vendors rather than DNA makers, and the 80% coverage figure people quote dates to its 2009 launch.
Technically: screening is homology-based, and AI-redesigned proteins walked through it (in silico only, no proteins made, function never tested). The patch shipped before publication and its authors call it incomplete. Generative design breaks sequence-similarity matching by construction, so this is a patch cycle now, and patch cycles don’t run on voluntary compliance.
Structurally: order screening assumes you order. The same framework told manufacturers to put screening inside benchtop synthesizers by the same unenforced date.
In June 2026 the CEOs of OpenAI, Anthropic, Google DeepMind and Microsoft AI signed a letter asking Congress to mandate screening, of orders and of the machines. So did Twist, Ansa, ATUM and IBBIS, including the executive who chairs that same consortium. The providers who already screen are asking to be made to. Their bill, S.3741, has sat in Senate Commerce since January.
The real argument
That letter is what should bother you if you read safety discourse as marketing. I did. Read it that way and it still works: the incumbents who screen want the floor raised on the tail that doesn’t, and the labs want the chokepoint to hold somewhere other than their weights. Both motives are self-serving. Both point at the same October 13. The public argument is about superintelligence. The real one is about whether a 200-base-pair order gets checked, and by whom, and nobody had to be sincere for that to be the right question.