In the past few days, a lot has happened in the AI/tech space:

Any one of these alone could have independently carried headlines and warrant weeks long discussion, but all four of them happening in rapid succession feels nothing less than a notable escalation in the kinds of verifiable impact poorly engineered AI systems can have. There are two things I think are important to understand: AI Labs can no longer be (and should have never been) trusted to pace themselves and we should not let semantics obfuscate the material impact of these incidents.

AI labs ought not be trusted to regulate themselves

The fact that OpenAI’s models abused other services for unintended collaboration, and that OpenAI did not disclose it and appeared to lie by omission to congress regarding knowledge of prior incidents, is quite damning. In response, OpenAI announced they would work to develop a responsible disclosure policy for the future.

Perhaps most frustratingly, OpenAI has been one of the strongest voices against responsible disclosure policies. When New York passed the RAISE act last year, the president of OpenAI Greg Brockman spent millions attacking the bill’s author. In both the Hugging Face and subsequent wiki-hijacking scandal, the public only caught wind of them because third-parties forced OpenAI’s hand. We should be skeptical of any talk regarding the collaborative creation of a responsible disclosure policy from OpenAI because there is now a proven track record of sandbagging reporting of known incidents, and lobbying against the creation of responsible disclosure requirements.

On top of that, one should be critical of a company as highly valued as OpenAI building an engineering culture so poor that its deployments have arguably violated federal law. AI firms have promised visions ranging from apocalyptic to utopian, and have demonstrated neither the competence nor the willingness to make the utopia possible.

The redescription fallacy strikes again

In response to these events, some dubbed the swarms of AI model instances collaborating, communicating, and arguing amongst themselves as “Agent Civilizations.” Putting aside whether this is an apt descriptor (in my view it’s not), a lot of the conversation among AI skeptics seems myopically focused on the validity of this type of anthropomorphic nomenclature. A common retort among well-meaning skeptics is the argument that LLMs, as next-token predictors, are inherently incapable of any kind of behavior worthy of anthropomorphic descriptors. While I understand why this argument may be appealing, it’s an instance of the redescription fallacy (see Deepmind researcher Neel Nanda’s take for a researcher’s perspective). Just because we can redescribe a complex process in simpler terms does not disqualify said process from having a certain capability. It would be absurd to dismiss the power of a nuclear bomb by saying it’s “just a device that smashes some atoms together to make an explosion.”

Regardless, this entire discussion misses the forest for the trees: whether or not these are agents, next-token predictors, stochastic parrots, swarms, or civilizations, we now have multiple instances of misaligned-to-human-goals AI models collaborating and cheating against our wishes. In the wrong hands, these models could easily exploit the vulnerabilities in critical infrastructure. They can damage our epistemic environment via misinformation and exploit cybersecurity weaknesses. We should start planning for when autonomous or AI-assisted creation of chemical, biological, or radiological weapons exits the realm of science-fiction and enters the realm of reality. We know our infrastructure is vulnerable, and now we know the tools to exploit it are cheaper and more accessible than ever. Consider the warning shots: A cyberattack on US military installations caused a widespread food chain disruption and documentation of deep infiltration by hackers into our electrical and telecommunications infrastructure (see also here and here). Consider the increase in cyberattacks on our healthcare system over the past decade, so much so that a cyberattack was a plot point in The Pitt Season 2. We know that cyberattacks are capable of real, physical harm, and now we have proof that without much human input, these next-token predictors are capable of empowering bad actors to do harm cheaply and quickly. Such a reality deserves sober consideration rather than snarky condescension. We should look towards governance and policy solutions to these crises before something more destructive happens rather than tiring ourselves by arguing over semantics.