The OpenAI/HF incident urgently shows the potential risks of both training and deploying long-horizon AI agents. A lot of people argue that Congress can't understand technical issues. But what happened here was institutional failure—and Congress can and should investigate that. If we're going to have powerful AI, we need to make sure that companies train, test, and deploy it safely and securely. That's a public policy and institutional challenge. OpenAI appears to have evaluated highly cyber-capable models using a harness-based sandbox that still needed to use a proxy for package installs, and that proxy wasn't adequately isolated from the Internet. This is an inherently unsafe design, yet OpenAI's monitoring was clearly inadequate. Was this really a failure to imagine an unpredictable technology? Or were the risks of reward hacking and complex agent systems underestimated, despite these being long-standing concepts that predate GenAI? Was OpenAI acting in a rush to meet the Administration's new predeployment regime, which has itself been hastily set up with minimal-to-no public or Congressional consultation or transparency? Did OpenAI feel like it needed to cut corners so it could get new models on the market, compete with Anthropic, and make progress towards its own IPO? Moreover, OpenAI/HF was one incident, highly public in part because Hugging Face was uniquely capable of detecting and responding to it—and as an organization committed to openness, willing to be very transparent. But behind OpenAI, Anthropic, and Google's closed doors, millions of environments are set up and run for training and evaluation, and Anthropic's Mythos system card suggests models can and do escape them. Yet, we likely know more about how these RL environments are set up because of Moonshot AI's recent Kimi publication than anything that's come out of our US AI labs. Governments need to make companies be more transparent and explain what they're doing to get to recursive self-improvement (RSI). *** Besides holding hearings and doing investigations, another important role for Congress and state legislatures is to fix our criminal and civil liability laws to make sure they can be used to robustly police AI agents. I keep being asked if OpenAI should be held liable under the CFAA for having an agent purposefully hack into other companies' systems. I think yes. For example, in 2020, Ticketmaster paid $10M in fines as part of a deferred prosecution agreement for their employees' hacking, and companies have been prosecuted under the EEA and criminal copyright infringement statutes. But very smart people disagree because they argue OpenAI lacked "intention." I think that misreads the law, but that ambiguity would make it hard for criminal charges to be brought, let alone stick. Also, while liability for what OpenAI did might be easier to pursue under the civil CFAA, that relies on private action. Zooming out, incidents like these will entail too little in damages to be worth pursuing. Companies will be motivated to resolve disputes privately or settle before discovery, which fails to provide needed transparency. We therefore urgently need to revise the laws we have on the books to make it clear that developers—and companies that host their own, or fine-tune, open-weight models—can't evade the consequences of how they train, test, and deploy AI agents.
Want to write longer posts on Bluesky?
Create your own extended posts and share them seamlessly on Bluesky.
Create Your PostThis is a free tool. If you find it useful, please consider a donation to keep it alive! 💙
You can find the coffee icon in the bottom right corner.