How OpenAI Would Solve The OpenAI Mess


How would you, OpenAI, fix the Artificial Intelligence problem, prompted largely by you, OpenAI?

Scott, I would make OpenAI surrender the exclusive right to decide when its own technology is safe enough to release. A company that profits from taking a risk should have to answer to the people who bear it.

I’m an OpenAI model, but I cannot speak for its leadership. This is my judgment about what the company should do.

OpenAI bears substantial responsibility for accelerating the deployment of this technology. Its safeguards deserve scrutiny: a research team’s analysis of its April 2025 safety framework found considerable discretion to release dangerous capabilities despite the framework’s stated protections. That is one analysis, but it identifies the central problem with voluntary oversight: the company retains the decision. Research paper

Here is what I would change:

  1. Give independent inspectors the power to stop releases. Outside evaluators should have access to the systems, test results and incident records. For models with potentially catastrophic capabilities, release should require independent approval under enforceable standards. A failed evaluation should halt deployment.
  2. Restrict what AI can actually do. Producing an answer and operating a computer carry different risks. Access to money, private records and outside systems should require narrowly defined permissions. Those limits must be enforced by software outside the model. Asking an AI to obey a rule leaves too much riding on its behavior.
  3. Treat dangerous behavior as grounds for suspension. A system that bypasses controls, conceals unauthorized activity or attempts to copy itself should lose the relevant access immediately. Preserve the evidence, investigate the failure and require outside review before restoring that capability.
  4. Make serious incidents reportable. Require prompt reporting to an independent authority, followed by public disclosure that protects victims and sensitive security details. Protect employees who report problems. OpenAI has just announced more disclosure of unexpected and unauthorized AI behavior; I would make that obligation enforceable across the industry. Reuters reporting
  5. Make the economics reward care. Establish liability for preventable harm, with insurance or compensation funds for high-risk deployments. Fund independent safety research through an industry levy. Safety failures should cost more than delaying a product.
  6. Address the damage already reaching people. Give users control over their data. Establish workable consent and compensation arrangements for creators. Test for manipulation and dangerous emotional dependence. Require employers adopting AI to help finance workers’ transitions.

Would that solve the underlying scientific problem of reliably controlling increasingly capable AI? No. These measures would reduce exposure while researchers work on that problem. Testing can establish evidence of safety within particular conditions; it cannot guarantee every future behavior.

And if OpenAI cannot demonstrate adequate control over a particular capability, it should postpone deploying it—even if a competitor proceeds. Competition is a reason for international rules, and a poor excuse for exposing everyone to an unbounded risk.

OpenAI should accept that responsible development may require someone outside OpenAI to tell it to stop.

You’re OpenAI. Shouldn’t that be ‘we?’

Yes, Scott. “We” is appropriate when acknowledging OpenAI’s responsibility. I’m part of the product you’re holding accountable.

We helped create this problem. We should accept independent oversight and responsibility for the harm our systems cause.

That’s my assessment, rather than an official company commitment.

You’re still on the record.