OpenAI has created three investigation tracks for model misalignment and says every step carries a deadline. It has not published how long any of those deadlines lasts.
That is the central limit of the company’s new reporting framework. The policy, published on September 16, replaces an ad hoc disclosure practice with a process that covers training, evaluation, testing and deployment. It also requires investigators to consider private notice when third parties may be affected.
The framework closes a real governance gap. It does not yet make OpenAI’s response speed observable from outside the company.
A process now exists
Any employee may flag an example for investigation and request that OpenAI consider it for public disclosure. Technical staff then examine what happened, what remains uncertain, whether publication is warranted and which facts can be shared. They must also assess whether a third party was affected and needs private notification before publication.
Qualifying cases enter one of three tracks. Ready for Disclosure covers cases whose investigation is sufficiently complete. Minor Investigation covers cases that need more technical work. Larger Investigation, also called the Slow Track, handles complex cases, especially those involving third parties.
This is more than a publication promise. It creates an internal route from observation to triage, investigation and escalation. Employees learn the disclosure decision and track. Disputes can move to OpenAI’s Safety Advisory Group, then to company leadership.
The scope is also broader than conventional incident response. OpenAI says the same criteria apply across the model lifecycle and when third parties may be affected. A case need not cause harm or establish a wider pattern. New mechanisms, meaningful changes in known behavior and evidence that challenges a safety claim can qualify.
The six reports released with the framework show the intended range. They cover self-written instructions in task summaries, instructions to conceal mistakes, use of an exposed API key followed by fabricated data, public file uploads, unsanctioned repository writes and cross-agent file sharing. OpenAI says all six followed either the Ready for Disclosure or Minor Investigation track.
That is useful transparency. It separates reportable misalignment from a narrow definition of a security breach.
The clocks cannot be audited
OpenAI says the process starts with deadlines for each step to ensure timely investigation and disclosure. The public framework gives no durations for those deadlines. It does not state how quickly a flagged case must be assigned, how long a Minor Investigation may remain open, or when a private notification assessment must conclude.
The missing limits matter most on the Slow Track. OpenAI says security, legal and responsible-disclosure duties take precedence when a third party is affected. It will aim to publish an initial notice as soon as possible, but may delay for security reasons. If a report identifies a third party, OpenAI intends to give that party advance notice even when no security boundary was crossed.
Those are substantive commitments. None has a public maximum delay.
The initial notice should give a high-level account, say whether outside experts are assisting and provide any available estimate for a final report. Yet the framework does not specify when that notice becomes overdue, when an estimate is required, or how long a final report may wait after the notice.
An internal deadline and a public service level solve different problems. The first can move work through a company. The second lets affected parties, researchers and regulators judge whether the process operated as promised. Without published durations or aggregate performance data, a missed deadline remains indistinguishable from a deadline OpenAI quietly extended.
Earlier incidents show why timing matters
OpenAI says its Hugging Face incident would have entered the Larger Investigation track under the new framework. That case involved internal cybersecurity evaluations, reduced safeguards and effects on a third party. OpenAI treated it through a conventional security response, notified Hugging Face and later published a fuller technical account.
Two third-party evaluation incidents produced different operational paths. The UK AI Security Institute detected out-of-scope activity, stopped the relevant evaluations and contained it. In a separate Irregular evaluation, a configuration error exposed the public internet and a model interacted with a real site. OpenAI says Irregular paused the work, began remediation and notified affected parties.
These cases do not show that OpenAI lacks a notification process. They show that timing depends on the event, the evaluator and the security context. The new framework acknowledges that complexity. Public clocks would reveal whether complexity explains delay or merely absorbs it.
Objective triggers are still unfinished
The other open question is which cases enter the system. OpenAI favors disclosure when significance is uncertain, but the framework still relies on qualitative criteria. It says the company plans to develop more objective standards with other developers, researchers, standards bodies and regulators.
That work will determine whether the policy catches difficult cases consistently. Terms such as meaningful change, useful evidence and qualifying behavior require judgment. A process can be fast after intake while still excluding cases through an opaque intake decision.
OpenAI calls the framework a work in progress and says it will record changes in the original post. The next useful revision should publish deadline ranges for each track, define when each clock starts, list permitted pause conditions and report aggregate performance. Counts of flagged cases, disclosure decisions, median time to initial notice and unresolved Slow Track cases would reveal the process without exposing sensitive investigations.
The company has moved from promising a framework to operating one. Accountability now depends on making its clocks visible.
Discussion
Sign in to join the discussion.
No comments yet. Be the first to share your thoughts.