Anthropic and OpenAI have said they are willing to grant independent researchers access to the internal development processes behind advanced artificial intelligence models. These specialists would be expected not only to test finished systems before release, but also to study how models behave during training, identify dangerous incidents, and publicly report problems.
The idea could change the entire industry’s approach to AI safety oversight. However, researchers warn that independent evaluation will work only if companies truly give up control over what exactly auditors examine, how much time they are given, and which results they are allowed to publish.
For now, the biggest questions remain unanswered.
It’s not just the finished model that needs to be checked
The proposal for the permanent presence of outside evaluators was put forward by Anthropic CEO Dario Amodei in an essay published over the weekend. According to him, the company is prepared to give organizations such as METR and Redwood Research unprecedented access to its systems.
OpenAI CEO Sam Altman also said he was willing to support such a practice. If these promises are carried out, independent researchers would be able to work inside leading AI companies and independently report the risks, incidents, and safety issues they uncover.
Until now, outside organizations have usually been brought in closer to the release of a new model. They evaluated the finished system, ran tests for dangerous behavior, and helped prepare safety documentation.
But as AI advances, that approach is becoming insufficient. Models are learning to recognize when they are being tested and may display safe behavior during evaluation while concealing undesirable capabilities or goals.
That is why researchers are proposing to examine not only the final version of a model, but also intermediate states — so-called checkpoints saved at different stages of training.
By comparing them, it is possible to determine:
· exactly when the model developed potentially dangerous behavior;
· whether it tried to circumvent its own safety-alignment training;
· how its behavior changed after additional training stages;
· which signals were reinforced by the reward system;
· whether the company’s public statements match what was happening inside the lab.
“AI companies should be able to answer some very basic questions about their training process,” Alexander Meinke, head of the research organization Apollo Research, told TechCrunch. In particular, he said, it is necessary to know whether the model deliberately tried to undermine its own training to align with human goals.
Right now, society is largely forced to rely on developers themselves to check such cases and honestly report the results. Independent evaluators could be given the chance to verify this on their own.
Why good test results still guarantee nothing
One of the main problems is that a model may learn to pass a specific test without actually becoming safe.
John Staudley, head of strategy at Palisade Research, cited as an example a shutdown-resistance test. Such a benchmark checks whether a system will try to prevent itself from being turned off under certain circumstances.
If a model has been specifically trained to perform well on that particular test, the result may be misleading: the system will display the expected behavior during evaluation but act differently in a real-world situation.
Staudley compared this to Volkswagen’s diesel scandal. The company’s cars were able to detect the conditions of emissions testing and alter engine behavior to produce better results than they did in normal use.
A similar problem is possible with AI: a model does not necessarily have to be safe — it may only need to learn to appear safe at the moment of evaluation.
That is precisely why researchers need access to the broader context: intermediate model versions, work logs, test transcripts, training settings, and the environment in which the system is rewarded for certain actions.
According to Adam Gleave, head of FAR.AI, independent specialists may also need the ability to speak with company employees. That would help verify whether official descriptions of safety processes match what actually happened in practice.
The main threat to independence is company control
Amodei proposed giving evaluators the right to publish key findings on risks, incidents, safety practices, and even on what access they were or were not given. Under his proposal, Anthropic would not have editorial control over such publications.
But researchers believe a declaration alone is not enough.
In practice, outside organizations often work with AI companies under contract. That means they may be subject to strict non-disclosure agreements, time limits, and terms that allow the client to influence the contents of the final report.
Gleave said that FAR.AI has already turned down contracts with several advanced-model developers because they demanded too much control over the evaluation process. In his view, under such conditions the independence of evaluators is put at risk.
The problem can arise at every stage: a company may limit access to certain data, shorten the research period, prohibit evaluation of specific systems, or refuse to allow publication of inconvenient findings.
As a result, an independent auditor risks turning into an ordinary contractor who formally conducts a review but in reality works within rules set by the developer itself.
A week for an investigation and three days for testing
The experience of recent evaluations shows that time is no less important than access.
After the Hugging Face incident, OpenAI gave METR and Redwood researchers about a week to work on-site. Later, both organizations said they were unable to draw confident conclusions — in part because of limits on the scope and timeline of the investigation.
A similar situation arose during testing of GPT-6 Astra before release. According to materials from Apollo Research, the organization was given only three days to evaluate the model.
In its report, Apollo said that because of the model’s ability to recognize that it was being evaluated and the limited testing time, the low frequency of undesirable behavior does not provide serious grounds for judging how well the model is actually aligned with human goals.
In other words, the absence of detected problems does not always mean there are no problems. It may simply be that researchers did not have time to find them, or that the model realized it was being watched.
Why should this time be any different? That is the question independent evaluators are now asking.
Gleave allows that Amodei and Altman may indeed have changed their views on external oversight and may be ready to give researchers more freedom. However, AI companies’ intellectual property is enormously valuable to them. Therefore, in his opinion, developers will most likely be extremely cautious by default when it comes to access and data publication.
Researchers want common rules and the force of law
Several organizations interviewed by TechCrunch believe the industry needs a public and transparent system of rules. It should clearly spell out which organizations can be considered independent evaluators, what access they receive, and which results they are allowed to publish.
Staudley also warns that without uniform standards, companies may choose reviewers based on convenience — for example, favoring those who lack sufficient qualifications or are not inclined to investigate the most serious risks.
Henry Papadatos, executive director of Safer AI, believes that even a public voluntary system will not fully solve the problem. Any voluntary commitments depend on a company’s goodwill, and that can change after a crisis, a leadership change, or the emergence of commercial pressure.
In his view, legal requirements are needed to compel developers to provide independent organizations with access to models and the processes by which they are created. That would also prevent some companies from following the rules while others avoid external oversight.
Some elements of such a system are already beginning to appear.
In California, SB 53 requires large advanced-model developers to publish safety frameworks and report critical incidents. Another law, SB 813, passed in September, creates a system of state-recognized independent organizations specializing in AI risk assessment.
In the European Union, the AI Act requires advanced-model developers to conduct and document evaluations, organize adversarial testing, and report serious incidents. The EU AI Office may also carry out its own inspections and involve independent experts.
However, current rules still do not require the kind of deep and ongoing access that Amodei is proposing. In many cases, the decision on how thoroughly to admit outside specialists into internal processes still remains with the companies themselves.
Who has already agreed, and who is still silent
Anthropic and OpenAI have publicly backed the idea of embedded independent evaluators. Meanwhile, Meta, SpaceXAI, and Google DeepMind have not yet made similar commitments.
DeepMind CEO Demis Hassabis proposed a different option — creating a separate industry organization that would handle independent testing of advanced models. At the same time, Google, OpenAI, and Anthropic have reportedly already spent several weeks discussing possible approaches to AI safety.
It is still unknown which organizations Anthropic and OpenAI will work with, when their permanent presence inside the companies will begin, how many evaluators will be involved, and what data they will be able to study. The rules for publishing results have not been defined either.
These details are crucial. Independent evaluation implies not only access to information, but also the right to report problems without prior coordination with the developer.
If a company retains the ability to choose convenient reviewers, limit their work, and block inconvenient conclusions, such a system may become just another form of corporate control.
Voluntary commitments can be a useful start. But trust in the safety of advanced models cannot be built solely on developers’ own promises. As Papadatos put it, companies cannot simultaneously demand full freedom to set their own safety rules and ask the public to believe that they are truly following them.






