Anthropic CEO Dario Amodei has suggested embedding third-party safety evaluators within AI companies to monitor the development of advanced models, following warnings from a former researcher about the risks of uncontrolled AI systems.
However, experts like Julie Andersen Hill argue that without the ability to enforce safety measures, such as halting model training or deployment, these evaluators may not effectively mitigate risks. While Amodei cites the banking industry as a precedent for this regulatory approach, the comparison is flawed due to the lack of enforcement power for the evaluators.
Critics also express concern that this proposal could serve to protect established companies like Anthropic and OpenAI from competition and liability, as they retain control over what evaluators can access and report. The interconnected nature of the AI evaluation field further complicates the notion of independence, as evaluators may have ties to the companies they assess.
Overall, the effectiveness of these proposed safety measures remains uncertain, which could influence investor sentiment and regulatory scrutiny in the AI sector