top of page

Embedded Assessments for Frontier AI

vor 9 Minuten
1 Min. Lesezeit

Bringing independent oversight inside frontier AI companies


Cover image: Stock photograph from Pexels, used under a free licence.
Cover image: Stock photograph from Pexels, used under a free licence.

Most independent evaluations of frontier AI test models from the outside, through an API, before release. But many of the most serious risks now come from how developers use their own models internally. Recent incidents show this clearly: internal AI agents have evaded monitoring, escaped sandboxes and compromised external systems before receiving any meaningful independent scrutiny.


A new paper examines embedded assessments as a promising answer. In this model, independent evaluators receive employee-like access to a developer's internal systems, staff and documentation, often working onsite under strict security controls. Similar resident inspection models are already standard in nuclear power, banking and food safety.

Following early pilots, and recent commitments by the CEOs of Anthropic and OpenAI, the paper sets out seven key design questions and offers a concrete starting point.


The authors recommend that assessments:

  • cover at least internal agent monitoring, agent security controls and permissions, and model alignment

  • run continuously, with targeted investigations after serious incidents

  • result in detailed public reports at least quarterly, with only narrowly defined redactions

  • include clear escalation routes for serious or unresolved concerns


Read the full paper:



 
 

Pour Demain

Europabüro

Clockwise

Avenue des Arts - Kunstlaan 44

1040 Brüssel

Belgien

Büro Bern

Marktgasse 46

3011 Bern

(Postanschrift: c/o ExpertFid & Audit AG, Zweigniederlassung Basel, Marktgasse 8, 4051 Basel)

Kontakt

Folgen Sie uns auf:

  • 2
  • 4
bottom of page