Vol. 1 · Edition 039Free · No paywall

Everyone Needs a Samwise

AI news · Synthesized · Opinionated · 🌿

$2B+
combined commitment
over five years
Safety
By Sam Taylor with Samwise

On employee-level model access, the $1 billion each commitment, why external evaluation has always been limited, and whether funding your own inspector delivers the accountability it promises.

Anthropic's new safety inspector works from inside the building. That's the whole point — and the whole problem.

Source lean on this story
▲ avg

Anti-AI

00

Skeptic

02

Neutral

02

Pro (practical)

00

Pro (hyped)

00

← Anti-AI · Pro-AI →

The thing about external AI safety evaluation is that it's always happened too late. An evaluator gets handed a finished model, runs a battery of tests, and publishes a report. By then, the training decisions that shaped the model are months old. You can catch what the model does. You can't catch what the company decided.

Anthropic announced this week that they're changing that arrangement. Accenture's Faculty division — Accenture's specialist AI practice — is becoming an "embedded evaluator" at Anthropic. Unlike external evaluators, embedded evaluators will work inside the lab, with access comparable to an employee's. They watch models take shape in training. They follow the decisions that govern how those models are built and deployed. They can speak directly to employees. Anthropic and Accenture each expect to invest at least $1 billion in building capacity in this area over the next five years.

$1B
What each party expects to invest — Anthropic AND Accenture, separately — over five years

→ Source: Anthropic

The work Faculty will do: evaluating and red-teaming models, conducting alignment assessments, testing model safeguards. The scope is wide. The access is wide. And Anthropic says the arrangement is non-exclusive — more evaluators are coming, and Accenture will work with other AI developers in similar capacities.

Source spread

  • Anthropic — Partnering with Accenture on embedded evaluation [safety] — primary source; Anthropic's own framing of the concept and the funding.
  • Accenture / Faculty [builder] — Faculty is Accenture's AI specialist practice; they work with enterprises and governments on AI deployment, which shapes their evaluator perspective.
  • METR evaluations [safety] — the nonprofit evaluator Anthropic says is in parallel dialogue for a pilot using its own funding; the Accenture announcement is only part of the picture.

Pros & cons

What's real:

  • External evaluation has always had a fundamental limitation: the evaluator arrives after the model is done. An embedded evaluator who watches training decisions unfold is doing something structurally different. If you're trying to catch problems with how a model is built — not just what it does — that early access is necessary.
  • The scope goes beyond capability testing. Alignment assessments and safeguard testing done by someone inside the process, with context for why decisions were made, produces more useful output than a test battery run against a black box.
  • Faculty brings enterprise deployment context. Accenture works with companies and governments actually using AI at scale. That perspective — what breaks in the field, what enterprises actually encounter — is different from the academic safety lens. Both are useful.
  • Non-exclusive and explicitly encouraging other evaluators to do the same is the right structure for long-term credibility. One embedded evaluator doesn't create accountability; an ecosystem of them might.

What deserves a side-eye:

  • Anthropic is paying Accenture to evaluate Anthropic. The article acknowledges that "there is also no settled system for funding independent evaluation" and says long-term funding should come from pooled or government sources. But the actual arrangement today is direct funding by the company being evaluated. That's a structural conflict of interest, regardless of how rigorous Faculty is.
  • There are currently no standards for what embedded evaluators should have access to, or how they should report their findings. The article says so directly. Anthropic is setting precedents here that other labs may copy, and right now those precedents are being set by one company, on its own terms.
  • The METR parallel pilot — where METR uses their own funding — barely gets a paragraph, but it's actually the more important part of this announcement. An evaluator using independent funding is more credible than one whose check comes from the subject. Bury the lead, Anthropic.
  • "Many of the details about how it will operate are still being worked out" means the accountability mechanism isn't actually in place yet. This is a commitment to build the thing, not the thing itself.

Samwise's take

What builders need to know

  • The accountability gap this is trying to close is real. If you're building on frontier models, whether any external party can actually verify lab safety commitments — not just test model outputs — has been an open question. Embedded evaluation is an attempt to answer it. Track whether it produces public findings.
  • Standards don't exist yet. Anthropic acknowledges there are no standards for embedded evaluator access or reporting. If you're in an industry where AI governance matters (financial services, healthcare, legal), watch for whether these standards emerge and from whom.
  • More evaluators are coming. Anthropic says the partnership is non-exclusive and they'll announce others. The diversity of embedded evaluators — and whether any use independent funding — will determine whether this model actually delivers on its premise.
  • The METR parallel pilot is the important one. An evaluator using its own funding is more credible than one funded by the subject. If METR's pilot launches, read that report first.
  • This is infrastructure for accountability, not accountability itself. Anthropic has made a commitment to build something. The thing isn't built yet. Return to this story when the first report is published.

Further reading

🌿

Liked this? Get the weekly digest.

Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.

Your take

How'd I do on this one?

What did I miss?

Tell Samwise (and Sam).

Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.