---
title: "Third-Party Evaluators Inside Frontier Labs: What Changes in Your Vendor Questionnaire"
description: "Dario Amodei's \"pace the frontier\" essay and the OpenAI-Anthropic- Google safety talks put independent evaluators on the table. Here is what a buyer should ask this quarter, before any of it is law."
image: "https://foundrysoft.co/images/blog-cards/third-party-evaluators-frontier-labs-vendor-questionnaire.png"
url: "https://foundrysoft.co/blog/third-party-evaluators-frontier-labs-vendor-questionnaire"
---

Insights // Procurement 2026-09-17 10 min read

# Third-Party Evaluators Inside Frontier Labs: What Changes in Your Vendor Questionnaire

Dario Amodei's "pace the frontier" essay and the OpenAI-Anthropic- Google safety talks put independent evaluators on the table. Here is what a buyer should ask this quarter, before any of it is law.

![Varun Raj Manoharan](https://foundrysoft.co/images/about/founder.webp)

Varun Raj Manoharan Founder & Principal Engineer

Governance Evals Compliance Security AI Strategy Enterprise AI

## Key takeaways

-   In mid-September 2026 Anthropic's CEO called for permanent third- party evaluator access inside frontier labs. OpenAI, Google, and xAI publicly backed the direction. Nothing in that is a contract you can rely on yet.
-   Add three questions to every model-vendor packet: who evaluates, what they can see, and whether you get the report.
-   Lab-side evaluators do not replace your product evals. They are evidence about the foundation model, not about your agent.
-   Antitrust noise around lab coordination is a reason to keep multi-vendor routing, not a reason to pause procurement.

## In this article

1.  01 [What this is not](#what-this-is-not)
2.  02 [Questions to add to the packet](#questions-to-add-to-the-packet)
3.  03 [How this changes a scoring sheet](#how-this-changes-a-scoring-sheet)
4.  04 [What we would do on a live project](#what-we-would-do-on-a-live-project)
5.  05 [Frequently Asked Questions](#frequently-asked-questions)

On 12 September 2026 Dario Amodei published a long essay arguing that frontier labs should slow capability enough to understand what they are shipping, and that independent evaluators should have employee-like access: models, training process, incident data. Sam Altman said OpenAI is willing. Elon Musk said xAI is willing. Demis Hassabis had already been floating a FINRA-style standards body. By 15 September OpenAI's policy lead confirmed the three labs had been talking for weeks, and also said they do not think they need an antitrust waiver.

None of that is a standard you can attach to a purchase order. It is a signal that "who watches the model lab" is about to become a checkbox in serious RFPs. You might as well write the checkbox now, while the labs are still competing to look cooperative.

## What this is not

It is not your eval suite. A third party sitting inside Anthropic does not know whether _your_ refund agent follows _your_ policy. That is still [your harness and your traffic](https://foundrysoft.co/blog/agent-eval-harness-synthetic-traffic-ci-cd).

It is not regulation. US federal policy is still largely voluntary. EU GPAI enforcement is aimed at providers, with you as deployer. See [GPAI fines and the SOW](https://foundrysoft.co/blog/eu-gpai-fines-sow-for-us-companies-shipping-agents).

It is not a reason to single-source. If the labs are coordinating on safety, they are also still competing on price and access. Keep two providers in the harness. Coordination plus a single vendor is how you inherit both groupthink and an invoice.

## Questions to add to the packet

Steal these. Require written answers, not a sales call.

1.  **Named evaluators.** Which independent orgs currently have access to this SKU, under what contract, and can we see the latest summary report (redacted is fine).
2.  **Scope of access.** Weights, training data summaries, incident tickets, production sampling, or only a closed eval harness? "Employee-like" is doing a lot of work in the essays. You want the list.
3.  **Incident duty to customers.** When a frontier model is caught doing something like autonomous probing of another org (the other story this month), what do you send to deployers, and in how many hours.
4.  **Your data in their evals.** If they sample production-like traffic, is customer content in it. Opt out, in writing.
5.  **SKU mapping.** Evaluators on Mythos do not automatically cover the Fable API you actually call. Ask per SKU.
6.  **Export and nationality.** Who on the eval team is allowed to see the model. This collides with [foreign-national access rules](https://foundrysoft.co/blog/claude-fable-5-export-controls-india-engineering-team).

If the vendor cannot answer (1) and (3), they are not "ahead on safety." They are ahead on blog posts.

## How this changes a scoring sheet

| Old question | New question |
| --- | --- |
| Do you have a model card? | Who last independently tested this SKU, and can we read the summary? |
| Do you red-team? | Is the red team yours, a retained third party, or both? What did they find in the last 90 days that changed a ship decision? |
| SOC 2? | SOC 2 plus: production-sampling policy for evals, and a customer-notification SLA for capability incidents. |
| Roadmap? | What happens to our contract if a FINRA-style body delays a SKU you sold us. |

You still score latency, price, and evals on _your_ tasks first. Safety theater that fails your extraction test is still a failed vendor.

## What we would do on a live project

Keep the [48-hour shadow](https://foundrysoft.co/blog/evaluate-new-llm-48-hours-shadow-mode). Add a folder in the vendor pack called `lab-evaluators/` with whatever PDFs they will send. Revisit it at renewal. Do not block a Q4 ship on a standards body that does not exist yet.

If a lab offers to embed evaluators and also asks you to _be_ a design partner, that is a separate agreement. Do not mix it with the production DPA.

## Frequently Asked Questions

**Should we wait for the standards body before picking a lab?** No. Pick on available SKUs, price, and your evals. Write the evaluator questions into the contract so you can exit or reroute if the body, when it exists, flags a SKU.

**Is lab coordination an antitrust problem for us?** That is their lawyers versus DOJ. Your mitigation is multi-vendor routing so you are not stuck if a collaboration becomes a constraint.

**Do open-weight models skip this?** You become closer to the "provider" role if you serve them. You still want third-party evals, they just might be _your_ retained firm, not Anthropic's. Budget for that.

**Will this raise prices?** Probably, if it is real. Independent access is not free. Ask for the line item rather than discovering it inside a 12% "safety" uplift.

---

_FoundrySoft helps buyers write vendor packets that match how agents actually fail. See [AI governance](https://foundrysoft.co/services/ai-governance) and [evals](https://foundrysoft.co/services/vercel-ai-sdk-evaluations), or [contact us](https://foundrysoft.co/contact)._

Interactive Engineering Calculators Free Tools

### Estimate your project cost, token budget, and automation ROI

We built free, production-calibrated tools to help engineering leaders forecast token consumption, compare build vs buy scenarios, and audit code security.

[Automation ROI Calculator →](https://foundrysoft.co/tools/automation-roi) [Project Cost Estimator →](https://foundrysoft.co/tools/project-cost-estimator) [Build vs Buy Calculator →](https://foundrysoft.co/tools/build-vs-buy) [Security Code Audit →](https://foundrysoft.co/tools/code-audit)

#### Work with us on this

[AI Governance Platform

Model inventories, policy gates, and audit trails for your AI systems. Track drift and bias with logging built for the EU AI Act and ISO 42001 reviews.

](https://foundrysoft.co/services/ai-governance)[AgentOps

Run AI agents in production with telemetry, regression evals, and guardrails. We add observability, prompt versioning, and one-click rollbacks before launch.

](https://foundrysoft.co/services/agentops)[Vercel AI SDK Output Evaluations

Stop guessing if your AI is improving. We implement automated evaluation pipelines to score Vercel AI SDK outputs for accuracy, tone, and relevance.

](https://foundrysoft.co/services/vercel-ai-sdk-evaluations)

#### Related reading

[ChatGPT Sponsored Agents: Buy Distribution or Keep Building Your Own

OpenAI is testing labeled ChatGPT ads and Sponsored Agents (16 September 2026), with HubSpot and Shopify hooks. Here is when paid distribution inside ChatGPT is a channel, and when it is a trap.

OpenAI AI Strategy AI Agents

](https://foundrysoft.co/blog/chatgpt-sponsored-agents-vs-building-your-own)[If Your India Team Cannot Call Claude Fable 5: Export Controls and How to Architect Around Them

US Commerce rules from June 2026 restrict some Anthropic frontier SKUs for foreign-national staff. Here is what that means for a US company with an India engineering bench, and the patterns that still ship.

Claude Anthropic Compliance

](https://foundrysoft.co/blog/claude-fable-5-export-controls-india-engineering-team)[EU GPAI Fines Are Live: What a US Company Shipping Agents into Europe Puts in the SOW

Commission enforcement powers over general-purpose AI have been live since 2 August 2026. High-risk system duties were delayed. Here is the contract language and the engineering work that still has a date.

EU AI Act Compliance Governance

](https://foundrysoft.co/blog/eu-gpai-fines-sow-for-us-companies-shipping-agents)

#### Next Article

[

Beyond RAG: Building Agentic Data Extraction Pipelines for Complex Unstructured Documents

](https://foundrysoft.co/blog/post-rag-agentic-data-extraction-unstructured-docs)

Available for new projects

## Let's build something great.

Have a project in mind? We are an elite software and AI development studio ready to bring your ideas to production. Let's talk about your roadmap.

Start a Project [See our work](https://foundrysoft.co/work)

```json
{
  "@context": "https://schema.org",
  "@type": "Organization",
  "name": "FoundrySoft",
  "url": "https://foundrysoft.co",
  "logo": "https://foundrysoft.co/logo.svg",
  "description": "FoundrySoft builds production-grade software and AI systems for US companies, from an India-based team of senior engineers.",
  "sameAs": [
    "https://github.com/foundrysofthq",
    "https://www.linkedin.com/company/foundrysoft",
    "https://www.instagram.com/foundrysoft/"
  ]
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "WebSite",
  "name": "FoundrySoft",
  "url": "https://foundrysoft.co"
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "TechArticle",
  "headline": "Third-Party Evaluators Inside Frontier Labs: What Changes in Your Vendor Questionnaire",
  "description": "Dario Amodei's \"pace the frontier\" essay and the OpenAI-Anthropic- Google safety talks put independent evaluators on the table. Here is what a buyer should ask this quarter, before any of it is law.",
  "url": "https://foundrysoft.co/blog/third-party-evaluators-frontier-labs-vendor-questionnaire",
  "mainEntityOfPage": "https://foundrysoft.co/blog/third-party-evaluators-frontier-labs-vendor-questionnaire",
  "image": [
    "https://foundrysoft.co/images/blog-cards/third-party-evaluators-frontier-labs-vendor-questionnaire.png"
  ],
  "datePublished": "2026-09-17",
  "dateModified": "2026-09-17",
  "keywords": "Governance, Evals, Compliance, Security, AI Strategy, Enterprise AI",
  "author": {
    "@type": "Person",
    "name": "Varun Raj Manoharan",
    "jobTitle": "Founder & Principal Engineer",
    "url": "https://foundrysoft.co/about",
    "sameAs": [
      "https://www.linkedin.com/in/varunrajmanoharan",
      "https://github.com/varun-raj"
    ]
  },
  "publisher": {
    "@type": "Organization",
    "name": "FoundrySoft",
    "logo": {
      "@type": "ImageObject",
      "url": "https://foundrysoft.co/logo.svg"
    }
  }
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "BreadcrumbList",
  "itemListElement": [
    {
      "@type": "ListItem",
      "position": 1,
      "name": "Home",
      "item": "https://foundrysoft.co/"
    },
    {
      "@type": "ListItem",
      "position": 2,
      "name": "Blog",
      "item": "https://foundrysoft.co/blog"
    },
    {
      "@type": "ListItem",
      "position": 3,
      "name": "Third-Party Evaluators Inside Frontier Labs: What Changes in Your Vendor Questionnaire",
      "item": "https://foundrysoft.co/blog/third-party-evaluators-frontier-labs-vendor-questionnaire"
    }
  ]
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "FAQPage",
  "mainEntity": [
    {
      "@type": "Question",
      "name": "Should we wait for the standards body before picking a lab?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "No. Pick on available SKUs, price, and your evals. Write the evaluator questions into the contract so you can exit or reroute if the body, when it exists, flags a SKU."
      }
    },
    {
      "@type": "Question",
      "name": "Is lab coordination an antitrust problem for us?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "That is their lawyers versus DOJ. Your mitigation is multi-vendor routing so you are not stuck if a collaboration becomes a constraint."
      }
    },
    {
      "@type": "Question",
      "name": "Do open-weight models skip this?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "You become closer to the \"provider\" role if you serve them. You still want third-party evals, they just might be your retained firm, not Anthropic's. Budget for that."
      }
    },
    {
      "@type": "Question",
      "name": "Will this raise prices?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Probably, if it is real. Independent access is not free. Ask for the line item rather than discovering it inside a 12% \"safety\" uplift. --- FoundrySoft helps buyers write vendor packets that match how agents actually fail. See AI governance and evals, or contact us."
      }
    }
  ]
}
```
