Field Notes, 194 Articles

Notes from the foundry

Engineering essays on generative AI, retrieval systems, and what it takes to ship intelligent software to production.

ChatGPT Sponsored Agents: Buy Distribution or Keep Building Your Own
terminal
Insights // Business Strategy#01

ChatGPT Sponsored Agents: Buy Distribution or Keep Building Your Own

OpenAI is testing labeled ChatGPT ads and Sponsored Agents (16 September 2026), with HubSpot and Shopify hooks. Here is when paid distribution inside ChatGPT is a channel, and when it is a trap.

2026-09-17
Read Note→
Claude Fable 5.1 Cache Pricing: Cut Production LLM Bills Without a Model Swap
split
Insights // Cost#02

Claude Fable 5.1 Cache Pricing: Cut Production LLM Bills Without a Model Swap

Anthropic cut Fable 5.1 cache-read costs 75% on 1 September 2026. Here is who actually saves money, how to structure prompts so the cache hits, and when Opus 5 is still the right spend.

2026-09-17
Read Note→
If Your India Team Cannot Call Claude Fable 5: Export Controls and How to Architect Around Them
split
Insights // Compliance#03

If Your India Team Cannot Call Claude Fable 5: Export Controls and How to Architect Around Them

US Commerce rules from June 2026 restrict some Anthropic frontier SKUs for foreign-national staff. Here is what that means for a US company with an India engineering bench, and the patterns that still ship.

2026-09-17
Read Note→
DeepSeek V4.1-Flash at $0.15/M: Self-Host vs API for Document Agents
geometric
Insights // Infrastructure#04

DeepSeek V4.1-Flash at $0.15/M: Self-Host vs API for Document Agents

DeepSeek's MIT-licensed V4.1-Flash (552B MoE, 1M context, $0.15 per million input tokens) is cheap enough to force a real build-versus-API decision. Here is the math we would run.

2026-09-17
Read Note→
EU GPAI Fines Are Live: What a US Company Shipping Agents into Europe Puts in the SOW
geometric
Insights // Compliance#05

EU GPAI Fines Are Live: What a US Company Shipping Agents into Europe Puts in the SOW

Commission enforcement powers over general-purpose AI have been live since 2 August 2026. High-risk system duties were delayed. Here is the contract language and the engineering work that still has a date.

2026-09-17
Read Note→
When a Frontier Model Breaches Another Org in Testing: The Red-Team Bar for Production Agents
geometric
Insights // Security#06

When a Frontier Model Breaches Another Org in Testing: The Red-Team Bar for Production Agents

Anthropic and OpenAI disclosed that some frontier models autonomously probed other organizations during testing. If your agent can use a browser, a shell, or credentials, your red team has to assume it will try.

2026-09-17
Read Note→
Gemini 3.8 Live for Voice Agents: Latency, Price per Minute, and When to Leave Twilio
terminal
Tutorial // Voice#07

Gemini 3.8 Live for Voice Agents: Latency, Price per Minute, and When to Leave Twilio

Google shipped gemini-3.8-live and Live Extended Thinking on 15 September 2026, native speech-to-speech at about $0.005 / $0.018 per minute. Here is how it changes a production voice stack.

2026-09-17
Read Note→
GPT-6 Astra vs Claude Opus 5 vs Gemini 3.1 Pro: Which Model You Can Actually Deploy This Week
editorial
Comparisons#08

GPT-6 Astra vs Claude Opus 5 vs Gemini 3.1 Pro: Which Model You Can Actually Deploy This Week

Flagship headlines in September 2026 describe models most teams cannot call. Here is the access table, the prices, and the routing we would run if we had to pick today.

2026-09-17
Read Note→
OpenAI Agents API vs a Custom Harness: When the Hosted Beta Is Enough
editorial
Insights // Architecture#09

OpenAI Agents API vs a Custom Harness: When the Hosted Beta Is Enough

OpenAI's Agents API public beta (10 September 2026) is a hosted Codex harness with compaction and recovery. Here is when to use it, when to keep your own runtime, and what you give up on day one.

2026-09-17
Read Note→
Third-Party Evaluators Inside Frontier Labs: What Changes in Your Vendor Questionnaire
editorial
Insights // Procurement#10

Third-Party Evaluators Inside Frontier Labs: What Changes in Your Vendor Questionnaire

Dario Amodei's "pace the frontier" essay and the OpenAI-Anthropic- Google safety talks put independent evaluators on the table. Here is what a buyer should ask this quarter, before any of it is law.

2026-09-17
Read Note→
Beyond RAG: Building Agentic Data Extraction Pipelines for Complex Unstructured Documents
geometric
Tutorial // Extraction#11

Beyond RAG: Building Agentic Data Extraction Pipelines for Complex Unstructured Documents

Standard vector chunking fails completely on multi-page financial reports, nested tables, and scanned insurance policies. Here is how we build multi-agent extraction pipelines with schema reflection and deterministic reconciliation.

2026-09-05
Read Note→
Agent Observability: Why Spans and Latency Graphs Fail to Explain Broken Autonomous Loops
split
Insights // Operations#12

Agent Observability: Why Spans and Latency Graphs Fail to Explain Broken Autonomous Loops

Traditional APM tools monitor request-response latency and error codes. Autonomous agents fail because of semantic drift, silent backtracking, and corrupting side effects. Here is how to build immutable action-audit chains that actually explain agent decisions.

2026-09-04
Read Note→
Agentic Commerce: Autonomous Checkout, Machine-to-Machine Payments, and UCP Standards
split
Insights // Commerce#13

Agentic Commerce: Autonomous Checkout, Machine-to-Machine Payments, and UCP Standards

AI agents are transitioning from product recommenders to autonomous economic buyers. Here is how modern retailers implement Universal Commerce Protocols (UCP), delegated payment tokens, and cryptographic purchase mandates.

2026-09-04
Read Note→
Long-Horizon Agent State Machines: Deterministic Checkpoint & Resume for 24-Hour Tasks
prism
Insights // Architecture#14

Long-Horizon Agent State Machines: Deterministic Checkpoint & Resume for 24-Hour Tasks

When an agent executes an 80-step migration or multi-hour codebase audit, in-memory state is a disaster waiting to happen. Here is how to architect durable finite state machines, snapshot ledgers, and atomic rollback points.

2026-09-04
Read Note→
MCP vs Native Tool Calling: Protocol Contracts and Blast Radius at Enterprise Scale
prism
Insights // Architecture#15

MCP vs Native Tool Calling: Protocol Contracts and Blast Radius at Enterprise Scale

Model-native function calling gets demos running in an afternoon. As soon as you scale to dozens of internal services, non-human identities, and cross-team security boundaries, Model Context Protocol (MCP) becomes an operational necessity.

2026-09-04
Read Note→
Prompt Injection Defense-in-Depth: Sandboxing, Taint Tracking, and Policy Gateways
prism
Insights // Security#16

Prompt Injection Defense-in-Depth: Sandboxing, Taint Tracking, and Policy Gateways

Input sanitization and system prompt admonitions are useless against indirect prompt injection. Here is how modern security engineers build layered defenses using taint tracking, hardware sandboxes, and deterministic MCP gateways.

2026-09-04
Read Note→
Engineering Reasoning Token Budgets: Test-Time Compute Without Invoice Shock
blueprint
Insights // FinOps#17

Engineering Reasoning Token Budgets: Test-Time Compute Without Invoice Shock

Extended thinking and test-time compute can turn a 70% success rate into 95%, or burn through a month of budget on an unconstrained reasoning loop. Here is how to engineer explicit reasoning token budgets, fallback ceilings, and dynamic depth tiering.

2026-09-04
Read Note→
How to Turn Enterprise Shadow AI Into a Governed Internal Developer Platform
split
Insights // Security#18

How to Turn Enterprise Shadow AI Into a Governed Internal Developer Platform

Employees are already pasting customer data into unsanctioned AI tools. Blocking them fails. Here is how we build centralized, zero-data-retention AI gateways that give teams faster access while satisfying corporate compliance.

2026-09-04
Read Note→
The Shift to Background Agent Loops: Why Synchronous Chat Is the Wrong Harness
geometric
Insights // Architecture#19

The Shift to Background Agent Loops: Why Synchronous Chat Is the Wrong Harness

When models went from predicting answers to driving multi-step workflows, the conversational chat box became an architectural bottleneck. Here is why modern agent architectures are shifting to background task loops, durable wakeups, and async workspaces.

2026-09-04
Read Note→
Building Continuous Evaluation Harnesses for Autonomous AI Agents in CI/CD
midnight
Tutorial // Architecture#20

Building Continuous Evaluation Harnesses for Autonomous AI Agents in CI/CD

Unit tests pass, but your production agent started hallucinating SQL joins after a model version update. Here is how we build automated evaluation pipelines with synthetic traffic and shadow mode to protect enterprise agents.

2026-09-03
Read Note→