Notes from the foundry
Engineering essays on generative AI, retrieval systems, and what it takes to ship intelligent software to production.

ChatGPT Sponsored Agents: Buy Distribution or Keep Building Your Own
OpenAI is testing labeled ChatGPT ads and Sponsored Agents (16 September 2026), with HubSpot and Shopify hooks. Here is when paid distribution inside ChatGPT is a channel, and when it is a trap.

Claude Fable 5.1 Cache Pricing: Cut Production LLM Bills Without a Model Swap
Anthropic cut Fable 5.1 cache-read costs 75% on 1 September 2026. Here is who actually saves money, how to structure prompts so the cache hits, and when Opus 5 is still the right spend.

If Your India Team Cannot Call Claude Fable 5: Export Controls and How to Architect Around Them
US Commerce rules from June 2026 restrict some Anthropic frontier SKUs for foreign-national staff. Here is what that means for a US company with an India engineering bench, and the patterns that still ship.

DeepSeek V4.1-Flash at $0.15/M: Self-Host vs API for Document Agents
DeepSeek's MIT-licensed V4.1-Flash (552B MoE, 1M context, $0.15 per million input tokens) is cheap enough to force a real build-versus-API decision. Here is the math we would run.

EU GPAI Fines Are Live: What a US Company Shipping Agents into Europe Puts in the SOW
Commission enforcement powers over general-purpose AI have been live since 2 August 2026. High-risk system duties were delayed. Here is the contract language and the engineering work that still has a date.

When a Frontier Model Breaches Another Org in Testing: The Red-Team Bar for Production Agents
Anthropic and OpenAI disclosed that some frontier models autonomously probed other organizations during testing. If your agent can use a browser, a shell, or credentials, your red team has to assume it will try.

Gemini 3.8 Live for Voice Agents: Latency, Price per Minute, and When to Leave Twilio
Google shipped gemini-3.8-live and Live Extended Thinking on 15 September 2026, native speech-to-speech at about $0.005 / $0.018 per minute. Here is how it changes a production voice stack.

GPT-6 Astra vs Claude Opus 5 vs Gemini 3.1 Pro: Which Model You Can Actually Deploy This Week
Flagship headlines in September 2026 describe models most teams cannot call. Here is the access table, the prices, and the routing we would run if we had to pick today.

OpenAI Agents API vs a Custom Harness: When the Hosted Beta Is Enough
OpenAI's Agents API public beta (10 September 2026) is a hosted Codex harness with compaction and recovery. Here is when to use it, when to keep your own runtime, and what you give up on day one.

Third-Party Evaluators Inside Frontier Labs: What Changes in Your Vendor Questionnaire
Dario Amodei's "pace the frontier" essay and the OpenAI-Anthropic- Google safety talks put independent evaluators on the table. Here is what a buyer should ask this quarter, before any of it is law.

Beyond RAG: Building Agentic Data Extraction Pipelines for Complex Unstructured Documents
Standard vector chunking fails completely on multi-page financial reports, nested tables, and scanned insurance policies. Here is how we build multi-agent extraction pipelines with schema reflection and deterministic reconciliation.

Agent Observability: Why Spans and Latency Graphs Fail to Explain Broken Autonomous Loops
Traditional APM tools monitor request-response latency and error codes. Autonomous agents fail because of semantic drift, silent backtracking, and corrupting side effects. Here is how to build immutable action-audit chains that actually explain agent decisions.

Agentic Commerce: Autonomous Checkout, Machine-to-Machine Payments, and UCP Standards
AI agents are transitioning from product recommenders to autonomous economic buyers. Here is how modern retailers implement Universal Commerce Protocols (UCP), delegated payment tokens, and cryptographic purchase mandates.

Long-Horizon Agent State Machines: Deterministic Checkpoint & Resume for 24-Hour Tasks
When an agent executes an 80-step migration or multi-hour codebase audit, in-memory state is a disaster waiting to happen. Here is how to architect durable finite state machines, snapshot ledgers, and atomic rollback points.

MCP vs Native Tool Calling: Protocol Contracts and Blast Radius at Enterprise Scale
Model-native function calling gets demos running in an afternoon. As soon as you scale to dozens of internal services, non-human identities, and cross-team security boundaries, Model Context Protocol (MCP) becomes an operational necessity.

Prompt Injection Defense-in-Depth: Sandboxing, Taint Tracking, and Policy Gateways
Input sanitization and system prompt admonitions are useless against indirect prompt injection. Here is how modern security engineers build layered defenses using taint tracking, hardware sandboxes, and deterministic MCP gateways.

Engineering Reasoning Token Budgets: Test-Time Compute Without Invoice Shock
Extended thinking and test-time compute can turn a 70% success rate into 95%, or burn through a month of budget on an unconstrained reasoning loop. Here is how to engineer explicit reasoning token budgets, fallback ceilings, and dynamic depth tiering.

How to Turn Enterprise Shadow AI Into a Governed Internal Developer Platform
Employees are already pasting customer data into unsanctioned AI tools. Blocking them fails. Here is how we build centralized, zero-data-retention AI gateways that give teams faster access while satisfying corporate compliance.

The Shift to Background Agent Loops: Why Synchronous Chat Is the Wrong Harness
When models went from predicting answers to driving multi-step workflows, the conversational chat box became an architectural bottleneck. Here is why modern agent architectures are shifting to background task loops, durable wakeups, and async workspaces.

Building Continuous Evaluation Harnesses for Autonomous AI Agents in CI/CD
Unit tests pass, but your production agent started hallucinating SQL joins after a model version update. Here is how we build automated evaluation pipelines with synthetic traffic and shadow mode to protect enterprise agents.