Daily AI Zine
Wednesday, August 19, 2026
Issue No. 034 · Tokyo · full edition

✦ Today's Big Thing

Codex turns migration debt into scoped work

The useful shift today is not a new model. It is a concrete example of agentic coding being applied to old engineering work with cost and time boundaries.

5 min read · 7 sections

In Brief
  1. Asana used Codex to replace an outdated testing system in two weeks, a job OpenAI says had been estimated at five years.
  2. Anthropic is putting more attention on containment, meaning limits that reduce how much damage an agent can cause if it acts in the wrong place.
  3. GitHub Copilot keeps widening model choice across editors, the command line, and the Copilot app.
  4. Cloudflare’s Kitesurf points to browsers built for agents rather than humans.

Today's Big Thing

The one thing that matters

Test it

Hot signal: Codex compresses migration debt into budget-sized chunks

OpenAI says Asana used Codex to replace an outdated testing system in two weeks, completing work that had been expected to take five years for about $12K. Treat this as a useful shift, not proof that every old codebase can be cleared by an agent. The important pattern is narrower: stale internal engineering work can be broken into bounded Codex jobs with a measurable budget, a defined target, and a clear before-and-after test. For Adrian’s stack, the action is to stop treating technical debt as background cleanup and start packaging it as a short queue of agent-executable migrations, each with a pass/fail outcome.

My AI Ecosystem

Your actual stack

Use it

Anthropic centers agent containment across products

Anthropic’s engineering blog is now foregrounding containment, meaning limits that cap an agent’s blast radius, or the damage it can cause if it acts in the wrong account, file, or tool. The summary is thin, so the practical takeaway is simple: stronger agents need narrower rooms. For daily agent work, permissions, tool scope, and task boundaries matter as much as prompt quality.

Test it

Copilot keeps widening model choice across the coding lane

GitHub’s August 10 Copilot weekly release points to new models, portable plugins, and smoother agent workflows across editors, the command line, and the Copilot app. Portable plugins are add-ons that can travel across more than one coding surface. The useful shift is operational: model choice is becoming a repo-level workflow decision, not a personal editor preference.

Watch it

Cloudflare’s Kitesurf treats browsing as agent infrastructure

Cloudflare introduced Kitesurf, an agent-first browser that runs on Workers inside V8 isolates, which are lightweight separated runtimes used to run code safely at scale. The interesting part is not a better human browser. It is a browser designed for agents that need repeated web actions without dragging around a desktop session. This is a watch item for research, monitoring, and source-gathering workflows.

CoWork Corner

Claude CoWork, day to day

CoWork: put containment before capability

Claude CoWork is the workspace Adrian uses for persistent projects, operational records, and AI-assisted workflows. The useful move today is procedural, not a new CoWork product change. Anthropic’s containment framing is a reminder that every persistent workspace needs a narrow operating room: name the task, name the allowed files or records, name the tools the agent may touch, and name the stop condition. Revisit one active room and add a short “allowed surface” note before giving it another broad instruction.

GPT Desk

OpenAI, ChatGPT, Codex

Test it

Use Codex for bounded migration work, not vague cleanup

What changed: OpenAI reports that Asana used Codex to replace an outdated testing system in two weeks, with the work previously estimated at five years and costing about $12K. Why Adrian cares: Grey Group OS, Goodsense, and production workflow tools likely contain small pieces of old process debt that never become urgent enough for a normal sprint. Do this today: pick one stale test, script, or data-cleanup job, write a pass/fail target, and give Codex only that lane.

Use it

Turn OpenAI’s cyber pacing note into a Codex task gate

What changed: OpenAI says it is strengthening monitoring, alignment, and security for frontier models as cyber-critical capabilities improve. Cyber-critical means a model could affect security-sensitive systems or workflows. Why Adrian cares: Codex jobs for Grey Group OS or client-facing production infrastructure should not all receive the same autonomy. Do this today: add one required field to the Codex task template: “Could this touch credentials, deployment, payments, private files, or external systems?” If yes, require human approval before execution.

Small Money Systems

Small, repeatable, real

Test it

Small system: two-week test-system replacement offer

System: Build a small Codex-powered service that finds one outdated test, build, or reporting workflow and replaces it with a cleaner version. Customer: Japan-facing founders, agencies, or production vendors with old internal tooling. Offer: a scoped repo review, one replacement plan, one Codex implementation pass, and a short handoff note. Price: ¥120,000 for the first fixed-scope package. Existing assets: Grey Group OS, Codex habits, GitHub workflows, and proposal templates. AI workflow: Codex drafts and edits, Claude Code reviews the change, and ChatGPT turns the result into a client-readable report. First action: list three old scripts or tests inside Grey Group OS and choose the easiest one. Repeatability: each finished job becomes a before-and-after proof for the next buyer. Effort: half day. Expected value: small revenue, cleaner internal tooling, and a reusable sales proof.

Test it

Small system: agent-browser research pack for activations

System: Build a repeatable source-gathering pack that uses an agent browser workflow to collect venue, sponsor, competitor, and local-market references for a campaign. Customer: brands, agencies, and visiting producers planning Japan work. Offer: a short research memo, checked source list, and practical activation angles. Price: ¥45,000 per snapshot. Existing assets: Street Attack Japan, production planning files, Japan-English briefing experience, and existing deck formats. AI workflow: a browser-agent pipeline inspired by Kitesurf-style stateless browsing gathers sources, then ChatGPT and Claude summarize and clean the client memo. First action: make one sample snapshot for a current Tokyo category. Repeatability: the same template can run per district, category, sponsor, or event. Effort: one hour. Expected value: lead generation and small paid research work without building a large product first.

Build Next

Deployable now

Use it

Build a preflight breaker for agent tasks

What to build: a small task-intake form that blocks Claude Code or Codex from starting until the user names the allowed files, tools, credentials, and stop condition. Why now: Anthropic is explicitly framing containment as a product engineering problem for stronger agents. Effort: one hour. Expected impact: fewer unsafe agent runs and less cleanup after over-broad instructions. Dependencies: an existing repo task template, a place to store task notes, and agreement that blocked tasks must be rewritten before execution.

Try This Today

One action, right now

Pick one stale test, script, or deploy chore. Write the current pain, target behavior, allowed files, stop condition, and pass/fail check. Then ask Codex for a plan only. No edits until the scope is clean.