Daily AI Zine
Monday, September 21, 2026
Issue No. 066 · Tokyo · full edition

✦ Today's Big Thing

Quiet day, useful agent discipline

No single launch dominates today. The useful move is tighter operating practice around coding agents, context, credentials, and reusable project instructions.

5 min read · 7 sections

In Brief
  1. Anthropic’s agent-eval work points to a practical rule: measure noisy infrastructure before judging model quality.
  2. Claude Code’s AGENTS.md support reduces duplicate instruction files across mixed-agent repos.
  3. A small MCP server pattern shows how agents can branch on typed probabilities instead of free-text guesses.
  4. ChatGPT in Word is worth testing for proposal and deck cleanup where the document itself is the workspace.

Today's Big Thing

The one thing that matters

Use it

Quiet day: agent evals need noise control

The practical lead is not a launch. It is a measurement habit.

No single new item clears the bar as the day’s major development. The useful shift is Anthropic’s focus on quantifying infrastructure noise in agentic coding evals. In plain English: before deciding that one coding agent, model, or prompt is better, separate model performance from flaky test runs, slow tools, changing dependencies, and harness behavior. For Adrian, the takeaway is operational. When Claude Code, Codex, or another agent has a bad run, do not immediately rewrite the prompt. First record the repo state, task, tools, test result, and failure cause. Verdict: use this as the default review habit for agent work.

My AI Ecosystem

Your actual stack

Use it

Claude Code adds AGENTS.md as a fallback instruction file

Claude Code version 2.1.277 reportedly checks AGENTS.md when a folder has no CLAUDE.md. AGENTS.md is a project-instruction file used by agent tooling, so this reduces duplicate setup across mixed Claude Code, Codex, and other agent repos. Treat it as a standardization win, not a prompt rewrite. Verdict: use it for new repos, while leaving existing CLAUDE.md files alone until cleanup time.

Watch it

Netlify keeps coding agents on the platform map

Netlify’s blog index now surfaces a coding-agents topic. The candidate summary gives no product detail, so do not treat this as a feature release. Still, it matters because Adrian’s shipping stack depends on static deploys, previews, and agent-built changes landing cleanly. Verdict: watch for concrete Netlify releases around agent-created deploys, preview review, and rollback safety.

Test it

Hot signal: Claude-readable video is becoming a practical wrapper

A new open-source video wrapper reports a simple pattern: turn video into scene-aware keyframes, timestamps, and local transcription so Claude Code or another LLM can inspect it as readable evidence. This is not a finished production suite, but the capability is immediately relevant to production review, cut notes, and shoot research. Verdict: test it on one non-sensitive clip before trusting it for client work.

CoWork Corner

Claude CoWork, day to day

CoWork can borrow typed decisions from MCP

Claude CoWork is the workspace Adrian uses for persistent projects, operational records, and AI-assisted workflows. A reported MCP server called TypeSafe exposes one decision tool that lets Claude Desktop, Claude Code, and Codex call a model and receive probabilities they can branch on, rather than relying only on free-text judgment. For CoWork, the useful workflow is narrow: use typed decisions for low-risk routing, such as classify this note as sales, production, admin, or research, then require human review before any external action.

GPT Desk

OpenAI, ChatGPT, Codex

Use it

Put ChatGPT in the document, not beside it

ChatGPT is reportedly available in Microsoft Word for drafting from notes, summarizing a document, revising selected text, and adjusting headings or formatting from a sidebar. Adrian should care because Grey Group, Goodsense, SET, and Street Attack Japan all depend on proposals, treatments, recaps, and bilingual business writing. Today’s action: take one existing proposal, open it in Word, and ask ChatGPT to produce a tighter one-page executive version while preserving the original claims. Verdict: use.

Test it

OpenAI’s ad agents belong in one controlled sales mockup

OpenAI describes new AI-powered advertising experiences, including Sponsored Agents, marketer tools, and HubSpot and Shopify integrations. Adrian should care because this points to ads that behave more like guided buying or briefing flows than static banners. Today’s action: make one Goodsense or Street Attack Japan mockup where the agent collects a brief, recommends a package, and hands off to a human, with no automated pricing commitment. Verdict: test.

Small Money Systems

Small, repeatable, real

Use it

Small system: proposal polish retainer

System: build a same-day proposal polish desk for treatments, sponsor decks, and Japan-entry one-pagers. Customer: small production companies, agencies, and founders selling in English and Japanese. Offer: a tightened executive summary, cleaner section structure, risk notes, and a short follow-up email. Price: ¥35,000 per document or ¥90,000 monthly for three. Existing assets: Grey Group OS proposal patterns, Goodsense writing standards, and Adrian’s production vocabulary. AI workflow: ChatGPT in Word drafts and restructures, then Claude Code or Codex can turn repeat checks into a reusable checklist tool. First action: choose one old deck and create the before-after sample today. Repeatability: every finished proposal becomes a reusable pattern. Effort: one hour. Expected value: recurring service revenue and stronger lead conversion without building a new product.

Test it

Small system: shoot footage briefing pack

System: create a quick video-to-brief pack from reference clips, location reels, or event footage. Customer: producers, agencies, and local teams preparing a shoot or recap. Offer: timestamped scene notes, visible production risks, interview pull-quotes where audio exists, and a one-page creative summary. Price: ¥25,000 per clip batch, with ¥75,000 for a three-batch project. Existing assets: Street Attack Japan field knowledge, SET production planning habits, and existing recap formats. AI workflow: a video wrapper extracts keyframes and transcript text for Claude Code to analyze, then Adrian reviews the final brief. First action: test one non-sensitive clip and compare the notes against human viewing. Repeatability: the same template works for each shoot, venue, or event. Effort: half day. Expected value: paid prep support and time saved before edit or planning meetings.

Build Next

Deployable now

Test it

Build an API-key handoff helper for remote agent sessions

What to build: a tiny internal helper that lets a remote coding session request a needed API key without pasting the secret into a chat transcript. Why now: a new llm-keys-ui plugin was released for exactly this Codex Remote phone-control problem, and OpenAI’s cloud-agent direction makes remote sessions more common. Effort: one hour for a local trial. Expected impact: lower credential leakage risk and less friction when agents need test keys. Dependencies: a non-production key, one test machine, and a clear rule that secrets never enter chat.

Use it

Build a repo instruction normalizer

What to build: a simple repo check that reports whether a project has CLAUDE.md, AGENTS.md, both, or neither, then suggests the preferred instruction source. Why now: Claude Code reportedly falls back to AGENTS.md when CLAUDE.md is absent, which makes mixed-agent repo instructions easier to standardize. Effort: 15 minutes for a shell script or one hour for a GitHub Action. Expected impact: fewer agent mistakes caused by missing project rules. Dependencies: active repos and agreement on which instruction file wins when both exist.

Try This Today

One action, right now

Pick one active repo with no project instruction file. Add a short AGENTS.md with purpose, install command, test command, forbidden files, and review rule. Then run one Claude Code or Codex task and check whether the agent follows it.