Daily AI Zine
Sunday, August 9, 2026
Issue No. 025 · Tokyo · full edition

✦ Today's Big Thing

Agent speed now needs risk control

No single launch dominates today. The useful shift is clearer: agent-built code and agent-connected tools need review layers that measure risk, not just output volume.

4 min read · 7 sections

In Brief
  1. Meta describes an AI risk score for code changes that predicts whether a diff may cause a production incident.
  2. A new prompt-injection red-team paper points to a practical safety test for agent tools and MCP-style workflows.
  3. A community Claude Code benchmark is a warning against accepting token-saving claims without task-level measurement.
  4. Codex handoff across local, cloud, GitHub, and iOS is worth testing as a small dispatch workflow, not a full replacement for desktop work.

Today's Big Thing

The one thing that matters

Test it

Agent work shifts from speed to risk control

Quiet day for major product launches, but a useful pattern is getting sharper. Meta describes Diff Risk Score, an AI system that predicts whether a code change may cause a production incident and highlights risky parts of the change. Separately, PIMiner frames prompt injection, where a malicious instruction tricks an agent, as something builders can red-team with another agent. The important shift is not more autonomous coding. It is adding small, repeatable checks before agent work touches a live repo, site, or client-facing workflow. Verdict: test this now as a lightweight operating habit, not a research project.

My AI Ecosystem

Your actual stack

Test it

Diff risk belongs in agent code review

Meta’s Diff Risk Score uses AI to judge whether a code change is likely to cause a production incident, then points reviewers toward risky areas. Treat this as a design pattern for agent-built pull requests: every AI change should ship with a short risk note, affected files, and rollback path. Verdict: test on one repo before formalizing it.

Test it

Prompt-injection testing becomes builder work

PIMiner is a research system for red-teaming prompt injection, the attack where hidden text or tool output tricks an agent into doing the wrong thing. The practical takeaway is simple: any agent that reads web pages, files, or client text needs a small hostile-input test before it is trusted. Verdict: test on one MCP-style workflow.

Test it

Claude Code token savers need local proof

A community benchmark says several token-saving tools did not match their headline claims on real Codex and Claude Code workloads. Treat that as a signal, not a final verdict. The action is to measure your own repeated task before adding another compression layer, because bad context saving can cost time as well as tokens. Verdict: use only after a local benchmark.

CoWork Corner

Claude CoWork, day to day

CoWork: refresh the handoff note

Claude CoWork is the workspace Adrian uses for persistent projects, operational records, and AI-assisted workflows. No meaningful CoWork product change showed up in today’s candidates. The useful workflow to revisit is context discipline: keep one handoff note per active project with current objective, active files, blocked decisions, and what not to touch. Anthropic’s context-engineering guidance supports the same operating idea: agents perform better when the working context is explicit and narrow.

Tier 4 · quiet day, honest fallback

GPT Desk

OpenAI, ChatGPT, Codex

Test it

Use ChatGPT iOS as a Codex dispatch pad

What changed: reported ChatGPT Business notes say Codex can work across terminal, IDE, web, GitHub, and the ChatGPT iOS app, with account state connecting the handoff. Why Adrian should care: this fits small, queued engineering tasks for Grey Group OS or Daily AI Zine tooling when you are away from the desk. What to do: pick one low-risk GitHub issue, dispatch it from iOS, and require Codex to return a branch, test note, and review summary before merge.

Small Money Systems

Small, repeatable, real

Test it

Small system: diff-risk notes for site launches

System: build a one-page AI diff-risk note for each small website or automation change before launch. Customer: Japan-based founders, agencies, or production teams with GitHub and Netlify work but no formal QA layer. Offer: a short risk summary, files changed, likely breakpoints, rollback note, and reviewer checklist. Price: 25,000 JPY per launch note. Existing assets: Grey Group OS, GitHub repos, Netlify workflow knowledge, and Adrian’s production-approval instincts. AI workflow: Codex or Claude Code reads the diff and drafts the note, then Adrian reviews. First action: run one past pull request through the template today. Repeatability: every launch creates another paid checkpoint. Effort: one hour. Expected value: small launch revenue and fewer avoidable fixes.

Test it

Small system: MCP prompt-injection smoke test

System: sell a tiny prompt-injection smoke test for AI tools that read documents, websites, or shared folders. Customer: small companies starting with ChatGPT, Claude, or MCP connectors but unsure what can go wrong. Offer: five hostile-input tests, a plain-English risk note, and three safer operating rules. Price: 45,000 JPY per test pack. Existing assets: Grey Group OS, Claude Code, Codex, MCP experience, and Adrian’s Japan-English workflow knowledge. AI workflow: Claude Code generates test cases and logs results, then Adrian rewrites the findings for executives. First action: draft five test prompts against one internal workflow. Repeatability: every new connector or folder creates another paid check. Effort: half day. Expected value: lead generation and paid trust work before larger AI retainers.

Build Next

Deployable now

Test it

Build a pull-request risk note bot

What to build: a GitHub Action that asks Codex or Claude Code to summarize each pull request as risk level, affected workflow, manual test, and rollback note. Why now: Meta’s Diff Risk Score shows the practical value of using AI to flag risky code changes before production incidents. Effort: half day. Expected impact: faster review and fewer preventable launch mistakes. Dependencies: a GitHub repo with pull requests, an API account, and permission to run the action on non-secret diffs.

Try This Today

One action, right now

Open one merged pull request from the last month. Ask Codex or Claude Code for a four-line risk note: what could have broken, what test should have run, what rollback would be, and what review question was missing.