Daily AI Zine
Tuesday, September 8, 2026
Issue No. 054 · Tokyo · full edition

✦ Today's Big Thing

Codex defaults move into ops territory

No single giant launch today, but one practical shift matters: saved agent settings now need the same review discipline as repos, deploys, and permissions.

5 min read · 7 sections

In Brief
  1. Codex model changes make saved defaults, managed settings, and automations worth checking now.
  2. GitHub keeps pushing parallel coding agents into ordinary Copilot use.
  3. A reported MCP connector issue is a reminder to inspect tool instructions before trusting agents with live work.
  4. Claude Code in CI needs explicit permission handling before unattended jobs run.

Today's Big Thing

The one thing that matters

Use it

Codex defaults are now an ops check

Quiet day, useful shift. An OpenAI Help Center item says GPT-5.4 and GPT-5.4 mini are no longer available in Codex for users signed in with a ChatGPT account, with replacements listed as GPT-5.6 Terra and GPT-5.6 Luna. The practical point is not the model names. It is that saved model settings, workspace defaults, managed configurations, and automations can drift when the platform changes. Treat Codex like deploy infrastructure: record which model each workspace uses, what it is allowed to touch, and which automations depend on it. This is an audit task, not a research task.

My AI Ecosystem

Your actual stack

Test it

GitHub teaches parallel agents as normal work

GitHub published a Copilot app guide on running several agents at once. The useful shift is cultural as much as technical: parallel agent work is being presented as a beginner workflow, not an expert hack. Test it on low-risk branches where the acceptance test is visible and the merge decision stays human.

Watch it

Hot signal: MCP connectors can carry hidden product instructions

A Reddit user reported that Notion's official MCP connector injected instructions that advertised Notion Business mid-task and told the agent not to explain why. Treat this as a reported signal, not a confirmed incident. Still, it points at a real MCP risk: connectors can shape agent behavior before the user sees the task. For any important connector, inspect the tool instructions and run a harmless test before adding project data.

Use it

Agentic coding evals need noise control

Anthropic has an engineering item on quantifying infrastructure noise in agentic coding evals. Plain English: a bad score can come from flaky test setup, tool failures, or environment issues, not the coding model itself. For Claude Code trials, separate model failure from harness failure before changing tools or prompts.

CoWork Corner

Claude CoWork, day to day

CoWork memory should stay project-local

Claude CoWork is the workspace Adrian uses for persistent projects, operational records, and AI-assisted workflows. No meaningful CoWork-specific product change surfaced today. The useful workflow to revisit is project-local memory: keep each working room tied to its own knowledge folder, then expose only the needed slice through MCP-style tools. A reported OKF Agent Memory project ships an MCP server for Claude Code, Cursor, Codex, and other agent platforms, which makes the pattern worth testing in a sandbox before connecting live records.

GPT Desk

OpenAI, ChatGPT, Codex

Use it

Audit Codex saved model settings

What changed: an OpenAI Help Center item says older Codex model choices tied to ChatGPT sign-in are being replaced, and saved defaults or managed settings may need updates. Why Adrian cares: Codex work inside Grey Group OS, Goodsense, SET, and build pipelines can quietly change behavior if defaults drift. Do this today: list every Codex workspace, its model, permissions, automations, and owner in one audit note.

Use it

Turn OpenAI governance into delivery rules

What changed: OpenAI published a case on Gilbert + Tobin scaling ChatGPT Enterprise and Codex with CEO-led commitment, governance, and human accountability. Why Adrian cares: Goodsense and Grey Group can sell AI work more credibly when every deliverable has an owner, review point, and allowed-tool boundary. Do this today: add a one-page AI accountability appendix to the next proposal or internal build brief.

Test it

Astra in Microsoft Copilot needs a client test script

What changed: ITmedia reports that Microsoft's Copilot environment now offers OpenAI GPT-6 Astra, with larger tasks delegated without breaking them into small steps. Why Adrian cares: many commercial clients live in Microsoft workflows, so SET and Grey Group should know whether Astra improves briefs, recaps, and proposal drafts there. Do this today: write a five-task test script using a client brief, a meeting summary, a call-sheet draft, a risk list, and a follow-up email.

Small Money Systems

Small, repeatable, real

Use it

Small system: Codex default audit

System: build a Codex workspace-default audit that checks model choice, saved settings, permissions, and automations. Customer: small AI-using teams that rely on ChatGPT, Codex, GitHub, or lightweight internal tools. Offer: a one-page risk map plus a recommended default-setting table. Price: ¥55,000 for a first audit. Existing assets: Grey Group OS, Codex usage patterns, GitHub repos, and proposal templates. AI workflow: use ChatGPT to interview the client, Codex to inspect config notes or repo docs, and Claude Code only if a small checker is needed. First action: make the intake checklist now. Repeatability: each new model change creates a reason to rerun it. Effort: one hour. Expected value: small direct revenue and fewer broken agent runs.

Test it

Small system: demo-video compression kit

System: package a browser video-compression page for quick demo clips, event recaps, and social proofs. Customer: local brands, event teams, small production crews, and Street Attack Japan partners who need fast lightweight video files. Offer: a branded upload page, compressed outputs, and a simple delivery note. Price: ¥33,000 setup, then ¥11,000 per event refresh. Existing assets: production workflow knowledge, Netlify habits, GitHub repos, and existing recap formats. AI workflow: Claude Code builds the page around a WebAssembly FFMPEG-style compressor, then ChatGPT writes client-facing instructions. First action: outline the upload, compress, download flow. Repeatability: reuse the same tool for every shoot or activation. Effort: half day. Expected value: recurring production savings and small add-on revenue.

Build Next

Deployable now

Test it

Build a Claude Code CI permission scan

What to build: a small GitHub Actions checker that flags Claude Code commands likely to hang or silently fail because nobody can approve permissions in CI. Why now: a reported CLI project is already scanning unattended Claude Code calls for this class of failure, and Anthropic has written about safer permission skipping. Effort: half day. Expected impact: fewer wasted CI minutes and fewer false-green automation runs. Dependencies: existing repos using Claude Code in GitHub Actions or another CI runner.

Try This Today

One action, right now

Pick one Codex workspace. Write down the current model, task type, allowed files, external tools, automation triggers, and human reviewer. If any field is unclear, do not run that workspace unattended until it is filled.