Daily AI Zine
Wednesday, September 9, 2026
Issue No. 055 · Tokyo · full edition

✦ Today's Big Thing

Codex gets a security-shaped proof point

The strongest practical signal today is not raw agent speed. It is Codex being framed as productive while still operating inside strict security rules.

5 min read · 7 sections

In Brief
  1. 1Password reports a 21% engineering productivity gain from Codex while maintaining rigorous security policies.
  2. GitHub has made GPT-6 Astra generally available in Copilot for long-horizon coding and agentic tasks.
  3. GitHub is also previewing HydraFusion, a multi-model coding workflow that aims to lower estimated workflow cost.
  4. Runway’s Gen-4.5 listing keeps high-end video competition focused on motion quality, prompt adherence, and visual fidelity.

Today's Big Thing

The one thing that matters

Use it

1Password makes Codex a security-policy story

The useful part is not only the 21% productivity figure. It is Codex being used for production-ready work under strict security practice.

OpenAI says 1Password engineers use Codex to build new features and internal tools, reaching production-readiness while maintaining rigorous security policies. The reported productivity lift is 21%. That makes today’s signal practical: coding agents are being sold less as solo speed tools and more as controlled engineering capacity. For an operator shipping AI-built systems, the takeaway is to make Codex prove three things on every task: what changed, why it is safe, and what review evidence exists before merge.

My AI Ecosystem

Your actual stack

Test it

Copilot adds Astra for long-horizon coding work

GitHub says GPT-6 Astra is generally available in GitHub Copilot. The model is described as OpenAI’s latest general-purpose model, designed for long-horizon autonomous coding and agentic tasks. Treat it as a new branch-level coding lane, not a default replacement. Test it on one contained implementation task with clear acceptance criteria.

Watch it

Copilot research preview routes coding work across models

GitHub’s HydraFusion preview uses selective coding workflows across models. GitHub says its controlled offline evaluations matched or exceeded the evaluated Opus 5 baseline while reducing estimated workflow cost. The practical point is routing: one agent may not be the cheapest or best worker for every step. Use this as a watch item until it proves itself on real repositories.

Test it

Hot signal: Runway keeps video pressure on motion quality

Runway Research lists Gen-4.5 as a video model focused on motion quality, prompt adherence, and visual fidelity. The source is thin, so do not treat this as a procurement decision. Treat it as a creative-tools signal: video models are competing on controllable movement, not only pretty frames. Run a short reference-shot test before using it in any commercial concept work.

CoWork Corner

Claude CoWork, day to day

CoWork stays steady, so tighten the handoff ritual

Claude CoWork is the workspace Adrian uses for persistent projects, operational records, and AI-assisted workflows. No meaningful CoWork product change appears in today’s candidate set. Revisit one useful habit: before a coding agent starts, write a three-line handoff in the project room covering the goal, the forbidden changes, and the review owner. Anthropic’s Claude Code best-practices material remains the better anchor today than any claimed CoWork feature news.

Tier 4 · quiet day, honest fallback

GPT Desk

OpenAI, ChatGPT, Codex

Use it

Use Codex only where it can show its safety trail

What changed: OpenAI says 1Password uses Codex for new features and internal tools, with a reported 21% engineering productivity increase while maintaining rigorous security policies. Why Adrian cares: Grey Group OS should not treat Codex output as finished just because it runs. What to do: pick one low-risk Goodsense or SET repo task and require Codex to produce implementation notes, risk notes, and a rollback note before review.

Test it

Polimill turns Japan knowledge search into a product pattern

What changed: Polimill is using OpenAI GPT models and Codex to help municipalities search and use administrative knowledge while accelerating development. Why Adrian cares: Goodsense and SET can turn Japan-English knowledge work into a visible demo rather than a consulting explanation. What to do: build a two-screen mockup for one municipal-style knowledge workflow, search on the left and draft response on the right, then use it as a sales conversation starter.

Small Money Systems

Small, repeatable, real

Test it

Small system: municipal knowledge demo pack

System: build a small Japan knowledge-search demo pack inspired by Polimill’s GPT and Codex municipal work. Customer: local agencies, civic-tech vendors, or Japan-facing consultancies that need to explain AI to conservative buyers. Offer: a clickable demo, one workflow map, and a bilingual briefing slide. Price: ¥88,000. Existing assets: Goodsense positioning, SET production discipline, Japan-English workflow, and Grey Group OS templates. AI workflow: Codex builds the demo shell, ChatGPT drafts bilingual copy, and Claude Code can clean the UI. First action: create the one-page offer today. Repeatability: swap the buyer segment and source documents without rebuilding the system. Effort: half day. Expected value: paid discovery and qualified meetings, not a guaranteed software contract.

Use it

Small system: Codex secure rollout mini-audit

System: package a lightweight Codex rollout audit for small teams that want speed without loose review. Customer: founders, studios, and small software vendors already experimenting with AI coding. Offer: one repo review, a Codex task policy, and a production-readiness checklist. Price: ¥55,000. Existing assets: Grey Group OS operating templates, Codex practice, GitHub habits, and Adrian’s delivery judgment from production workflows. AI workflow: Codex inspects repo structure and draft tasks, ChatGPT turns findings into client language, and Claude Code can add checklist files. First action: make the audit intake form in one page. Repeatability: each audit reuses the same checklist and only changes the repo-specific notes. Effort: one hour. Expected value: small service revenue plus better leads for larger implementation work.

Build Next

Deployable now

Test it

Build a Codex production-readiness gate

What to build: a GitHub Actions check that asks for purpose, files touched, security policy notes, test evidence, and rollback notes on any Codex-assisted pull request. Why now: 1Password’s Codex case pairs productivity with rigorous security policies, which is the right operating pattern. Effort: half day. Expected impact: fewer unreviewable AI changes and faster human approval. Dependencies: an active GitHub repo, Codex use on implementation tasks, and a simple PR template.

Test it

Build a video-model shot fidelity test page

What to build: a small Netlify page that compares short generated video tests by prompt, reference frame, motion note, and pass or fail comment. Why now: Runway is positioning Gen-4.5 around motion quality, prompt adherence, and visual fidelity, which makes testing more concrete. Effort: half day. Expected impact: better creative judgment before pitching AI video as part of a concept. Dependencies: Runway access, a few approved reference prompts, and a Netlify project.

Try This Today

One action, right now

Before the next Codex task, add five lines to the brief: goal, files allowed, files off-limits, security concern, and rollback plan. Then judge the output against those lines before reading the code in detail.