Daily AI Zine
Saturday, September 5, 2026
Issue No. 051 · Tokyo · full edition

✦ Today's Big Thing

GitHub tests coding by model mix, not model loyalty

HydraFusion is the strongest new signal today: Copilot is moving toward selective multi-model coding workflows, with cost and quality judged at the task level.

7 min read · 7 sections

In Brief
  1. GitHub’s HydraFusion research preview matched or beat an evaluated Opus 5 baseline in offline coding tests while reducing estimated workflow cost.
  2. GPT-6 Astra is now generally available in GitHub Copilot for long-horizon coding and agentic tasks.
  3. Gemini 3.8 Flash also landed in GitHub Copilot, giving builders more practical model choice inside one coding surface.
  4. Polimill’s OpenAI and Codex work in Japan points to a repeatable knowledge-search offer for public-sector and regulated teams.

Today's Big Thing

The one thing that matters

Test it

Hot signal: GitHub tries multi-model coding inside Copilot

GitHub’s Project HydraFusion is a research preview inside Copilot that tests a practical shift: stop asking one model to do every coding job, and route parts of the workflow across models. GitHub says its selective coding workflows matched or exceeded the evaluated Opus 5 baseline in controlled offline evaluations while reducing estimated workflow cost. The useful part is not the benchmark claim alone. It is the operating pattern: judge model choice by finished task quality and cost, not brand preference. For a stack that already uses coding agents, this is worth a controlled test on repo work where planning, implementation, and review can be separated. Verdict: test it on one contained issue before trusting it on production-critical work.

My AI Ecosystem

Your actual stack

Test it

Astra is now a Copilot model, not just an OpenAI launch

GPT-6 Astra from OpenAI is generally available in GitHub Copilot. GitHub describes it as a general-purpose model designed for long-horizon autonomous coding and agentic tasks. Treat this as a repo-level trial, not a default switch. Give Astra one task with clear acceptance checks and compare output against the current Claude Code or Codex path.

Test it

Gemini Flash adds another Copilot coding lane

Gemini 3.8 Flash is now available in GitHub Copilot. GitHub says early testing showed strength on complex terminal-based coding tasks. The practical move is to keep it in a lower-cost or exploratory lane until it proves better than the existing model for one repeatable job, such as dependency cleanup or terminal-heavy setup work.

Use it

Copilot’s weekly release focuses on control, not only models

GitHub’s weekly Copilot release expands model choice and content protections, while VS Code adds ways to manage agent sessions and get pull requests merge-ready. The useful shift is operational: more agents only help if sessions are visible, bounded, and reviewable before merge.

Watch it

Editable visual design points back to layered creative output

A new paper on Editable Visual Design argues that image models still tend to produce flattened bitmaps, while code-based design can preserve editable layers but lacks stronger visual taste. The immediate lesson for production design is simple: keep AI-generated visuals editable whenever they may face client revisions.

CoWork Corner

Claude CoWork, day to day

CoWork: no new product move, recheck the decision record

Claude CoWork is the workspace Adrian uses for persistent projects, operational records, and AI-assisted workflows. No meaningful CoWork product change showed up in today’s candidates. The useful workflow to revisit is the decision record: each room should keep one current objective, one latest decision, one blocked question, and one next handoff. That keeps persistent context from becoming a storage pile. It also mirrors the discipline needed for coding agents, where the hard part is not producing more output but keeping the work bounded and reviewable.

Tier 4 · quiet day, honest fallback

GPT Desk

OpenAI, ChatGPT, Codex

Test it

Make Astra prove itself on one production-shaped branch

What changed: GPT-6 Astra is generally available in GitHub Copilot and is positioned for long-horizon autonomous coding and agentic tasks. Why Adrian should care: Grey Group OS and Goodsense work often needs a model to hold intent across files, not just answer one prompt. What to do: pick one low-risk repo issue, ask Astra for a plan plus implementation path, then run the same task through the current Claude Code or Codex route and compare review burden.

Use it

Package Japan knowledge search as a sales demo

What changed: Polimill is using OpenAI GPT models and Codex to help Japanese municipalities search and use administrative knowledge while accelerating development. Why Adrian should care: this is close to the Goodsense and SET lane, where Japan-facing knowledge, bilingual workflow, and practical delivery matter more than abstract AI demos. What to do: draft a one-page demo brief showing how a client’s scattered policies, FAQs, and project notes become a searchable operating layer.

Small Money Systems

Small, repeatable, real

Use it

Small system: Japan knowledge-search readiness scan

System: build a fixed-scope readiness scan for turning messy Japanese organizational knowledge into an AI-searchable operating layer. Customer: small municipalities, public-sector vendors, schools, or Japan-market teams with scattered policy and admin documents. Offer: a short bilingual report, three sample search questions, and a recommended build path. Price: ¥45,000. Existing assets: Goodsense positioning, SET production discipline, Japan-English workflow, and Grey Group OS templates. AI workflow: ChatGPT drafts the question set, Codex prepares a lightweight demo structure, and Claude Code turns the process into a reusable checklist. First action: write the one-page offer today. Repeatability: every scan uses the same intake form and report shell. Effort: one hour. Expected value: paid discovery that can become a larger implementation lead.

Test it

Small system: repo agent bakeoff pack

System: sell a compact agent comparison pack for one real GitHub repo task. Customer: founders or small technical teams unsure which coding agent or model to trust. Offer: run the same issue through two or three lanes, then deliver a plain report on quality, review time, and risk. Price: ¥60,000. Existing assets: Grey Group OS, Claude Code, Codex, GitHub workflow, and Adrian’s production habit of comparing outputs before delivery. AI workflow: use Copilot model options, Codex, and Claude Code to produce candidate approaches, then summarize the decision. First action: choose one recent low-risk issue as the sample case. Repeatability: the template becomes a monthly model-choice check. Effort: half day. Expected value: consulting revenue plus better internal model routing.

Test it

Small system: fast event recap page

System: create a paid recap package that turns a small event, activation, or shoot into a same-week branded web page and social cutdown plan. Customer: local brands, venues, agencies, and Street Attack Japan partners who need visible proof after an activation. Offer: a Netlify-hosted recap page, short copy, selected stills or clips, and three follow-up post angles. Price: ¥80,000. Existing assets: Street Attack Japan, production planning, decks, GitHub, and Netlify. AI workflow: ChatGPT drafts the recap copy, Claude Code or Codex builds the page shell, and current video-generation tools can support style frames when needed. First action: clone one prior recap format into a reusable template. Repeatability: every event reuses the same intake, page, and delivery checklist. Effort: half day. Expected value: repeatable post-event revenue and stronger sponsor follow-up.

Build Next

Deployable now

Use it

Build a Copilot model-change watcher

What to build: a tiny internal note generator that watches GitHub Copilot model and agent-session changes, then outputs a one-page decision note for which models deserve a test. Why now: Copilot added GPT-6 Astra, Gemini 3.8 Flash, model-choice expansion, content protections, and better agent-session management in the same release window. Effort: one hour. Expected impact: less manual tracking and fewer accidental default-model decisions. Dependencies: an existing GitHub account, a place to store the note, and one owner who updates the test result.

Test it

Build a three-lane agent task card

What to build: a reusable GitHub issue template that sends one task through planning, implementation, and review lanes, with the model used in each lane recorded. Why now: HydraFusion’s preview makes model orchestration a practical workflow to test, not only a research idea. Effort: half day. Expected impact: better output quality and clearer review burden on agent-built changes. Dependencies: a repo with issues, access to at least two coding-agent lanes, and a small task that can be judged within days.

Try This Today

One action, right now

Pick one small repo task. Ask one model to plan, one to implement, and one to review. Record time, corrections, and which lane added the most value. The goal is not speed today. The goal is a repeatable routing rule.