Daily AI Zine
Tuesday, September 15, 2026
Issue No. 060 · Tokyo · full edition

✦ Today's Big Thing

Copilot adds a cost dial for agent work

GitHub’s auto model selection now lets teams choose efficiency, balance, or intelligence, turning model routing into an operating choice instead of a hidden default.

5 min read · 7 sections

In Brief
  1. GitHub Copilot auto model selection now has three tiers for cost, quality, and response time.
  2. OpenAI’s Fyxer case study is a useful pattern for trusted executive-assistant workflows built on memory and feedback.
  3. Anthropic’s Model Hardware Standard preview is a watch item for physical-world agent safety.
  4. Claude Desktop’s one-click MCP server installation remains the cleanest CoWork-adjacent tool pattern to revisit today.

Today's Big Thing

The one thing that matters

Use it

Copilot turns model choice into an operating control

Auto model selection is no longer just automatic. It now carries a budget and quality preference.

GitHub Copilot auto model selection now offers three tiers: efficiency, balance, and intelligence. That matters because agentic coding work is becoming less about picking one model forever and more about routing each task to the right cost, quality, and speed profile. For daily shipping, this is a useful shift rather than a grand announcement. Small fixes, documentation cleanup, and first-pass scaffolding can default toward efficiency, while high-risk architecture or release work can justify intelligence. The practical move is to stop treating Copilot as one undifferentiated assistant and start logging which tier handled which class of task, then compare review burden and output quality after a week.

My AI Ecosystem

Your actual stack

Test it

OpenAI’s Fyxer example makes memory accountable

Fyxer uses OpenAI models, fine-tuning, memory, and real user feedback to organize inboxes and draft emails in each user’s voice. The useful part is not “AI email.” It is the feedback loop: the assistant improves because users keep correcting it. Treat this as a pattern for any trusted writing workflow where tone, timing, and discretion matter.

Test it

Claude’s new model line targets coding and knowledge work

Anthropic introduced Claude Fable 5.1 and Claude Mythos 5.1 as its most advanced models for coding and knowledge work, with research capabilities described as an early glimpse of AI progress. The summary is thin, so the action should stay narrow: test only against one known coding task and one knowledge synthesis task before changing defaults.

Watch it

Hot signal: physical-agent safety starts to get a shared format

Anthropic previewed the Model Hardware Standard, described as a shared specification for AI agents to safely operate physical systems. This is not a workflow tool yet. It is a signal that browser and software agents are no longer the only safety surface. For production and event work, watch for when camera, robot, venue, or device control starts requiring formal agent permissions.

CoWork Corner

Claude CoWork, day to day

CoWork: revisit the one-click MCP install path

Claude CoWork is the workspace Adrian uses for persistent projects, operational records, and AI-assisted workflows. No meaningful CoWork-specific product change showed up in today’s candidates, but Claude Desktop’s one-click MCP server installation remains the closest practical lever. Use it as a boundary exercise: one tool server for one recurring workspace job, installed cleanly, named clearly, and removed if it does not earn its place.

GPT Desk

OpenAI, ChatGPT, Codex

Test it

Use the Fyxer pattern for one trusted follow-up lane

OpenAI says Fyxer uses models, fine-tuning, memory, and user feedback to organize inboxes and draft in each user’s voice. Adrian should care because Goodsense, SET, Street Attack Japan, and Grey Group all depend on timely, tone-sensitive follow-up after proposals, meetings, and production conversations. Today’s action: choose one inbox label, export ten past replies, and ask ChatGPT to draft a voice guide plus three follow-up templates for review, not automatic sending.

Test it

Give GPT-6 Astra one difficult business-writing trial

OpenAI describes GPT-6 Astra as its most capable model for business, with advanced reasoning, computer use, and stronger writing and design judgment. Adrian should care because Grey Group OS needs repeatable judgment on proposals, decks, treatment notes, and client-facing language. Today’s action: give Astra one finished proposal and one rough brief, then ask for a revised executive summary, risk notes, and missing-decision list. Keep it advisory until it beats the current workflow twice.

Small Money Systems

Small, repeatable, real

Use it

Small system: inbox lead-revival desk

System: build a weekly AI-assisted follow-up desk that finds stalled warm leads and drafts human-reviewed replies. Customer: small production companies, creative studios, and Japan-market operators who lose work in inbox gaps. Offer: ten reviewed follow-up drafts, a priority list, and a voice guide. Price: ¥35,000 per weekly batch. Existing assets: Adrian’s Goodsense, SET, Street Attack Japan, and Grey Group OS proposal and communication patterns. AI workflow: ChatGPT creates the voice guide and drafts, with Claude Code or Codex later packaging the intake and review page. First action: label ten stale threads this hour and draft the first batch. Repeatability: the same checklist runs every week per client. Effort: one hour. Expected value: revived leads, faster replies, and fewer missed commercial openings.

Test it

Small system: coding-tier cost audit

System: create a lightweight audit that maps coding tasks to efficiency, balance, or intelligence tiers before an AI coding sprint. Customer: small teams using GitHub Copilot or agentic coding without cost discipline. Offer: a one-page routing policy, ten sample task classifications, and a review checklist. Price: ¥50,000 for a half-day setup. Existing assets: Grey Group OS delivery checklists, GitHub repos, and Adrian’s Claude Code and Codex operating habits. AI workflow: Codex or Claude Code reviews recent issues and pull requests, then drafts the routing table. First action: export ten recent coding tasks and classify them manually. Repeatability: every new sprint reuses the same table. Effort: half day. Expected value: lower wasted model spend and cleaner review decisions.

Build Next

Deployable now

Test it

Build an email voice review queue

What to build: a small review page where drafted replies are scored against a saved voice guide before a human sends them. Why now: OpenAI’s Fyxer example points to memory plus real feedback as the trust layer for inbox automation. Effort: half day. Expected impact: faster follow-up with fewer off-tone drafts reaching clients or partners. Dependencies: a ChatGPT workspace, a small set of approved past replies, and a simple place to store feedback.

Use it

Build a Copilot tier log for real tasks

What to build: a simple repo note or dashboard that records each Copilot-assisted task, selected tier, review time, and whether the output was accepted. Why now: GitHub just exposed efficiency, balance, and intelligence settings for auto model selection. Effort: one hour. Expected impact: less guessing about when higher-cost reasoning is worth it. Dependencies: an active GitHub Copilot setup and at least one week of comparable coding tasks.

Try This Today

One action, right now

Open one active repo and list ten recent or pending coding tasks. Mark each as efficiency, balance, or intelligence before running an assistant. After the work, add one note: accepted, revised, or rejected.