Daily AI Zine
Tuesday, August 4, 2026
Issue No. 020 · Tokyo · full edition

✦ Today's Big Thing

Agents are getting their own computers

Today’s practical signal is infrastructure, not model drama: agent runtimes are moving beyond plain containers toward controlled computers that can run real work.

4 min read · 7 sections

In Brief
  1. Cloudflare introduced an agent runtime that shifts between fast isolates and full Linux containers.
  2. Meta says its ads foundation model now trains at LLM scale and doubled end-to-end training efficiency.
  3. Claude Code’s auto mode remains the useful safety pattern for lower-friction coding delegation.
  4. OpenAI’s small-business push is the clearest GPT Desk move to turn into a paid automation offer.

Today's Big Thing

The one thing that matters

Test it

Cloudflare gives agents a fuller machine to work inside

The important move is not a smarter chatbot. It is better operating space for agents.

Cloudflare introduced @cloudflare/computer, an agent runtime that gives each agent a “computer” rather than only a container. The summary says it can dynamically choose between fast isolates, which are lightweight execution spaces, and full Linux containers, which are heavier environments for real software tasks. That matters because serious agents increasingly need browsers, files, tools, and controlled execution, not just text responses. For Adrian, the useful read is infrastructure direction: agent work will need clearer sandboxes, logs, and failure boundaries before it can touch production workflows. Verdict: test one contained, non-client task before treating this as production infrastructure.

My AI Ecosystem

Your actual stack

Test it

Claude Code auto mode is still the safer autonomy pattern

Anthropic’s Claude Code auto mode work is worth revisiting because it frames permission-skipping as a safety design problem, not a convenience toggle. The useful pattern is simple: let the agent move faster only inside pre-approved task lanes, then keep review gates around files, secrets, and deploy steps. Verdict: test it on a low-risk repo before expanding.

Watch it

Hot signal: Meta is making ad AI a foundation-model workload

Meta says its Generative Ads Recommendation Model, used behind ads recommendations across Instagram and Facebook, now trains at LLM scale on several thousand latest-generation GPUs. The reported shift is efficiency: Meta says it doubled end-to-end training efficiency to 20 to 25 percent MFU while scaling training FLOPs 4x. For creative operators, watch how ad platforms fold more AI into targeting and recommendation systems. Verdict: watch.

Test it

Document extraction is getting a more useful test

ExtractBench tests schema-guided enterprise document extraction, meaning an agent reads documents and fills a defined structure with evidence. The benchmark covers 4,869 pages across 370 enterprise documents and scores accuracy, completeness, grounding, and cost together. This matters for production briefs, contracts, decks, and call-sheet inputs because cheap extraction without source evidence is operationally risky. Verdict: test one internal document type against this pattern.

CoWork Corner

Claude CoWork, day to day

CoWork: reset the evidence folder, not the whole room

Claude CoWork is the workspace Adrian uses for persistent projects, operational records, and AI-assisted workflows. No meaningful CoWork product change was found in today’s candidates. The useful revisit is a filesystem memory habit: one directory of markdown notes per active project, with stale assumptions removed and current evidence kept close to the task. A recent paper notes that deployed agents often keep long-term memory as a directory tree of files. Use that pattern deliberately: fewer notes, clearer names, and one source-of-truth project brief.

GPT Desk

OpenAI, ChatGPT, Codex

Use it

Use OpenAI’s small-business push as a sellable demo

What changed: OpenAI launched a ChatGPT for Small Businesses program focused on helping entrepreneurs build AI skills, automate work, and grow with ChatGPT Work. Why Adrian should care: Goodsense and Grey Group can turn that broad message into concrete Japan-English operating demos for small teams. What to do: build one 20-minute demo showing ChatGPT turning a messy inquiry, meeting note, and product page into a follow-up email, task list, and bilingual proposal outline.

Small Money Systems

Small, repeatable, real

Use it

Small system: ChatGPT Work starter audit for local teams

System: build a fixed-scope ChatGPT Work starter audit. Customer: Tokyo small businesses, production vendors, and bilingual service teams that know they should use AI but have no operating map. Offer: a one-page workflow map, three prompt templates, and one live automation demo. Price: ¥35,000. Existing assets: Goodsense positioning, Grey Group OS process notes, Adrian’s proposal workflow, and Japan-English business experience. AI workflow: ChatGPT drafts the map and templates, while Claude Code or Codex turns repeat steps into reusable files. First action: write the audit intake form this hour. Repeatability: each audit reuses the same questionnaire, demo script, and delivery template. Effort: one hour. Expected value: small paid revenue plus qualified leads for larger workflow retainers.

Build Next

Deployable now

Test it

Build a contained agent-computer runner for one safe task

What to build: a small runner that sends one non-client research or QA task into a controlled agent environment, saves inputs, outputs, and errors, then refuses deploy access. Why now: Cloudflare introduced @cloudflare/computer for agents that need more than a container. Effort: half day. Expected impact: safer testing of browser-like and file-based agent work before it reaches live workflows. Dependencies: a Cloudflare setup, one safe task, and an existing repo where Claude Code or Codex can wire the logging.

Try This Today

One action, right now

Pick one safe, annoying task that needs files or browser state. Write the allowed inputs, forbidden actions, success check, and stop condition. If the card feels vague, the agent runtime is not the bottleneck yet.