Quiet day: agent evals need noise control
The practical lead is not a launch. It is a measurement habit.
No single new item clears the bar as the day’s major development. The useful shift is Anthropic’s focus on quantifying infrastructure noise in agentic coding evals. In plain English: before deciding that one coding agent, model, or prompt is better, separate model performance from flaky test runs, slow tools, changing dependencies, and harness behavior. For Adrian, the takeaway is operational. When Claude Code, Codex, or another agent has a bad run, do not immediately rewrite the prompt. First record the repo state, task, tools, test result, and failure cause. Verdict: use this as the default review habit for agent work.