Tag
Coding Agents
10 articles

Creative Tooling
Figma's Agent Skills Sell Personalization the Data Doesn't Back
Figma's 13 August skill-authoring launch is pitched on capturing personal taste, but a study three days earlier found generic skills beat personalized ones.

Development
SWE-Touch: The Edit You Make While an Agent Still Runs
SWE-Touch's 3 August 2026 benchmark found resolve rates fall 7.7 points on average when a user edits code an agent is still working on, and the agent often finishes anyway.

Development
Merge Rate Became AI Coding's Scoreboard, and It Doesn't Agree
Four 2026 studies score coding agents by pull-request merge rate, but rankings flip between papers, suggesting the metric tracks the repo, not the agent.

Design Engineering
Design Systems Need Evals to Check if AI Agents Obey Them
A practitioner argues design systems need CI-run evals, since evidence on AI instruction-following suggests agents may ignore documented rules more than teams assume.

Development
GitHub Copilot's Two Modes Work Great Separately, Badly Together
A July 2026 field study finds mixing Copilot's autocomplete and chat modes in one task erodes their gains, even as Microsoft logs a durable 24% PR lift.

Development
GhostCommit Shows AI Reviewers and Agents Don't See Alike
GhostCommit hides prompt injection inside a PNG that AI reviewers skip and coding agents read, exposing a harness-level blind spot rather than a broken model.

Creative Tooling
DESIGN.md Turns Brand Identity Into a Forkable File
Community projects now package Apple, Stripe and Nike's visual identity into MIT-licensed DESIGN.md files that any coding agent can install to generate on-brand UI.

Development
Open Source's No-More-Pull-Requests Moment
Ladybird, tldraw, and the whole Jazzband collective have stopped taking public pull requests. It isn't a verdict on AI code quality — it's open source rebuilding its trust model from scratch.

Development
The End of Code Review, or Just Its Relocation?
A provocative paper declares human code review obsolete now that agents can do it faster. The evidence suggests something narrower and more interesting is actually happening.

Development
FrontierCode: The Benchmark That Asks Whether AI Code Is Ready to Merge
A new benchmark built with more than 20 open-source maintainers deflates the record-breaking numbers behind coding agents: even the best model clears only 13% of the hardest tasks.