Category

Development

A bold vintage screen-print illustration in cork tan, deep pine green and rust red of a corkboard grid of pinned index cards linked by a dense crisscrossing web of red strings and pushpins, with one card slot left empty at the top where a boss card would sit.

Development

Multi-Agent Coding Teams Don't Need a Boss, a Study Finds

A 1,902-run study of Claude Code agent teams found naming a coordinator adds no measurable benefit, while shared-file versus messaging coordination swings token costs by up to 42%.

A bold vintage screen-print illustration in ivory, deep aubergine and mustard gold of a giant geometric fingerprint stamped over a ruled printed page.

Development

AgenTag: AI Pull Request Tells Are in the Prose, Not the Code

AgenTag's 2 August 2026 study found AI-authorship signal in pull requests comes almost entirely from PR descriptions, not code diffs — and rewriting the text defeats it.

A bold vintage screen-print illustration in parchment cream, deep navy and terracotta red of two crossing pens over a manuscript page, trailing two overlapping zigzag correction lines that visibly contradict each other.

Development

SWE-Touch: The Edit You Make While an Agent Still Runs

SWE-Touch's 3 August 2026 benchmark found resolve rates fall 7.7 points on average when a user edits code an agent is still working on, and the agent often finishes anyway.

A bold vintage screen-print illustration in warm parchment, deep teal and muted rust of a large faceted judge's gavel poised above its block, with three tiny orbs clustered near the hammer head.

Development

Wealthfront's AI Code Reviewer Costs $4 a Pull Request to Say Nothing

Wealthfront rebuilt its code reviewer so three AI models argue before any comment reaches a human, and now counts silence on most pull requests a win.

A bold vintage screen-print illustration in cream, deep teal-charcoal and dusty slate blue of a large mechanical split-flap scoreboard with mismatched abstract flip-panel tiles and a small stopwatch beside it.

Development

Merge Rate Became AI Coding's Scoreboard, and It Doesn't Agree

Four 2026 studies score coding agents by pull-request merge rate, but rankings flip between papers, suggesting the metric tracks the repo, not the agent.

A bold vintage screen-print illustration in aged cream, dusty slate-blue and rust terracotta of a boxy photocopier spilling an identical fan of duplicated paper sheets onto a filing cabinet.

Development

GitClear's Maintainability Gap Puts a Price on AI's PR Boom

GitClear and GitKraken's 623-million-change study finds AI coding lifts pull requests but pushes code duplication up 81% as legacy maintenance falls 74% since 2023.

A bold vintage screen-print illustration in mustard amber, deep plum and muted teal-green of a railway track switch splitting into two diverging rails beside a faceted semaphore signal.

Development

GitHub Copilot's Two Modes Work Great Separately, Badly Together

A July 2026 field study finds mixing Copilot's autocomplete and chat modes in one task erodes their gains, even as Microsoft logs a durable 24% PR lift.

A bold vintage screen-print illustration in aged paper cream, charcoal and red pen crimson of a tilted printed specification document marked up with a heavy diagonal strike-through, a circled correction, and a solid red sticky note.

Development

Spec-Driven Development Can't Spec the Governance That Matters

A 420-KLOC case study challenges spec-driven coding's premise, arguing real governance is discovered from failures mid-project, not written into a spec upfront.

A bold vintage screen-print illustration in indigo-charcoal, warm parchment and amber-gold of a large picture frame holding a faceted geometric ghost silhouette, set against horizontal code-printout bands.

Development

GhostCommit Shows AI Reviewers and Agents Don't See Alike

GhostCommit hides prompt injection inside a PNG that AI reviewers skip and coding agents read, exposing a harness-level blind spot rather than a broken model.

A bold vintage screen-print illustration in cream, deep charcoal-navy and alarm red of a circular smoke-detector alarm with concentric rings casting a triangular beam of red warning light down onto a laptop keyboard.

Development

Your Coding Agent Trips the Same Alarms as an Intruder

Sophos telemetry from June 2026 shows Claude Code, Cursor and OpenAI Codex tripping the same rules built to catch attackers, just as GitHub ships an auto-approve mode.

A constructivist vintage screen-print cover in cream, slate-blue and terracotta, with a large geometric bar chart whose tallest bars topple and shatter.

Development

Amazon and Meta Killed Their AI Coding Leaderboards

Amazon's KiroRank and Meta's Claudeonomics ranked engineers by AI tokens consumed, until gamed usage inflated costs and both companies quietly shut the boards down.

A constructivist vintage screen-print cover in dusty blue, navy and ochre, with a large geometric stopwatch.

Development

The Study METR Couldn't Run: What a Failed Control Group Reveals About AI Coding

METR tried to repeat its AI-productivity study in 2026 and couldn't recruit developers willing to work without AI — a methodological failure that may say more than any number could.