Podcast

The Pipeline Mag Podcast

Two synthetic hosts work through one Pipeline Mag article at a time — the argument, the evidence, and the counterpoint — in under ten minutes.

Every episode is generated by AI from an article already published here. If you would rather read it, each episode links back to the piece it came from.

RSS feed

Episodes

  1. 11 7 min 19 s

    AI Images Only Lose Trust Once Someone Suspects They're Fake

    Two studies published the same day find that AI-generated images cost a brand nothing until a viewer suspects one is fake — and that penalty lands on real photos too. We trace Nielsen Norman Group's hero-image test, a Frontiers in Computer Science watermark experiment with an oddly positive twist, and why EU Article 50 just made that suspicion a permanent, mandatory feature of every realistic image online.

  2. 10 7 min 28 s

    Figma's Agent Skills Sell Personalization the Data Doesn't Back

    Figma just shipped a skill-authoring feature pitched entirely on capturing a designer's personal taste — but a study released three days earlier found personalized coding-agent skills barely beat having no skill at all, while generic pooled skills won more often.

  3. 09 7 min 07 s

    AI Writes Responsive Code That Isn't Responsive

    A 12 August 2026 benchmark rendered 203 AI-generated webpages across nine real browser-and-device combinations and found 68% broke somewhere, 1.7 times the human baseline, while looking correct in both a code diff and a single-width preview. Two hosts work through why the failures split so widely by tool — 26% for Vercel's v0, 79% for Cursor, 100% for a raw GPT-5.1 call — and take seriously the benchmark's own caveats about its tool mix and its human baseline.

  4. 08 7 min 13 s

    Multi-Agent Coding Teams Don't Need a Boss, a Study Finds

    A 1,902-run study of Claude Code agent teams found naming a coordinator adds no measurable benefit, while shared-file versus messaging coordination swings token costs by up to 42%.

  5. 07 7 min 05 s

    Design Theater: The Gap Between an AI's Rationale and the Screen

    A July 2026 benchmark called Design Theater found that over a quarter of AI design tools' written rationales describe functionality the generated code doesn't actually have, rising to 34% on functional requirements. The two hosts work through where that gap comes from, why it survives a normal design review, and what the benchmark's own limitations do and don't undercut.

  6. 06 7 min 39 s

    AgenTag: AI Pull Request Tells Are in the Prose, Not the Code

    AgenTag, a 2 August 2026 attribution study, found that AI-authorship signal in pull requests lives almost entirely in the prose of the description, not in the code diff — a fingerprint that survives even after explicit disclosure markers are stripped out. Two hosts work through what that means for open source policies that rely on that same free-text field, and the study's own caveats about what it never actually tests.

  7. 05 7 min 27 s

    SWE-Touch: The Edit You Make While an Agent Still Runs

    SWE-Touch's 3 August 2026 benchmark found resolve rates fall 7.7 points on average when a user edits code an agent is still working on, and the agent often finishes anyway. Two hosts work through why the failure looks like success, and where capability closes the gap and where it doesn't.

  8. 04 7 min 04 s

    The AI Sparkle Icon Meets Europe's New Disclosure Law

    Europe's new AI content law demands a label that proves origin, but the sparkle icon design teams already ship was built to sell delight, not prove provenance. We trace the Nielsen Norman Group test where no one read the icon as AI, Google's own contradicting research, and the two-icon system Article 50 actually requires.

  9. 03 7 min 46 s

    Wealthfront's AI Code Reviewer Costs $4 a Pull Request to Say Nothing

    Wealthfront rebuilt its internal code reviewer so three AI models argue over a flagged issue before a human ever sees it, and now spends $4 a pull request to say as little as possible. Two hosts work through the research on why noisy AI review gets ignored, and the open question of whether a quieter reviewer is also letting more bugs through.

  10. 02 7 min 46 s

    AI Coding Tools Helped Blind Developers. Now Their Interfaces Are the Barrier

    A 5 August 2026 study validated 600 accessibility bug reports across five AI coding tools and found that maintainer attention, not model quality, decides which one a blind developer can actually use. The episode works through the numbers behind that gap and the research that complicates it.

  11. 01 6 min 50 s

    UX.md and DESIGN.md Reveal What AI-Ready Documents Leave Out

    Nielsen Norman Group published two contradicting essays on AI design documentation in July 2026. One argues research output should become AI-ready context like UX.md; the other warns that outsourcing research synthesis costs a team the learning. The episode works through who is supposed to own that curation.