When AI generates work, who’s accountable?
AI can improve the mean quality of work on a team, but it makes accountability harder. Eric and John argue that it's not just about QA of the output, but who's doing the hard thinking before using AI.
Show content
Listen to the audio, watch the video, or check out the Show Notes for a summary, key takeaways, and links to people, content, and tools we mention.
Audio
Listen on Spotify:
Listen on Apple Podcasts:
Video
Summary
Eric and John start with the “taste and judgment” framing that has dominated AI conversations in tech circles, then make the case for a third differentiator that almost no one is talking about: accountability. When AI raises the floor on everyone’s work product, the real competitive edge shifts to who actually owns the output and who has the judgment to evaluate it.
They ground the conversation in fundamentals, tracing the lineage from Peter Drucker’s management by objectives through OKRs, and examine what accountability looks like at a practical day-to-day level. From there, the episode gets specific: what happens when a mid-level employee suddenly shines with AI, or when a strong performer starts making more mistakes? John’s answer cuts through the noise: AI rewards right thinking, not just fast doing.
The episode closes on two memorable examples. Eric describes an internal memo at Vercel called “Agent Responsibly,” which drew clear lines of jurisdiction: if you push code to production, you own it, whether you wrote it or an agent did. And a new hire’s insight about their content agent stops Eric cold: the tool is powerful, but it only works if people do the hardest part first, which is bringing a well-formed argument to the table.
Key takeaways
Accountability is the undersold differentiator: Taste and judgment get all the attention in AI conversations, but ownership of output is the harder and more important factor, especially as AI raises the floor for everyone.
AI rewards right thinking: When mistakes increase after someone starts using AI, the diagnosis is almost always the same: not enough time invested in research and planning before building.
Drawing lines of jurisdiction is a leader’s job now: “Agent Responsibly” frames it simply: if you push it to production, you own it. Leaders need to define those lines explicitly rather than letting them stay ambiguous.
Token budget allocation shapes output quality: Thinking of AI work as a pie chart of time spent across research, planning, prototyping, building, and QA reveals where most teams are underinvesting. Underspending on planning usually shows up as expensive QA.
Evals are accountability built into the machine: Defining what an AI agent should be able to do, writing test questions, and running them every time something changes is the structural equivalent of management by objectives for automated work.
AI erodes the skills required to use it well: L.M. Sacasas put it plainly: AI use tends to erode the formation of the virtue and expertise required to use it well. That makes deliberately practicing the hard, non-AI version of your work a form of self-preservation.
Don’t shortcut the front end: The quality of what comes out of an AI agent is a direct function of how well-formed the input is. Building systems that help people do the hard upstream thinking, not just execute downstream, is the real unlock.
Notable mentions and links
Peter Drucker is referenced as the father of modern management, whose methodology of management by objectives (MBOs) laid the groundwork for OKRs and the way most companies structure accountability today.
OKRs (Objectives and Key Results) are framed as the tech-world evolution of Drucker’s MBOs, and the episode uses them to define the difference between being accountable for outcomes versus being accountable for inputs.
WordPress comes up as the canonical example of how technological barriers used to constrain taste: when everyone had to use templates, design differentiation was impossible, but AI removes that constraint entirely.
Vercel is Eric’s workplace context throughout the episode, and the standards his team holds around human review of all published content are used to illustrate what accountability at scale looks like in a high-throughput AI environment.
“Agent Responsibly” is a Vercel blog post, based on an internal talk by engineer Matthew Binshtok, that drew clear lines of jurisdiction for AI-generated code: if you push it to production, you own it, regardless of who or what wrote it.
L.M. Sacasas is a technology writer whose tweet Eric brings into the conversation: “its use tends to erode the formation of the virtue and expertise required to use it well,” a property Sacasas argues is unique to AI and that shapes how leaders should think about building skills on their teams.
Dan Shipper, CEO and cofounder of Every, is mentioned in the context of his appearance on Lenny’s Podcast, where his thinking about how to structure AI work without losing the human ability to operate independently shaped John’s advice to his own team.
Whispr Flow is mentioned as a voice-first capture tool that John uses to start the planning process in a more analog, unstructured way before bringing AI into the loop.
Claude and ChatGPT are both cited as examples of AI tools people use daily, with the episode noting that both are building memory features that make agents non-deterministic in new and harder-to-predict ways.


