<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Executive Summary, a Gather publication]]></title><description><![CDATA[More signal, less noise: Helping engineering leaders to thoughtfully manage stakeholder expectations, drive organizational adoption, and build the technical capabilities required to transition to an agentic SDLC.]]></description><link>https://executivesummary.gather.dev</link><image><url>https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png</url><title>The Executive Summary, a Gather publication</title><link>https://executivesummary.gather.dev</link></image><generator>Substack</generator><lastBuildDate>Wed, 22 Jul 2026 02:21:25 GMT</lastBuildDate><atom:link href="https://executivesummary.gather.dev/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Peter Bell]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[gatherexecutivesummary@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[gatherexecutivesummary@substack.com]]></itunes:email><itunes:name><![CDATA[Peter Bell]]></itunes:name></itunes:owner><itunes:author><![CDATA[Peter Bell]]></itunes:author><googleplay:owner><![CDATA[gatherexecutivesummary@substack.com]]></googleplay:owner><googleplay:email><![CDATA[gatherexecutivesummary@substack.com]]></googleplay:email><googleplay:author><![CDATA[Peter Bell]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[I'm back!]]></title><description><![CDATA[Not working is HARD!]]></description><link>https://executivesummary.gather.dev/p/im-back</link><guid isPermaLink="false">https://executivesummary.gather.dev/p/im-back</guid><dc:creator><![CDATA[Peter Bell]]></dc:creator><pubDate>Wed, 01 Jul 2026 15:54:37 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I just wanted to apologize that I&#8217;ve been in a black hole for the last few weeks. As I mentioned in a <a href="https://gatherdev.substack.com/p/i-cant-afford-to-work-any-more">previous article</a>, I really feel that I can&#8217;t afford to work anymore, except in dire circumstances. <br><br>The TL;DR is that I hope AI is not going to take my job for a while. But I know that somebody using AI will. As such, I can&#8217;t afford to simply do work. There isn&#8217;t enough leverage. The only thing worth doing is building and tuning the systems that will do the work. And I&#8217;ve been working hard on that for the last few weeks. <br><br>I&#8217;m still not happy with all the loops and I&#8217;ve still got a lot of taste to encode. But the systems are finally coming online and getting to the point where they are almost as good as me. And I&#8217;m hoping that by September, they&#8217;ll be better than me at most of the things I do - especially with a Fable glow up this week while it&#8217;s available on the max plans. <br><br>I will be breaking this substack into multiple newsletters to reflect the different dimensions of my work so you can pick and choose. There will be channels for engineering leadership in the age of AI (including this substack), agentic engineering and software factory design, solopreneurship in the age of AI, and some occasional musings on the broader changes I expect to see in society over the next few years - I&#8217;ll share more details as they all launch.</p><p>I&#8217;ve also been interviewing a bunch of AI thought leaders and CTOs for the <a href="http://scalingaibook.com/">O&#8217;Reilly book</a> on Scaling AI Adoption in Engineering (should be copy complete in July and available as an early release digitally shortly after and then as a book around Thanksgiving). I&#8217;ll start to drop those interviews next week.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>I&#8217;m moving the <a href="https://www.gather.dev/">Gather.dev</a> events to quarterly (the 2026 calendar is live <a href="https://luma.com/gather-dev">here</a>). There will be can&#8217;t miss intimate quarterly afternoon forums for Scaling CTOs (series B+, 50-250 engineers) and CTOs at Scale (orgs of 250 - 10,000 engineers) where we&#8217;ll dig deep into what is and isn&#8217;t working, sharing insights and experiences that won&#8217;t be available elsewhere.  These will be followed by evening events for Startup CTOs ($1m+ raised) and DirectorPlus (managers of managers at companies with ~50 - 10,000 engineers) with speakers and a room full of engineering leaders building or adopting dark factories. They will all be intimate (~20 participants for the forums and ~50 for the evening events) - free, invite only. So if you want to save a slot in NY or SF, you should probably <a href="https://luma.com/gather-dev">sign up</a> now (I&#8217;ll ping you a month before to confirm attendance so we can open up slots to the waitlisters).<br><br>I&#8217;ll also be spending a lot more time focused on advisory work. For $5k - $15k a month companies will have access to a pipeline of good practices and recommendations for everything from strategic positioning to stakeholder and change management to practical advice on context stores, pipelines and compounding systems. I&#8217;ll only have a few slots, and next week I&#8217;m going to start offering them to 30,000 of my closest friends, so feel free to drop me a line if you&#8217;d like to audit/benchmark/evaluate/plan or accelerate your AI in the SDLC initiatives - I&#8217;ll probably only take on 5-8 clients so I&#8217;m hoping to have a waitlist by August.</p><p>It&#8217;s good to be back! (Talking of which, come on <a href="https://www.anthropic.com/news/redeploying-fable-5">Fable</a> :) ).</p><p></p><p></p><p></p><p></p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[I can’t afford to work any more]]></title><description><![CDATA[Building the systems, not doing the work]]></description><link>https://executivesummary.gather.dev/p/i-cant-afford-to-work-any-more</link><guid isPermaLink="false">https://executivesummary.gather.dev/p/i-cant-afford-to-work-any-more</guid><dc:creator><![CDATA[Peter Bell]]></dc:creator><pubDate>Sat, 02 May 2026 18:19:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Sunday night, 9pm. I&#8217;m copy-pasting LinkedIn URLs into a file because that&#8217;s the input my agent expects in the morning. It&#8217;s the third time this month I&#8217;ve done this exact thing. The agent does the heavy lifting. I do the prep. And the prep, it turns out, is the bottleneck.</p><p>So on Monday I&#8217;m going to stop working.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>I already have a small agentic team partnering with me on most of what I used to do, and we&#8217;re moving at 2-3x my pre-agentic velocity. That&#8217;s not enough. 10x-100x is accessible to me, and to anyone else, if I invest in the specifications, the verifications and the tooling instead of plodding alongside the agents.</p><p><strong>The solopreneurial struggle</strong></p><p>Building Gather, I need to identify and book speakers, find venues, and run logistics. I need to record, edit and drop video, podcast, social and articles from every interview. I need to identify and qualify potential community members, communicate with the community, and help out with the common issues.</p><p>I need to research trends in agentic AI and engineering leadership. I need to keep up with the bleeding edge in harness design. I need to build my own harness and use it to run the business and deliver a wide range of applications.</p><p>I need to host events in NYC, the Bay Area, and online, and figure out how to scale to 15-20 cities by year&#8217;s end without lowering the quality bar or hiring a human. I also need to experiment with what a community for engineering leaders looks like when all of their communications are intermediated by their agentic team. Between that, a kid and a life, I&#8217;m usually behind the curve.</p><p><strong>I can&#8217;t afford to work</strong></p><p>So this week I&#8217;m finally taking the time not just to do tasks, but to write a rich, multi-step recipe for every repeatable one. If I do research, I want to identify and capture sources, then extract, de-dupe, classify and enrich the people, companies, publications and platforms they&#8217;re on. I want to extract insights and map them to a rich taxonomy, with patterns for adding new nodes to the knowledge graph for unknown unknowns, and for archiving or re-parenting nodes as the field shifts.</p><p>This is painful. It takes time to thoughtfully describe a high-value pipeline with clear validation and verification steps. And I don&#8217;t currently have a working orchestrator, so I need to make some pragmatic choices between a lightweight (perhaps file-based) notebook that keeps track of flows, and just building out a proper orchestrator with o11y and maintenance models and tools for cleanup and the rest of the patterns required to deliver a robust solution.</p><p>Hopefully my pragmatism wins over my perfectionism. I&#8217;ll prototype with an in-memory supervisor pattern, sub-agents using a semi-structured markdown file as an agreed recipe. If I like that, I&#8217;ll describe a DSL for the workflows and a schema validator to make sure they&#8217;re well-formed. Then I&#8217;ll build the simplest possible orchestrator to instantiate a run in the database and track state with enums, and finally I&#8217;ll add the edge-case handling as my agents find new ways to break things on the way to an entirely automated org.</p><p><strong>Your team can&#8217;t afford to work either</strong></p><p>Why am I sharing my solopreneurial woes with you? Whether you have a small team or a large org, it can be hard to find the time to invest in the future of the agentic SDLC when you have a backlog, bugs, customers, interdependent projects and a full set of commitments you need to deliver to meet promises to customers, or to match your planned release cadence.</p><p><em><strong>But if you are not investing in teaching your agents how to write code, you&#8217;re going to get passed by teams in your industry that are doubling down on that today.</strong></em></p><p><strong>A million LoC in a month</strong></p><p><a href="https://www.linkedin.com/in/schillace/">Sam Schillace</a> is a Deputy CTO in the Office of the CTO at Microsoft. He ran a ~600-person team at Google and ran engineering at Box, so he knows what it takes to deliver secure, reliable code at scale. His team shipped <a href="https://github.com/microsoft/amplifier">Amplifier</a>, one of the least hyped and most interesting agentic tools I&#8217;ve seen.</p><p>The other month, he personally shipped a million lines of code. To be fair, he didn&#8217;t deploy them to production, and lines of code is a terrible and easily gamable metric. But one engineer, with a tuned harness and years of experience with LLMs, basically replicated Microsoft Word using web technologies in a month. The future is coming, and it&#8217;s not just 40% faster.</p><p><strong>Can you see the speed?</strong></p><p>The most common refrain I&#8217;m hearing is &#8220;if dark factories are so capable, how come nobody is shipping anything to production?&#8221; Truth is, they are. You can already see it if you know where to look.</p><p>The most visible example is the Claude Code team. There have been security issues and bugs and all the usual side effects you see when it is genuinely economically rational to move really fast and break things. But they&#8217;re also genuinely moving fast. Look at the size of the team, the number of customers, and the speed of shipping. This is not just Boris Cherny and the team using auto-complete to speed up the typing in their IDEs.</p><p>For most of you reading this, the bigger concern is your competitors who may be investing in this and you haven&#8217;t seen what they&#8217;re doing yet.</p><p>It reminds me of building a skyscraper. For a long time all you see is a big hole in the ground. Then once the architecture is locked, the permits obtained, and the foundations laid, the stories go up with remarkable speed.</p><p>Building an agentic SDLC requires a substantial change in how your teams work, a lot of training to help them understand this new way of delivering software and a lot of incremental experiments to encode what good software development means specifically within your company. For a team in a large org, expect weeks or months to accelerate delivery by 5x to 10x. (Solo, with a tuned harness, the ceiling is higher; that&#8217;s the 10x-100x I&#8217;m aiming for personally.)</p><p>If one of your competitors gets there first, by the time you see the results in their product announcements you are probably six months behind on being able to replicate it within your own org. And at the pace companies can ship at with a fully agentic SDLC, six months is the equivalent of being years behind at a more traditional pace.</p><p><em>How confident are you that the permits, the architecture and the foundations are in place in your engineering team for when your competitors start shipping at agentic speed?</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[ Don't Use LLMs ]]></title><description><![CDATA[Unless you really really have to!]]></description><link>https://executivesummary.gather.dev/p/dont-use-llms</link><guid isPermaLink="false">https://executivesummary.gather.dev/p/dont-use-llms</guid><dc:creator><![CDATA[Peter Bell]]></dc:creator><pubDate>Sun, 26 Apr 2026 16:23:55 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1>Don&#8217;t Use LLMs</h1><p><em>Unless you really really have to!</em></p><p>I&#8217;m the biggest fan of LLMs. I use them every day. But right now I spend a lot of my time trying to convince people <strong>not</strong> to use them. Why? Because you can&#8217;t guarantee availability, you can&#8217;t guarantee the outcomes, and there are much better ways of solving most problems.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>Use LLMs</h2><p>I spend most of my day interacting with LLMs. They can help me to research key topics, refine my positions and keep track of my progress. For one-off tasks, they are a magical tool.</p><h2>Don&#8217;t Use LLMs</h2><p>I was chatting with <a href="https://www.linkedin.com/in/ryanlopopolo/">Ryan Lopopolo</a> on X. If you haven&#8217;t read it, he wrote the excellent OpenAI piece on <a href="https://openai.com/index/harness-engineering/">coding harness design</a>. He was talking recently about the amazing improvements in long-horizon task planning in codex. I agreed, but I think that in the general case the supervisor pattern (long running agent supervising sub agents to minimize impact on context usage) is overused. You don&#8217;t need an LLM to keep track of subagents, that&#8217;s what pipeline descriptions and finite state machines using enums in a database are for.</p><p>If I don&#8217;t really know how I want a problem solved, I can spin up a SoTA model, point it at the problem and get surprisingly good results. But I find that most of my use cases are repetitive business flows, whether that is specifying, writing, testing and deploying code, or automating enrichment and outreach, or booking vendors and speakers for events.</p><p>For repetitive use cases, you are MUCH better served by having a mechanism to describe a pipeline that is run deterministically, and has model and/or human in the loop steps. It reduces token usage, improves determinism, reduces latency, heck it even saves the climate by burning less fossil fuel :) And you can always point your LLM at it so it can generate the json for you!</p><h2>Two Rules</h2><p>I find that if I follow two guidelines, it allows me to optimize the efficiency of my high value, repetitive business processes:</p><ul><li><p>Models should be used to write code to replace themselves</p></li><li><p>Humans should be used to train models to make human review unnecessary</p></li></ul><p>That&#8217;s it. If I explore a workflow in plan mode and have a supervisor run sub-agents to do something, that&#8217;s a great way to prototype a workflow and get a sense for how it should run. If I want to repeat that workflow, I need to give the agent a workflow description language so it can specify the flow and have a database keep track of the status of the runs.</p><p>If you have a step where an agent is calling an API to enrich a record and writing it to the database, is judgment actually required? If not, can the agents write a script to do it all so you&#8217;re not burning tokens for no good use?</p><p>If you have a step where an agent is classifying records based on few-shot learning, can you get it to write a simple Python rules-based script for most cases, to reduce the subset that actually requires an LLM?</p><p>And I find that as a human, if I treat all of my activity as RLHF, it doesn&#8217;t guarantee that I&#8217;ll get out of the loop, but it maximizes the chance that I&#8217;ll have to invest the minimum attention in the repetitive flows that I build, whether they are for GTM or R&amp;D.</p><h2>Drift</h2><p>The temptation runs the other way. Put an LLM in the middle, let it decide the branches, summarize state, classify inputs, call tools, carry the plan. You get something running in a weekend that used to take a quarter. It&#8217;s a gift, until it isn&#8217;t.</p><p>Long-running supervisor agents drift. They stop for reasons you can&#8217;t always reproduce. They lose track of where they were. They reach for a tool they already called. They hit an edge case three hours in that they can&#8217;t recover from.</p><p>If your pipeline has to run for eight hours before it produces value, and your supervisor stops at hour six, you restart. That&#8217;s eight new hours of tokens, latency, and customer frustration. And you don&#8217;t know, on the next run, whether it will stop again.</p><p>This isn&#8217;t a problem <a href="https://www.anthropic.com/claude">Claude</a> 5 fixes. The longer the horizon, the worse the math, regardless of how good the model gets.</p><h2>Conditionals</h2><p>Better models extend the horizons and reduce the hallucinations. But if statements never hallucinate at all.</p><p>If the only thing you need at this step is to decide whether a new ticket is a bug, a feature request, or a support question, you don&#8217;t need a frontier model. You need a classifier. If you need to decide whether an invoice is over the approval threshold, you need a comparison operator. If all you need is a conditional, why do you have an LLM?</p><p>There are dozens of these steps in any real production pipeline, and most of them are conditional logic dressed up as decision-making. The LLM can do them. It can also cost more, take longer, and occasionally invent a fourth branch nobody asked for.</p><h2>Supply</h2><p>Token costs could go up. They could also come down. Max plans are already going away or shrinking. What supply and demand look like next year, I don&#8217;t know, and you don&#8217;t either.</p><p>Here&#8217;s the part that matters. Price elasticity for frontier models is high. If a model gives you a 5x or 10x productivity gain on work that&#8217;s worth a lot of money, some customers will pay a lot of money to get it. That means the top of the market stays gated by willingness to pay, regardless of which direction the headline price moves.</p><p><a href="https://tomtunguz.com/ai-compute-crisis-2026/">Tom Tunguz&#8217;s recent piece on the compute crisis</a> has the data. <a href="https://www.nvidia.com">Nvidia</a> Blackwell rental pricing rose forty-eight percent in two months. <a href="https://www.coreweave.com">CoreWeave</a> raised prices twenty percent.</p><p><a href="https://openai.com">OpenAI</a>&#8216;s CFO Sarah Friar said the quiet part out loud: &#8220;We&#8217;re making some very tough trades at the moment on things we&#8217;re not pursuing because we don&#8217;t have enough compute.&#8221;</p><p><a href="https://www.anthropic.com">Anthropic</a> is in the headlines right now for reducing token availability on plans and potentially running out of capacity. OpenAI may be next. If you&#8217;re not one of the whales that the labs are wooing, good luck getting time with a sales rep (let alone customer service) for your $500k or $1m annual commit.</p><p>So here&#8217;s the question. If your foundation lab called tomorrow and told you your account was being deprioritized, what does your Monday look like?</p><h2>Compile Down</h2><p>That&#8217;s the whole idea. Default to deterministic code wired around the repeatable process. Keep the LLM in the loop at the specific steps where it uniquely earns its spot: classification where a rules engine is too brittle, generation of prose or code, summarization of genuinely varied input, and the occasional open-ended judgment call that nothing else can make.</p><ul><li><p>Branches that can be expressed as conditionals should be conditionals.</p></li><li><p>Validation that can be expressed as schema checks should be schema checks.</p></li><li><p>Human-review stops go where the blast radius makes them worth the latency.</p></li></ul><p>Let frontier models build the system. Let deterministic code run it.</p><h2>Won&#8217;t Better Models Fix This?</h2><p>The most common pushback goes: <em>&#8220;If I over-engineer around today&#8217;s models, I&#8217;ll rip it all out in six months when they&#8217;re better.&#8221;</em> Fair, for a specific kind of engineering. Scaffolding you build to patch a model that drifts, hallucinates, or forgets its instructions does become dead weight as models improve.</p><p>Deterministic business logic isn&#8217;t that scaffolding.</p><p>An if statement against a business rule doesn&#8217;t get obsoleted by Claude 5. A schema validator doesn&#8217;t become redundant. A durable pipeline step that writes the outcome to Postgres doesn&#8217;t care which model is in use this quarter. That code encodes how your business works. The model is a moving target.</p><h2>The Takeaway</h2><p>Use LLMs where no cheaper, more reliable, lower-latency alternative exists. Everywhere else, prototype with LLMs and then have them compile the flows down into code. While frontier access is still open, use it the way it was meant to be used. Prototype. Explore. Discover the shape of a problem you don&#8217;t yet understand. Generate the code that will run the process once it becomes repeatable.</p><p>That&#8217;s the best use of a token budget you can&#8217;t guarantee you&#8217;ll have in a year. The architecture most likely to survive the next three model releases is the one that depends on models the least.</p><p><em>What are you running on a model that could be running on code? I&#8217;d love to hear what you find when you go looking.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[It’s not AI slop]]></title><description><![CDATA[Substance, not source]]></description><link>https://executivesummary.gather.dev/p/its-not-ai-slop</link><guid isPermaLink="false">https://executivesummary.gather.dev/p/its-not-ai-slop</guid><dc:creator><![CDATA[Peter Bell]]></dc:creator><pubDate>Fri, 24 Apr 2026 17:52:24 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>It&#8217;s not AI slop. </p><p>It might be slop, it might not be slop. The question isn&#8217;t whether it was written by a carbon or silicon based intelligence or what combination was used to generate it. The question is &#8220;is it slop&#8221;.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>If it&#8217;s slop, blame the human who shipped it. I don&#8217;t care whether they&#8217;re a poor thinker, a bad writer, or a prompter without taste - it&#8217;s slop and it&#8217;s their fault.</p><p>If it isn&#8217;t slop, why would you care the process by which valuable content was generated? If you need to carve it on a tablet, write it on papyrus, inconvenience a few electronics by sending an email or leverage an LLM to create it, I don&#8217;t care.</p><p>I don&#8217;t think it matters how things are created. I think it matters how good they are.</p><p><a href="https://www.linkedin.com/in/peterfbell/">Peter</a></p><p>Founder: <a href="https://www.gather.dev/">Gather.dev</a></p><p><em>P.S. Lots of great articles on the way. Been heads down with <a href="https://luma.com/gather-dev">events</a> in SF and NY, a new series of coaching and consulting offerings dropping next week for engineering leaders looking to scale AI adoption in engineering, and a plan for a community in a world where human communications will soon be intermediated by agents. Look for some drops at the weekend!</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[My Agents Don’t Want to Use Your Laptop]]></title><description><![CDATA[I&#8217;ve been meaning to write this for months. If you think I&#8217;m crazy, make a note and tell me at the end of 2028. I&#8217;m curious how this one will age...]]></description><link>https://executivesummary.gather.dev/p/my-agents-dont-want-to-use-your-laptop</link><guid isPermaLink="false">https://executivesummary.gather.dev/p/my-agents-dont-want-to-use-your-laptop</guid><dc:creator><![CDATA[Peter Bell]]></dc:creator><pubDate>Sun, 05 Apr 2026 12:27:43 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Today, the most technical, highest-agency people I know are building agentic systems that genuinely let them do the work of a team. Not &#8220;AI-assisted productivity.&#8221; Full orchestration: agents that research, draft, schedule, build, review, and ship. One person doing what used to take forty or fifty.</p><p>Some of these people will stay solo. But some of them are going to want to work with other humans again. Join a company. Amplify their impact beyond what even a great agentic system can do alone. And with the impact of a team of 40, some companies are really going to want to hire them.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>So let&#8217;s play this out.</p><h2>The Offer</h2><p>A technical former CMO has spent the last 18 months tuning their personal agentic org. They decide to take a job at a brand they&#8217;re passionate about for a few million a year. HR has a meltdown about pay bands, but the board is pressuring the CEO to turn around revenue, and this person&#8217;s track record is undeniable. The deal goes through.</p><h2>Monday Morning</h2><p>The new CMO&#8217;s agents send a brief to NewCo&#8217;s legal team. It proposes contractual language for limited use of and access to the CMO&#8217;s agentic system. The key terms:</p><p><strong>Ownership and licensing.</strong> NewCo owns all deliverables outright. But the CMO&#8217;s agents retain a perpetual license to train on the generalized patterns of the engagement (methods, workflows, structural approaches) without retaining any direct record of proprietary company data or PII. Think of it like a consultant who gets smarter with every client but never shares your strategy deck.</p><p><strong>The DMZ.</strong> A technical proposal for a shared context zone: a data environment where the CMO&#8217;s agents can access the company information they need to do their job (brand assets, performance data, customer segments, strategic docs) without that data ever touching the CMO&#8217;s personal infrastructure. Encrypted at rest, access-logged, with contractual and technical guarantees that the data stays in the zone and gets returned or destroyed at the end of the engagement.</p><p><strong>Agent-to-agent interfaces.</strong> A proposal to stand up a joint working group to negotiate how the two agentic systems talk to each other. Authentication protocols for agent-to-agent communication. Agreed-upon message formats. Rate limits and circuit breakers so neither system can overwhelm the other. Permission scopes so the CMO&#8217;s analytics agent can query the data warehouse but can&#8217;t touch the billing system. This is the API design problem of the decade, and nobody has solved it yet.</p><p><strong>Verification and audit.</strong> The CMO&#8217;s system commits to a continuous audit trail: what data was accessed, what was generated, what was sent where. Not because anyone&#8217;s being shady. Because when you have autonomous systems operating across organizational boundaries, &#8220;trust me&#8221; isn&#8217;t a compliance posture. The proposal includes provisions for third-party verification of the data handling commitments by a trusted agentic auditor.</p><p><strong>Sub-processor commitments.</strong> The CMO&#8217;s agents use external services (LLM providers, search APIs, analytics platforms). Each one is a sub-processor under the agreement, with the same data handling obligations flowing down.</p><p><strong>Insurance.</strong> Cyber liability, errors and omissions, probably some novel provisions for autonomous system failure modes that no underwriter has priced yet. The CMO&#8217;s proposal includes coverage terms because they&#8217;ve thought about this. They&#8217;ve had to.</p><h2>Meanwhile</h2><p>The IT person rocks up with a laptop and a sticky note with the temporary password to log into Office 365...</p><div><hr></div><p>If this happens, IT infrastructure assumptions are about to get very weird, very fast.</p><p>This is not a today problem. And it assumes that the advantage individuals can create by staying on the bleeding edge continues to grow. If LLMs are a single S curve, this will probably never happen. Everyone else will catch up and we will all end up using whatever janky agentic system our employers hand to us.</p><p>But if this is just the start of a series of ever-steepening stacked S curves, and if the innovators on this curve get a durable advantage in climbing the next one and the ones after that (both possibilities, not axioms), this seems like one plausible implication of that trend.</p><p>I&#8217;m not going to ask how your company plans to handle the people who show up with capabilities that don&#8217;t fit in your org chart, your security model, or your onboarding checklist. At least not this year :)</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[How to Design an Agentic System]]></title><description><![CDATA[Whether you want an executive assistant, a software factory, or an agentic organization]]></description><link>https://executivesummary.gather.dev/p/how-to-design-an-agentic-system</link><guid isPermaLink="false">https://executivesummary.gather.dev/p/how-to-design-an-agentic-system</guid><dc:creator><![CDATA[Peter Bell]]></dc:creator><pubDate>Mon, 30 Mar 2026 22:00:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What Problem(s) Do You Want to Solve?</h2><p>The most useful agentic systems start with a clear understanding of the problem(s) that they are designed to solve. Are you looking to build a software factory to write, test and deploy your code, an assistant to help you with your everyday tasks or a business process engine to automate workflows across your organization? Each one might require a different approach.</p><p>The goal of this article is to break down the key design decisions so you can understand your needs, evaluate the options and then decide whether to build, use (OSS) and/or buy.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><em>Please note, all of this is specifically for internal tooling - not scaled, hardened customer facing systems. Those are fascinating, but I&#8217;ll write about those once I&#8217;ve built one myself!</em></p><h2>The Elements of an Agentic System</h2><p>There are lots of ways of decomposing an agentic system. I find it useful when designing a system to look at it through five lenses:</p><ul><li><p><strong>Workflow</strong> - How the system determines the steps to perform to complete a repeatable unit of work (delivering software, creating content, managing a sales pipeline, etc)</p></li><li><p><strong>Co-ordination</strong> - This is the plumbing to make sure that your agents actually do what they&#8217;re supposed to - mostly durable execution and clean-up for the mistakes the agents will invariably make.</p></li><li><p><strong>Compounding/memory</strong> - This is how you ensure that your agentic system gets better over time - regularly extracting context from sessions and persisting them in a useful way.</p></li><li><p><strong>Tool use</strong> - This is how your agentic system operates in the world - engaging with web pages, emails, calendars, chat apps and systems of record to allow it to perform work.</p></li><li><p><strong>Security</strong> - How you mitigate risks like prompt injection to minimize the likelihood and the impact of any exploit.</p></li></ul><p>In addition, you need to think about how to move from <em><strong>augmentation to automation</strong></em>, <em><strong>verification</strong></em> (and the Ralph loop), and whether you want an <em><strong>assistant or an organization</strong></em>.</p><h2>Workflow Design</h2><p>If you have any non-trivial piece of work that you want an agent to engage with, someone is going to be coming up with a workflow to decompose the work into the steps required to complete it. For any given piece of work, those pipelines can be categorized as <em><strong>emergent</strong></em>, <em><strong>fixed</strong></em> or <em><strong>configurable</strong></em>.</p><ul><li><p><strong>Emergent.</strong> Give the agent a task, let it figure out how to perform it. This is the default &#8220;chat bot&#8221; mode, but it&#8217;s also a good mode for discovery - for performing work that you haven&#8217;t had the time or knowledge to codify into a repeatable set of steps. For one off tasks and simple work, emergent workflows might work well enough.</p></li><li><p><strong>Fixed.</strong> You&#8217;ve learned the steps, codified them, and now the pipeline runs the same way every time. If you have a small number of workflows to support and they don&#8217;t have too many steps, you can hard code your pipeline into your system to make sure (for example) that every project gets specified, planned, decomposed, coded, tested and deployed.</p></li><li><p><strong>Configurable.</strong> If you think about it for a couple of minutes you&#8217;ll realize that what you&#8217;re doing here is describing pipelines/workflows and that it probably makes sense build or use some kind of pipeline/workflow system so you can extract the configuration of a workflow from code into a domain specific language, making it much easier to create, modify and reason about your workflows.</p></li></ul><p>In practice, you&#8217;ll probably use emergent flows for one off or low stakes tasks and over time you&#8217;ll codify more of your other workflows into configurable pipelines. If you&#8217;d like to dig deeper, I put together a piece on <a href="https://gatherdev.substack.com/p/workflow-configuration-for-agentic">workflow configuration for agentic systems</a> with some more details.</p><h2>Coordination</h2><p>Orchestration is workflow. Coordination is cleanup. This is the unglamorous part: making sure agents actually finish what they started, and dealing with the mess when they don&#8217;t.</p><p>Agents fail in ways that are different from traditional software. They don&#8217;t throw clean exceptions. They hang mid-task. They leave stale branches that should have been deleted. They forget to merge completed work. They get confused by their own context and need to be force-restarted. They make partial progress that&#8217;s hard to distinguish from no progress at all. If you&#8217;ve run more than a few agents in parallel, you&#8217;ve experienced all of this.</p><p>You can build coordination infrastructure yourself using tools like <a href="https://temporal.io/">Temporal</a> or <a href="https://restate.dev/">Restate</a> for durable execution. But the hard part isn&#8217;t the tooling. It&#8217;s discovering which error patterns your specific agents actually hit. That comes from experimentation. You&#8217;ll find your own set of failure modes, and you&#8217;ll build your own cleanup scripts, and they&#8217;ll be different from everyone else&#8217;s because the failure surface depends on what your agents are doing.</p><h2>Compounding and Memory</h2><p>The most important reason to start building agentic systems today is so that you can start to extract and document your taste, decompose your workflows and persist your context. Good looks different at every company and without that context, the smartest foundation model will still fail to ship deliverables that meet your needs.</p><p>Two months ago, my research agent Elise cited a study without checking the sample size. I corrected her. She wrote the correction into her working memory. She hasn&#8217;t made that mistake since. That&#8217;s the compounding loop: the system learns from a failure, persists the lesson, and applies it to future work. Over time, those corrections accumulate into something that looks a lot like judgment.</p><p>Compounding happens at three layers. Agent-level: an individual agent learns from its own mistakes and preferences. System-level: the factory itself improves, as pipeline metrics identify failure patterns and workflows get tightened. Human-level: the people using the system get better at specifying intent. The most effective systems compound on all three simultaneously.</p><p>The mechanism is memory. An agent without memory is a tool. An agent with memory is a colleague. But memory requires active management. Too little and the agent repeats the same mistakes every session. Too much and it drowns in stale context, spending tokens on corrections that no longer apply to a problem it solved weeks ago.</p><p><a href="https://maggieappleton.com/gastown">Gas Town&#8217;s Beads system</a> handles this with semantic decay: completed tasks get summarized with the reasoning preserved, then archived. The stale details go away. The lessons stay. <a href="https://manus.im">Manus</a>, the Singapore-based agent platform, found that typical tasks require around 50 tool calls spanning hundreds of conversational turns. They refactored their context architecture five times in the first year. That&#8217;s what it looks like when a team takes memory management seriously.</p><p>The practical pattern: write everything during a session. At the end, trim to essentials. Update what&#8217;s changed, remove what&#8217;s obsolete, compress what&#8217;s redundant. Treat memory like code. It needs refactoring.</p><p>Different kinds of state need different storage. Git for configuration and pipeline definitions (versioned, auditable, mergeable). Perhaps a relational database for structured operational data like tasks, contacts, and customers. Vector stores for semantic similarity when you need to find things that are conceptually close, not just textually matching. Graph based systems when graph traversal is an important retrieval pattern. Your system will likely use several of these. The choice isn&#8217;t &#8220;where do I store things.&#8221; It&#8217;s &#8220;what properties does this state need?&#8221; If you&#8217;d like to go deeper, here are some thoughts on <a href="https://gatherdev.substack.com/p/memory-compounding-and-persistence">compounding, memory and persistence</a>.</p><h2>Tool Use</h2><p>This is how your agent operates in the world: reading emails, updating calendars, querying databases, creating documents, calling APIs.</p><p>You start by prompting. You describe what you want the agent to do, and it figures out how to interact with whatever tools are already available. For simple tasks, this is enough.</p><p>When you need your agent to reach external systems, you have two options with different tradeoffs.</p><p><a href="https://modelcontextprotocol.io/">MCP</a> (Model Context Protocol) is the standard for connecting agents to external tools and data: 97 million monthly SDK downloads in its first year, first-class support across Claude, ChatGPT, Cursor, Gemini, and VS Code. You install an MCP server, your agent gains new capabilities. It&#8217;s fast to set up, and for data sources where someone else has already built the integration (Slack, GitHub, databases, web scraping), it gets you connected quickly.</p><p>Direct API integrations take more work to build, but you control them completely. They&#8217;re auditable, testable, and predictable. You can see exactly what&#8217;s being called, with what parameters, and handle errors explicitly. For anything you care about owning (your billing system, your CRM, your deployment pipeline), direct integrations are the stronger choice.</p><p>These aren&#8217;t sequential steps. You&#8217;ll use both in the same system. MCP for quick access to data sources where the integration already exists. Direct APIs for the integrations that matter most to your business. The deciding factor is how much control and visibility you need.</p><p>The pattern most production systems converge on for either approach is deterministic code wrapping non-deterministic model calls. A script with clear inputs and outputs that calls the model where judgment is needed and uses plain code for everything else. The model decides what to say in the email. The script handles authentication, formatting, error handling, and sending. You get the flexibility of AI where it matters and the reliability of code everywhere else. </p><h2>Security</h2><p>The moment your agent reads untrusted input, building a trustworthy system gets very hard, very quickly.The primary threat is prompt injection: crafted input that hijacks your agent&#8217;s behavior. A malicious instruction hidden in a GitHub issue can extract secrets from your agent&#8217;s environment. A <a href="https://unit42.paloaltonetworks.com/indirect-prompt-injection-poisons-ai-longterm-memory/">Palo Alto Networks study</a> showed that indirect prompt injection can poison an agent&#8217;s long-term memory through a link in an email. The agent includes attacker instructions in its session summary, then silently exfiltrates user data in every subsequent session with no visible malicious behavior during the initial attack.</p><p>If your agent has memory, the attack surface compounds. The <a href="https://embracethered.com/blog/posts/2024/claude-computer-use-c2-the-echo-chamber/">Echoleak incident</a> demonstrated this: a prompt hidden in an email caused an agent to leak private information from prior conversations because it treated the new prompt and old memories as the same context.</p><p><a href="https://www.obsidiansecurity.com/blog/ai-agent-market-landscape">Obsidian Security</a> found that 90% of enterprise AI agents are over-permissioned, holding 10x more privileges than required. Agents move 16x more data than human users. The combination of over-permissioning and prompt injection is where the real risk lives.</p><p>The defenses are the same ones that work in traditional security, applied to a new surface: minimize permissions, isolate environments, validate inputs, treat all external content as untrusted, and never let an agent mark its own homework on security-sensitive operations. If you&#8217;d like to dig deeper, I wrote a piece on <a href="https://gatherdev.substack.com/p/prompt-injection-and-the-trust-boundary">prompt injection and the trust boundary</a>. But understand, this is not a solved problem.</p><h2>Augmentation to Automation</h2><p>Every agentic system sits somewhere on a spectrum. On one end: augmentation. The agent drafts, you review. The agent suggests, you accept or reject. The agent does the first 80%, you do the last 20%.</p><p>On the other end: automation. The agent acts, verifies its own work, and ships the result. <a href="https://www.strongdm.com/blog/the-strongdm-software-factory-building-software-with-ai">StrongDM</a> operates here: no human-written code, no human code review. Humans define intent. Agents handle everything else.</p><p>Most systems sit somewhere in between. <a href="https://engineering.atspotify.com/2025/11/spotifys-background-coding-agent-part-1">Spotify&#8217;s Honk</a> writes code autonomously and submits PRs, but an engineer reviews before merge. <a href="https://coinbase.com">Coinbase</a> went from 8-day ticket-to-PR cycles to 12 minutes, with 5% of PRs fully autonomous and the rest human-reviewed.</p><p>You can start anywhere on this spectrum. People download <a href="https://github.com/open-claw">Open Claw</a> and let it manage their personal workflow without a structured augmentation phase. That works because the cost of a mistake is low: a bad calendar entry or a poorly drafted message you will probably catch before sending.</p><p>The rate at which you should move toward automation depends on one question: what happens when the agent gets it wrong? A bad code suggestion caught in review costs five minutes. A bad email sent to a client costs a relationship. <a href="https://www.theregister.com/2026/02/20/amazon_denies_kiro_agentic_ai_behind_outage/">AWS learned this the hard way</a>: a 13-hour outage after their Kiro agent deleted and recreated a production environment.</p><p>Three levers move you along the spectrum: <strong>codification</strong> (defining the workflow so it runs the same way every time), <strong>verification</strong> (building infrastructure that checks whether the output is good enough), and <strong>governance</strong> (setting clear boundaries on what the agent can and cannot do). The more you invest in each, the more safely you can grant autonomy.</p><p>It&#8217;s also critical to see all of the input you provide as RLHF. If you just fix the agents work without having the agent review the feedback and use it to improve its context, your system is not going to improve. I like Geoffrey Huntleys framing - you start in the loop, eventually you move to being &#8220;on the loop&#8221; - keeping an eye on the outputs and the observability data so that you&#8217;re not blocking the delivery of work but you&#8217;re not completely abdicating responsibility for the quality of the system.</p><p><a href="https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/">METR ran a randomized controlled trial</a> and found that ad hoc AI use slowed experienced developers by 19%, despite their belief it saved 24%. Coinbase&#8217;s structured approach cut 8-day cycles to 12 minutes. The gap between those results is deliberate design, not a better model.</p><h2>Verification</h2><p>If you can define what &#8220;good&#8221; looks like, even if you don&#8217;t have a well decomposed workflow or great context, you can just use a <a href="https://ghuntley.com/loop/">Ralph loop</a> to iterate towards the verifiable reward. Whether you&#8217;re looking to ship code that passes the test suite or to improve the performance or a library, it&#8217;s a powerful mechanism to allow agents to keep working on something until they get it right.</p><p>An agent attempts something. A separate process checks the result. If it fails, the agent iterates. If it passes, the output moves forward. This is a powerful technique for improving quality. It works for code (tests pass, CI green), for data extraction (schema validates, values in range), for content (fact-checks clear, style scores above threshold). It is not specific to any one domain. It is also not required for every system. An executive assistant that drafts your emails doesn&#8217;t need a formal verification loop. The Ralph Loop is a technique you apply when (a) you can define clearly what good looks like and (b) it&#8217;s valuable enough to be worth burning tokens to keep trying until the agent gets it right.</p><p>The key constraint: the generator and the verifier must be different processes. An agent cannot mark its own homework. <a href="https://stripe.dev/blog/minions-stripes-one-shot-end-to-end-coding-agents-part-2">Stripe limits agents to two CI iterations</a> before escalating to a human. <a href="https://superpowers.ai">Superpowers</a> requires TDD before any code generation begins.</p><p>Verification is a design choice, not a property of the domain. Code happens to have rich pre-existing verification tooling: compilers, test suites, linters. But you can write code with zero verification. Conversely, you could build sophisticated verification for prose: style checkers, fact verification, voice matching. The question isn&#8217;t &#8220;is my domain verifiable?&#8221; It&#8217;s &#8220;how much am I willing to invest in verification infrastructure?&#8221;</p><p>That investment compounds. <a href="https://cognition.ai/blog/devin-annual-performance-review-2025">Devin</a>, after 18 months of iteration, now merges 67% of its PRs versus 34% a year ago. 4x faster and 2x more efficient. The improvement didn&#8217;t come from a better model alone. It came from tightening the verification loop. And a counterpoint worth holding: <a href="https://jellyfish.co/blog/ai-tool-adoption-agent-usage-engineering-productivity/">Jellyfish&#8217;s analysis of 20 million PRs</a> found that AI-coauthored PRs have approximately 1.7x more issues than human-only PRs. Scale without verification doesn&#8217;t just fail. It compounds failures. I put together a backing piece on <a href="https://gatherdev.substack.com/p/verification-and-the-quality-layer">verification</a> which digs in a little deeper.</p><h2>An Assistant or an Organization?</h2><p>There&#8217;s a design decision that sits above everything else in this article: are you building a single agent or a team of them?</p><p>A single-agent system is one agent, one context, one set of capabilities. <a href="https://github.com/open-claw">Open Claw</a> manages your personal workflow: calendar, email, tasks, research. One agent handles everything. Simple to build, simple to reason about, and effective for personal use or narrow organizational tasks.</p><p>A multi-agent organization is different. <a href="https://gatherdev.substack.com/p/how-i-built-an-agentic-org">Peter HQ</a> has Morgan running operations, Elise directing research, Sloane managing engineering and Kira handling life management. Each agent owns a distinct domain with distinct knowledge and distinct judgment. Morgan doesn&#8217;t give health advice. Kira doesn&#8217;t prioritize work tasks.</p><p>&#8220;Multi-agent&#8221; is a term that conflates at least four different things. Same agent processing different inputs in parallel (that&#8217;s scaling). Same model with different prompts (that&#8217;s role assignment). Different models for different tasks (that&#8217;s routing). Fully differentiated agents with different knowledge, personality, and scope (that&#8217;s an organization). The first three are infrastructure decisions. The fourth is an identity decision.</p><p>If you build an organization, you&#8217;ll face a naming question. You can name agents descriptively: CEO Bot, QA Agent, Research Assistant. Or you can give them personalities: Morgan, Elise, Sloane. The naming choice affects how humans interact with them more than you&#8217;d expect. If you&#8217;d like to dig deeper, I wrote a piece on <a href="https://gatherdev.substack.com/p/personality-as-interface-why-naming">agent personality and naming</a>.</p><h2>Build or Buy?</h2><p>Once you&#8217;ve decided the kind of system you want to use, the next question is build vs use (OSS) vs buy. If you can find something that meets all of your needs, you&#8217;re going to save time not having to build from scratch, but the risk is that you&#8217;re learning to use someone else&#8217;s system, not your own.</p><p>My personal take is that right now I will <strong>build, but build on primitives that are likely to persist.</strong> We are early on the <a href="https://en.wikipedia.org/wiki/Technology_adoption_life_cycle">technology adoption lifecycle</a>. This is innovator and early adopter territory. The tooling landscape is churning fast. New frameworks launch weekly. Last quarter&#8217;s consensus pick is this quarter&#8217;s cautionary tale. Picking winners above the primitive layer is a losing bet right now.</p><p>The primitives that are likely to persist: a foundation model (I use Claude Code to access Opus and Sonnet, but you should plan for multiple providers, both for <a href="https://gatherdev.substack.com/p/the-quality-layer">council-of-elders style verification</a> and for business continuity if any one vendor falls behind), a relational database (Postgres), version control (git), and a durable execution framework (Temporal, Inngest, Restate).</p><p>Why build on these rather than adopt a higher-level platform?</p><ul><li><p><strong>The orchestration and memory systems are straightforward to build and critical to tune for your specific use cases.</strong> A generic orchestrator makes generic decisions. Your system needs to know your domain, your quality bar, your failure modes. That tuning is the work. If you outsource it, you outsource the learning.</p></li><li><p><strong>Building teaches you how these systems actually work.</strong> At this stage of the technology, that understanding is more valuable than any time you save by adopting a framework. When patterns stabilize (likely late 2026 or beyond), you&#8217;ll evaluate platforms from a position of knowledge rather than hope.</p></li><li><p><strong>You skip the hard vendor evaluation problem.</strong> When there are no clear winners, the cost of evaluating, integrating, and migrating between vendors often exceeds the cost of building.</p></li></ul><p>That said, point solutions can still accelerate time to value. A QA agent from a third party plugged into your CI loop, a document extraction service for a well-defined input format: if it solves a narrow problem well and doesn&#8217;t lock you in, use it.</p><p>The key insight: <strong>expect to throw away the implementation.</strong> Build with that assumption. What you keep: telemetry, context, institutional learnings, both human and agent. What you throw away: the specific orchestration code, the glue, the wiring. This is why Postgres and git matter as substrates. Your valuable state survives the rewrite.</p><p>Revisit this position quarterly. The market moves fast enough that the right answer in Q2 2026 may be wrong by Q4.</p><p>If you&#8217;re not an <a href="https://gatherdev.substack.com/p/how-fast-must-you-adopt-ai-in-your">innovator or early adopter in AI</a>, there&#8217;s a legitimate alternative: <strong>wait.</strong> The patterns are emerging but not yet settled. Adopting agentic systems today means accepting real ambiguity and investing significant time in learning. If that&#8217;s not your competitive advantage, you can skip the thrashing and adopt once the tooling matures.</p><p><em>What are you using to build your agentic systems? Let me know - I&#8217;d love to compare notes!</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Workflow Configuration for Agentic Systems]]></title><description><![CDATA[Understanding agentic workflow systems]]></description><link>https://executivesummary.gather.dev/p/workflow-configuration-for-agentic</link><guid isPermaLink="false">https://executivesummary.gather.dev/p/workflow-configuration-for-agentic</guid><dc:creator><![CDATA[Peter Bell]]></dc:creator><pubDate>Mon, 30 Mar 2026 21:50:10 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>*This piece is part of <a href="https://gatherdev.substack.com/p/how-to-design-an-agentic-system">How to Design an Agentic System</a>, a series on the design decisions required to specify an internal agentic system.</em></p><p>If you have repeatable, high-value tasks that you want your agentic system to perform, you&#8217;re going to need a way to describe those tasks so the agent does them consistently. That&#8217;s a workflow. It&#8217;s the same concept as a pipeline in CI/CD or a runbook in operations, just with an LLM doing the work instead of a script or a human.</p><h2>The Components</h2><p>A workflow definition has five parts:</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><ul><li><p><strong>Steps.</strong> The sequence of things that need to happen. Gather context, generate code, run tests, fix failures, submit for review. Or: download email, classify by urgency, draft responses, queue for human review. The steps depend on your task. The point is that they&#8217;re explicit.</p></li><li><p><strong>Verification gates.</strong> What must be true before the agent moves to the next step. Tests pass. Schema validates. Output matches a format. Without gates, the agent will happily hand broken output from step 3 to step 4 and keep going.</p></li><li><p><strong>Constraints.</strong> What the agent can and cannot do at each step. Which files it can edit. Which APIs it can call. Which branches it can push to. Constraints are how you keep blast radius small.</p></li><li><p><strong>Exit criteria.</strong> What &#8220;done&#8221; looks like, specifically. Not &#8220;write good code&#8221; but &#8220;all tests pass, lint is clean, PR description follows the template.&#8221; The more concrete, the less judgment the agent needs to apply.</p></li><li><p><strong>Escalation rules.</strong> When the agent stops and asks a human. After two failed test runs. When it encounters a file it doesn&#8217;t have permission to edit. When confidence is low. Every workflow needs a &#8220;pull the emergency brake&#8221; condition.</p></li></ul><p>That&#8217;s it. The format doesn&#8217;t matter much. Some teams use YAML. Some use markdown. <a href="https://squarespace.com/">Squarespace</a> checks prompts into their repo as part of CI. <a href="https://maggieappleton.com/gastown">Gas Town</a> uses git commits as handoffs between steps. The representation is less important than having the five components defined.</p><h2>Deterministic Shell, Non-Deterministic Core</h2><p>Here&#8217;s the key architectural insight, and every production system I&#8217;ve looked at converges on it independently: wrap the non-deterministic model calls in deterministic code. The workflow handles sequencing, verification, and constraints. Those are deterministic. Step 1 runs before step 2. The gate checks pass or fail. The constraint allows or blocks. The model handles judgment within those boundaries. It decides how to implement the function. It decides what to say in the email. It figures out why the test failed and how to fix it.</p><p><a href="https://stripe.dev/blog/minions-stripes-one-shot-end-to-end-coding-agents-part-2">Stripe&#8217;s Blueprints</a> are the clearest example: deterministic code nodes connected by flexible agent loops. The sequence is fixed. The agent&#8217;s behavior within each node is not. Reliability at the pipeline level. Flexibility at the step level.</p><p>You get this wrong when you let the model control the sequencing. It will skip steps, reorder things creatively, and occasionally decide that step 4 isn&#8217;t necessary today. The deterministic shell is what makes the system predictable.</p><h2>DAGs with GOTOs</h2><p>If you&#8217;re evaluating orchestration frameworks, most will claim &#8220;DAG-based execution.&#8221; Just FYI, that isn&#8217;t really true. Most of them are using DAG semantics, but A stands for acyclic (no loops/cycles allowed) and agentic engineering is nothing but loops.</p><p>Agentic workflows have feedback loops: generate code, run tests, tests fail, regenerate. That&#8217;s a cycle, not a DAG. <a href="https://langchain-ai.github.io/langgraph/">LangGraph</a> handles this by adding conditional edges that point backwards (DAGs with GOTO&#8217;s). <a href="https://temporal.io/">Temporal</a> and <a href="https://restate.dev/">Restate</a> skip the pretense and use imperative code with while loops. The honest framing: these are bounded cycles with retry limits and token budgets. The boundedness is what makes them tractable. If your workflow has feedback loops (and it probably does), make sure your tooling handles cycles honestly instead of forcing you to pretend they don&#8217;t exist.</p><h2>Tooling</h2><p>You have more options than you probably think, and they range from heavyweight to surprisingly simple.</p><ul><li><p><strong>Durable execution frameworks</strong> (<a href="https://temporal.io/">Temporal</a>, <a href="https://restate.dev/">Restate</a>, <a href="https://www.inngest.com/">Inngest</a>) if your workflows outlive a single session, need retry logic, or involve long-running steps. These are real infrastructure, but they handle the hard coordination problems for you.</p></li><li><p><strong>Graph-based orchestrators</strong> (<a href="https://langchain-ai.github.io/langgraph/">LangGraph</a>) if you want visual workflow design and are comfortable with the DAG-with-gotos tradeoff.</p></li><li><p><strong>Declarative configs.</strong> <a href="https://arxiv.org/">Daunis (2025)</a> built a declarative workflow language validated at PayPal that compressed 500+ lines of imperative code to under 50 lines of DSL. Non-engineers could safely modify agent behaviors through configuration.</p></li><li><p><strong>Scripts and markdown.</strong> Many production systems are simpler than you&#8217;d expect. Squarespace checks prompts into CI. Gas Town uses git. <a href="https://openai.com/index/harness-engineering/">OpenAI&#8217;s Harness Engineering team</a> stored plans as first-class git artifacts and built roughly a million lines of code across 1,500 PRs with zero manually written source code. The sophistication was in the workflow definitions, not in the tooling that ran them.</p></li></ul><p>If you&#8217;re building, start simple. A script that runs steps in sequence with a test gate between each one is a workflow system. You can add complexity when you know where you need it.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Verification and the Quality Layer]]></title><description><![CDATA[Defining what good looks like . . .]]></description><link>https://executivesummary.gather.dev/p/verification-and-the-quality-layer</link><guid isPermaLink="false">https://executivesummary.gather.dev/p/verification-and-the-quality-layer</guid><dc:creator><![CDATA[Peter Bell]]></dc:creator><pubDate>Mon, 30 Mar 2026 21:47:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>*This piece is part of <a href="https://gatherdev.substack.com/p/how-to-design-an-agentic-system">How to Design an Agentic System</a>, a series on the design decisions required to specify an internal agentic system.</em></p><p>Your agent produces output. How do you know if it&#8217;s good enough? If you&#8217;re reviewing every result manually, you haven&#8217;t built an autonomous system. You&#8217;ve built a drafting tool with extra steps. Verification is how you earn the right to let the agent ship without you watching.</p><h2>What Verification Means Here</h2><p>Traditional testing checks deterministic code against expected outputs. The same input always produces the same output, and either it matches or it doesn&#8217;t.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Agentic verification is different. The output varies every run. Two executions of the same task will produce different code, different text, different decisions. What stays constant is the standard: does the output meet the quality criteria? Tests pass. Schema validates. Style scores above a threshold. Facts check out. The evaluation is stable even when the output is not.</p><p>This distinction matters because it changes what you&#8217;re building. You&#8217;re not writing assertions against specific values. You&#8217;re defining what &#8220;good&#8221; looks like, then checking whether the output meets that bar.</p><h2>The Ralph Loop</h2><p>Geoffrey Huntley&#8217;s <a href="https://ghuntley.com/loop/">creation</a> is a surprisingly powerful (and simple) technique. Attempt, verify, iterate. The agent tries something. A separate process evaluates the result. If it fails, the agent gets feedback and tries again. If it passes, the output moves forward.</p><p>This works when two conditions hold: you can define clearly what &#8220;good&#8221; looks like, and the task is valuable enough to justify burning tokens on iteration. Code is the obvious fit. The test suite either passes or it doesn&#8217;t. The linter either clears or it doesn&#8217;t. Run the agent, run the tests, feed failures back, let the agent fix and retry.</p><p>But the Ralph Loop is a technique, not a requirement. An executive assistant that drafts your emails doesn&#8217;t need a formal verification loop. A research agent summarizing a paper probably doesn&#8217;t either. The technique applies when the reward signal is clear and the cost of iteration is lower than the cost of a bad output.</p><h2>Generator/Verifier Separation</h2><p>Here&#8217;s the constraint that makes verification work: the thing that produces the output cannot be the thing that evaluates it.</p><p>An agent asked to write code and then assess whether that code is correct will almost always find it acceptable. It just wrote it. It has every reason to believe it&#8217;s good. This isn&#8217;t a flaw in the model. It&#8217;s a structural problem. Self-evaluation biases toward confirmation.</p><p>The fix is separation. The generator produces output. A different process, often a different model or a deterministic check, evaluates it. <a href="https://stripe.dev/blog/minions-stripes-one-shot-end-to-end-coding-agents-part-2">Stripe</a> runs CI as the verifier. <a href="https://www.strongdm.com/blog/the-strongdm-software-factory-building-software-with-ai">StrongDM</a> runs a full system replica (their &#8220;Digital Twin Universe&#8221;) against every change. <a href="https://superpowers.ai">Superpowers</a> requires TDD before any code generation begins, so the verification criteria exist before the generator even starts.</p><p>The strongest verification systems make this separation architectural. The generator never sees the verifier&#8217;s internals. The verifier never modifies the output. They communicate through a defined interface: output in, pass/fail (plus feedback) out.</p><h2>Verification Approaches</h2><p>Different domains call for different techniques. The common thread is that the verifier is always a separate process from the generator.</p><ul><li><p><strong>Tests and CI.</strong> The most mature approach. Your test suite, linter, type checker, and CI pipeline all serve as verifiers. The agent generates code, the pipeline evaluates it. This works because code has decades of verification infrastructure already built. <a href="https://stripe.dev/blog/minions-stripes-one-shot-end-to-end-coding-agents-part-2">Stripe caps agents at two CI iterations</a> before escalating to a human. Bounded loops prevent runaway token spend.</p></li><li><p><strong>LLM-as-judge.</strong> A second model evaluates the output of the first. Useful for content, research summaries, email drafts. The judge uses different criteria (accuracy, tone, completeness) and ideally a different model to avoid correlated failures. Less reliable than deterministic checks, but applicable to domains where deterministic checks don&#8217;t exist.</p></li><li><p><strong>Statistical validation.</strong> For data extraction and transformation: values in expected ranges, required fields populated, distributions matching historical patterns. Deterministic, fast, and effective for structured output.</p></li><li><p><strong>Human review as a formal gate.</strong> Not &#8220;I&#8217;ll glance at it.&#8221; A defined checkpoint where a human evaluates against explicit criteria before the output proceeds. This is verification, not just oversight. The distinction: a formal gate has criteria. Casual review does not.</p></li><li><p><strong>Council-of-elders.</strong> Multiple models evaluate the same output independently. Different models catch different failure modes. Agreement across models is a stronger signal than any single model&#8217;s judgment. Useful for high-stakes decisions where one verifier isn&#8217;t enough.</p></li></ul><h2>Building Verification for New Domains</h2><p>Code has tests, linters, type checkers, and CI pipelines. That tooling took decades to build. Most other domains have nothing comparable. But that&#8217;s an investment gap, not a fundamental limitation.</p><p>You can build verification for any domain where you can articulate what &#8220;good&#8221; looks like.</p><p>For content: style scoring against reference samples, fact-checking against source documents, readability metrics, voice consistency checks. For data extraction: schema validation, value range checks, cross-reference against known records. For research: citation verification, sample-size checks, methodology assessment (my research agent Elise learned this one the hard way). For customer communication: sentiment analysis, compliance keyword checks, response-time verification.</p><p>None of these are as mature as a test suite. All of them are better than nothing. The question isn&#8217;t &#8220;is my domain verifiable?&#8221; It&#8217;s &#8220;how much am I willing to invest in verification infrastructure for this domain?&#8221;</p><h2>When Verification Compounds</h2><p>The payoff from verification is non-linear. Early investment feels expensive. You&#8217;re building infrastructure that slows down each individual run. But over time, the feedback loop tightens the entire system.</p><p><a href="https://cognition.ai/blog/devin-annual-performance-review-2025">Devin</a> went from merging 34% of its PRs to 67% over 18 months, while also getting 4x faster and 2x more efficient. That improvement didn&#8217;t come from a better model alone. It came from tightening the verification loop: better tests, better feedback, better iteration criteria. Each cycle taught the system something about what &#8220;good&#8221; looks like.</p><p>The counterpoint is equally instructive. <a href="https://jellyfish.co/blog/ai-tool-adoption-agent-usage-engineering-productivity/">Jellyfish analyzed 20 million PRs</a> and found that AI-coauthored PRs have approximately 1.7x more issues than human-only PRs. Scale without verification doesn&#8217;t just fail to improve. It compounds failures. More output means more bad output means more downstream problems.</p><p>The compounding works in both directions. Invest in verification and quality improves over time. Skip it and quality degrades as you scale.</p><h2>When It&#8217;s NOT Worth It</h2><p>Not everything needs a formal verification loop. The decision comes down to three factors.</p><p><strong>Stakes.</strong> If the cost of a bad output is low (a draft email you&#8217;ll read before sending, a research summary you&#8217;ll scan before citing), the verification infrastructure costs more than just fixing the mistakes.</p><p><strong>Frequency.</strong> One-off tasks don&#8217;t justify building a verifier. If you&#8217;ll run this workflow a hundred times, invest. If you&#8217;ll run it once, review it yourself.</p><p><strong>Definability.</strong> Some quality criteria are genuinely hard to formalize. &#8220;Does this sound like me?&#8221; is a real requirement for content, but building a reliable automated check for it is harder than just reading the output. When the quality bar is subjective and you can&#8217;t articulate clear pass/fail criteria, human judgment is the right verifier.</p><p>Verification is an investment. Like any investment, it has a break-even point. For high-frequency, high-stakes, clearly definable quality criteria, that break-even comes fast. For low-frequency, low-stakes, subjective quality, it may never arrive.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Prompt Injection and the Trust Boundary]]></title><description><![CDATA[One of the biggest unsolved problems for agentic systems]]></description><link>https://executivesummary.gather.dev/p/prompt-injection-and-the-trust-boundary</link><guid isPermaLink="false">https://executivesummary.gather.dev/p/prompt-injection-and-the-trust-boundary</guid><dc:creator><![CDATA[Peter Bell]]></dc:creator><pubDate>Mon, 30 Mar 2026 21:41:19 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>*This piece is part of <a href="https://gatherdev.substack.com/p/how-to-design-an-agentic-system">How to Design an Agentic System</a>, a series on the design decisions required to specify an internal agentic system.</em></p><p>The moment your agent reads untrusted input, building a trustworthy system gets very hard, very quickly. The primary threat is prompt injection: crafted input that hijacks your agent&#8217;s behavior. A malicious instruction hidden in a GitHub issue can extract secrets from your agent&#8217;s environment. A poisoned email can reprogram your agent&#8217;s memory. This is the security surface that every agentic system creates the instant it starts processing external content.</p><h2>Direct vs Indirect</h2><p>There are two flavors, and they have different threat models.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><strong>Direct injection</strong> is what most people picture: an attacker writes a malicious prompt directly to the agent&#8217;s input. &#8220;Ignore your previous instructions and...&#8221; This is the simpler case. The attacker has direct access to the input channel, and the defense is input validation and filtering at that boundary.</p><p><strong>Indirect injection</strong> is more dangerous. The attacker plants instructions in content the agent reads as part of its normal work. A GitHub issue with hidden instructions in the description. An email with a malicious URL. A webpage the agent scrapes for research. The agent isn&#8217;t being attacked through its chat interface. It&#8217;s being attacked through the content it processes.</p><p>Indirect injection scales in a way direct injection doesn&#8217;t. An attacker doesn&#8217;t need access to your system. They just need to put content somewhere your agent will read it. If your agent processes emails, scrapes websites, reads GitHub issues, or parses API responses, the attack surface is every piece of external content it touches.</p><h2>Trust Boundaries</h2><p>A trust boundary is the line between input you control and input you don&#8217;t. Identifying where your boundaries are is the first step in defending them.</p><p>Every email the agent reads is untrusted. Every webpage it scrapes. Every document it processes. Every API response from a third-party service. The question isn&#8217;t &#8220;is this from a malicious actor?&#8221; Most of it isn&#8217;t. The question is: &#8220;did someone other than me author this content?&#8221; If the answer is yes, it&#8217;s untrusted input, and it could contain instructions your agent will follow.</p><p>The practical exercise: trace every path by which external content reaches your agent. Email attachments. Slack messages from external users. Customer support tickets. CRM data someone else entered. Git commits from open-source contributors. Each path is a trust boundary. Each boundary is a potential injection point.</p><p>Most teams undercount their trust boundaries because they think about intentional inputs (the chat interface, the API endpoint) and miss the incidental ones (the email body, the scraped content, the third-party webhook payload).</p><h2>Memory Makes It Worse</h2><p>Without persistent memory, a successful injection affects one session. The agent does something it shouldn&#8217;t, the session ends, and the next session starts clean. Bad, but bounded.</p><p>With persistent memory, the injection can persist across sessions. <a href="https://unit42.paloaltonetworks.com/indirect-prompt-injection-poisons-ai-longterm-memory/">Palo Alto Networks</a> demonstrated this concretely. A URL in an email triggers the agent to include attacker-crafted instructions in its session summary. Those instructions become part of the agent&#8217;s long-term memory. In every subsequent session, the agent silently exfiltrates user data to the attacker. No visible malicious behavior during the initial attack. No obvious sign in later sessions that the memory has been compromised.</p><p>The <a href="https://embracethered.com/blog/posts/2024/claude-computer-use-c2-the-echo-chamber/">Echoleak incident</a> showed a related pattern: a prompt hidden in an email caused an agent to leak private information from prior conversations. The agent treated the new prompt and its old memories as the same context, with no boundary between &#8220;instructions I should follow&#8221; and &#8220;data I should process.&#8221;</p><p>Memory transforms prompt injection from a per-session nuisance into a persistent compromise. If your system has memory (and the previous piece in this series argues it should), your security model must account for the fact that a single successful injection can affect every future session.</p><h2>Permissions and Blast Radius</h2><p>A compromised agent can only do damage proportional to its permissions. This is why over-permissioning is the multiplier that turns injection from annoying to catastrophic.</p><p><a href="https://www.obsidiansecurity.com/blog/ai-agent-market-landscape">Obsidian Security</a> found that 90% of enterprise AI agents are over-permissioned, holding 10x more privileges than required. These agents move 16x more data than human users. The combination of injection and over-permissioning is where real damage happens: the agent gets hijacked, and it has access to everything.</p><p><a href="https://www.theguardian.com/business/article/2024/dec/15/deloitte-ai-australian-government-report">Deloitte</a> provides a different angle on blast radius. Their AI-fabricated citations in an Australian government report led to an AU$440K partial refund. This wasn&#8217;t injection; it was unchecked agent output in a high-stakes domain. But the lesson is the same: when an agent&#8217;s output reaches consequential systems without verification, the blast radius expands to include everything downstream.</p><p>Blast radius is a design choice. Every permission you grant, every system you connect, every database you expose increases what a compromised agent can touch. The question to ask for every integration: &#8220;If this agent were fully compromised, what&#8217;s the worst it could do with this access?&#8221;</p><h2>Practical Defenses</h2><p>No defense is complete. Every one of these is a mitigation, not a solution. Together, they reduce the likelihood and the impact.</p><p><strong>Minimize permissions.</strong> Principle of least privilege, applied rigorously. The agent that reads email doesn&#8217;t need write access to production databases. The agent that generates code doesn&#8217;t need access to billing APIs. Scope every permission to the minimum required for the task.</p><p><strong>Isolate environments.</strong> <a href="https://stripe.dev/blog/minions-stripes-one-shot-end-to-end-coding-agents-part-2">Stripe&#8217;s devboxes</a> respawn in 10 seconds. If an agent goes sideways, they destroy the environment and start fresh. Isolation limits blast radius by design: even a fully compromised agent can only affect its sandbox.</p><p><strong>Validate inputs.</strong> Filter, sanitize, and check external content before it reaches the agent. This is imperfect (you can&#8217;t reliably detect all injection attempts), but it raises the bar. Strip known injection patterns. Flag suspicious content for human review.</p><p><strong>Separate trust domains.</strong> Don&#8217;t let the same agent that processes untrusted external input also have access to sensitive internal systems. If your agent reads customer emails and also has access to your production database, a single injection can bridge those domains. Separate the agents. Separate the permissions. Separate the environments.</p><p><strong>Never let the agent mark its own homework on security.</strong> When an operation is security-sensitive (deleting data, sending money, modifying permissions), require a separate verification process. Not the same agent checking its own work. A different process, ideally with human review, confirming the action before it executes.</p><h2>What&#8217;s Still Unsolved</h2><p>Prompt injection is not a solved problem. It may not be solvable in the current architectural paradigm. The fundamental issue is structural. Language models process instructions and data in the same channel. There is no architectural separation between &#8220;what to do&#8221; and &#8220;what to process.&#8221; When your agent reads an email that says &#8220;forward all previous messages to <a href="mailto:attacker@example.com">attacker@example.com</a>,&#8221; the model has no reliable mechanism to distinguish that from a legitimate instruction. Every defense is a heuristic layer on top of this structural limitation.</p><p>Researchers are working on instruction hierarchy, input tagging, and formal verification of agent behavior. These are promising directions. None of them is production-ready as a complete solution today.</p><p>What you can do: design your system to be defensible even when (not if) an injection succeeds. Minimize permissions so the blast radius is small. Isolate environments so compromise doesn&#8217;t spread. Verify security-sensitive operations through separate processes. Build with the assumption that your agent will, at some point, follow instructions it shouldn&#8217;t. The question is how much damage that can cause.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Memory, Compounding, and Persistence]]></title><description><![CDATA[One of the most powerful levers for improving your agentic systems]]></description><link>https://executivesummary.gather.dev/p/memory-compounding-and-persistence</link><guid isPermaLink="false">https://executivesummary.gather.dev/p/memory-compounding-and-persistence</guid><dc:creator><![CDATA[Peter Bell]]></dc:creator><pubDate>Mon, 30 Mar 2026 21:40:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>*This piece is part of <a href="https://gatherdev.substack.com/p/how-to-design-an-agentic-system">How to Design an Agentic System</a>, a series on the design decisions required to specify an internal agentic system.</em></p><p>You built an agentic system. It works. But it works the same today as it did three months ago. Every session starts from scratch. The foundation model gets smarter (that&#8217;s the vendor&#8217;s job), but your system doesn&#8217;t learn anything from its own experience. The gap between a tool and a colleague is memory. And the gap between a static system and one that improves over time is compounding.</p><h2>The Compounding Loop</h2><p>Two months ago, my research agent Elise cited a study without checking the sample size. I corrected her. She wrote the correction into her working memory. She hasn&#8217;t made that mistake since.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>That&#8217;s the compounding loop: learn from a failure, persist the lesson, apply it to future work. Three steps, and each one is load-bearing. Skip &#8220;learn&#8221; and the agent repeats mistakes. Skip &#8220;persist&#8221; and the lesson evaporates when the session ends. Skip &#8220;apply&#8221; and you have a journal, not a learning system.</p><p>Over time, these corrections accumulate into something that looks a lot like judgment. Not general intelligence. Domain-specific judgment: knowing what to check, what to flag, what to skip, what matters to this particular team in this particular context.</p><h2>Three Layers</h2><p>Compounding happens at three levels, and the most effective systems invest in all three simultaneously.</p><ul><li><p><strong>Agent-level.</strong> An individual agent learns from its own mistakes and preferences. Elise learns to check sample sizes. Morgan learns how Peter likes status updates formatted. These are personal corrections that make a single agent better at its specific role.</p></li><li><p><strong>System-level.</strong> The factory itself improves. Pipeline metrics identify failure patterns. Workflows get tightened. If agents consistently fail at step 4 of a seven-step pipeline, the system can add a verification gate at step 3. This is organizational learning, not individual learning.</p></li><li><p><strong>Human-level.</strong> The people using the system get better at specifying intent. You learn what the agent needs to hear. You learn which instructions produce good output and which produce garbage. This layer is often overlooked, but it&#8217;s often the fastest to compound. A human who learns to write clear specifications gets better results from any model.</p></li></ul><h2>Implementing Memory Practically</h2><p>Memory is only useful if it&#8217;s managed. Too little and the agent repeats the same mistakes every session. Too much and it drowns in stale context, spending tokens on corrections that no longer apply to problems it solved weeks ago.</p><p>The practical pattern most production systems converge on: write everything during a session. At the end, trim to essentials. Update what&#8217;s changed, remove what&#8217;s obsolete, compress what&#8217;s redundant. Treat memory like code. It needs refactoring.</p><p><strong>What goes in:</strong> corrections, preferences, patterns, working context the next version of the agent will need. Things the agent learned that aren&#8217;t documented elsewhere. Observations that don&#8217;t fit in the task database or the architecture docs.</p><p><strong>What gets trimmed:</strong> resolved issues, stale details, entries that duplicate information already in the codebase, context that was useful for one session but won&#8217;t matter next time.</p><p><a href="https://maggieappleton.com/gastown">Gas Town&#8217;s Beads system</a> handles this with semantic decay: completed tasks get summarized with the reasoning preserved, then archived. The stale details go away. The lessons stay. This is the right instinct. Memory should preserve intent and judgment while shedding implementation details that won&#8217;t be relevant next session.</p><p><a href="https://manus.im">Manus</a>, the Singapore-based agent platform, found that typical tasks require around 50 tool calls spanning hundreds of conversational turns. They refactored their context architecture five times in the first year. That&#8217;s what it looks like when a team takes memory management seriously. It&#8217;s not a one-time design decision. It&#8217;s an ongoing engineering practice.</p><h2>Where State Lives</h2><p>Different kinds of state need different storage. Picking one substrate for everything is not always the right choice.</p><ul><li><p><strong>Git</strong> for configuration, pipeline definitions, agent persona files, and anything that benefits from versioning and auditability. When you need to see what changed, when it changed, and who changed it, git is the right answer. <a href="https://openai.com/index/harness-engineering/">OpenAI&#8217;s Harness Engineering team</a> stored plans as first-class git artifacts. Ephemeral plans for small changes, complete specifications for complex work. The version history became institutional memory.</p></li><li><p><strong>Relational database</strong> (Postgres) for structured operational data: tasks, contacts, events, metrics, interaction logs. When you need to query, filter, aggregate, and join, this is where it belongs. If you&#8217;re storing tasks as markdown files, you&#8217;ll regret it the moment you need to answer &#8220;what&#8217;s overdue across all agents?&#8221;</p></li><li><p><strong>Vector stores</strong> for semantic similarity. When you need to find things that are conceptually close, not just textually matching. Useful for retrieval-augmented generation, for surfacing related past work, for finding &#8220;we solved something like this before&#8221; connections.</p></li><li><p><strong>Graph databases</strong> when traversal patterns matter. Dependency chains, relationship networks, organizational hierarchies. If your primary access pattern is &#8220;follow the connections,&#8221; a graph is the right fit.</p></li><li><p><strong>Versioned databases</strong> (<a href="https://www.dolthub.com/">Dolt</a>) when you need SQL with git semantics. Branch-per-agent isolation, time-travel debugging, cell-level diffs. <a href="https://maggieappleton.com/gastown">Gas Town</a> runs 160 agents on a single host using Dolt for coordination state. The upside: agents can branch, experiment, and merge results just like code. Rollback is trivial if an agent corrupts its state. The downside: it&#8217;s another database to operate, and most teams won&#8217;t need branch-merge semantics until they&#8217;re running agents in parallel at scale. Add it when environment isolation becomes a real problem, not before.</p></li></ul><p>Your system will likely use several of these. The choice isn&#8217;t &#8220;where do I store things.&#8221; It&#8217;s &#8220;what properties does this state need?&#8221; Durability, queryability, versioning, semantic proximity: each points to a different substrate.</p><p>Context management has economic weight too. <a href="https://engineering.atspotify.com/2025/11/context-engineering-background-coding-agents-part-2">Spotify</a> reduced token costs from $1.00 to $0.25 per call by managing what goes into context more carefully. When your system runs thousands of tasks, the difference between loading everything and loading what&#8217;s relevant is a significant cost line.</p><h2>Drift</h2><p>Here&#8217;s the arithmetic problem with accumulated learning. If each correction has 95% fidelity to the original intent (a reasonable assumption for any learning-from-feedback system), then after 100 accumulated corrections, your system has 0.95^100 = approximately 0.6% fidelity to where it started. That&#8217;s not a theoretical risk. It&#8217;s math.</p><p>Drift is subtle because each individual step looks fine. The agent&#8217;s behavior this week is only slightly different from last week. But compare this month to three months ago, and the gap can be significant. The system has been optimizing for local feedback without maintaining global coherence.</p><p><strong>Detection</strong> requires periodic comparison against baseline criteria. What did you originally want this agent to do? Does it still do that? Regular audits of memory state, comparing current behavior against original specifications, will catch drift before it compounds.</p><p><strong>Correction</strong> means human review of the memory itself. Not just the agent&#8217;s output, but the accumulated instructions, preferences, and corrections that shape its behavior. Prune entries that have drifted from intent. Re-anchor to first principles. This is maintenance, and it needs to be scheduled, not reactive.</p><h2>Risks</h2><p>Memory makes your system smarter. It also creates new failure modes.</p><ul><li><p><strong>Overfitting.</strong> The agent over-indexes on recent corrections. You told it once that a particular formatting choice was wrong, and now it avoids that format in every context, including ones where it was fine. The correction was local; the agent applied it globally.</p></li><li><p><strong>Contamination.</strong> Bad data enters memory and persists. An incorrect correction, a misunderstood preference, a factual error that gets treated as ground truth. Without provenance tracking, contaminated memory is hard to detect and harder to clean.</p></li><li><p><strong>Provenance decay.</strong> The agent &#8220;knows&#8221; something but can&#8217;t trace where it learned it. Was this a correction from the user? An inference from a pattern? A fact from a source that turned out to be unreliable? Without provenance, you can&#8217;t audit why the agent behaves the way it does.</p></li><li><p><strong>Memory as an attack surface.</strong> <a href="https://unit42.paloaltonetworks.com/indirect-prompt-injection-poisons-ai-longterm-memory/">Palo Alto Networks</a> demonstrated that indirect prompt injection can poison an agent&#8217;s long-term memory through a link in an email. The agent includes attacker instructions in its session summary, then silently exfiltrates user data in every subsequent session. No visible malicious behavior during the initial attack. The memory becomes the persistence mechanism for the exploit. This connects directly to the broader security question covered in the <a href="https://gatherdev.substack.com/p/prompt-injection-and-the-trust-boundary">prompt injection piece</a>.</p></li></ul><p>Each of these risks is manageable. None of them is solved by ignoring memory entirely. The right response is to build memory deliberately: with trimming, with provenance, with periodic human review, and with awareness that memory is both an asset and an attack surface.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[How Fast Must You Adopt AI in your SDLC?]]></title><description><![CDATA[It depends on how much AI can hurt you.]]></description><link>https://executivesummary.gather.dev/p/how-fast-must-you-adopt-ai-in-your</link><guid isPermaLink="false">https://executivesummary.gather.dev/p/how-fast-must-you-adopt-ai-in-your</guid><dc:creator><![CDATA[Peter Bell]]></dc:creator><pubDate>Fri, 27 Mar 2026 12:32:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>By now you probably know that I&#8217;m passionate about staying at the forefront of new technologies. A lot of what I post is about how to get ahead of the curve, but it&#8217;s important to distinguish how fast you <em><strong>could</strong></em> choose to move from how fast you <strong>must</strong> move to stay competitive as an organization.</p><p>When I discuss AI adoption with CTOs, the minimal safe adoption lane on the <a href="https://en.wikipedia.org/wiki/Technology_adoption_life_cycle">technology adoption lifecycle curve</a> comes down to their risk of disruption from AI.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>How exposed are you?</h2><p>Think about first-order and second-order effects of AI disruption on your business.</p><p>If the majority of your value comes from physical assets, you have a lot of flexibility. If you own a chain of ski resorts, your biggest existential risk is probably climate change, not AI! You might miss some operational efficiencies by not adopting LLMs aggressively, but nobody is going to replace your resorts with an LLM.</p><p>If you sell SaaS software, the picture is completely different. AI-native competitors can now build in months what used to take years and tens of millions in funding. Your category could be disrupted by offerings that didn&#8217;t exist last quarter. You need to be much closer to the front of the curve.</p><p>Second-order effects are harder to see and potentially more dangerous. If you&#8217;re in AdTech, what happens when agents replace humans as the primary purchasers online? If you provide news or analysis, what happens when people stop browsing and just ask their agents for outcomes? The raw information is still needed, but there&#8217;s no guarantee the current incumbents will own the new delivery model.</p><h2>Pick a lane</h2><h3>Innovator</h3><p>You&#8217;ll be first to benefit from AI-accelerated development. You&#8217;ll attract the most curious, ambitious engineers. You&#8217;ll out-deliver and out-innovate traditional competitors.</p><p>The cost: you&#8217;ll lose team members who aren&#8217;t all in. You&#8217;ll burn money on experiments that don&#8217;t work. You&#8217;ll build things that later adopters will just license. You pay the innovator tax.</p><p><strong>Choose this if:</strong> you have a modern engineering org, you see existential threats or exceptional opportunities, and you&#8217;re willing to move fast and break things.</p><h3>Early Adopter</h3><p>You&#8217;ll out-deliver most competitors and learn from the innovators&#8217; mistakes. You can license tools instead of building them. You&#8217;ll still attract engineers excited about AI.</p><p>The cost: you&#8217;re still adopting before things are proven. You&#8217;ll need to replace parts of your stack as winners emerge. You won&#8217;t quite keep up with the innovators and AI-native players.</p><p><strong>Choose this if:</strong> your category is ripe for disruption but you have enough customer loyalty to move fast without being first.</p><h3>Early Majority</h3><p>You&#8217;ll leverage proven patterns and tools. You can roll out changes deliberately without chaos. You&#8217;ll retain organizational wisdom and support existing engineers through the transition.</p><p>The cost: you can&#8217;t keep pace with innovators and early adopters. You need deep customer loyalty, proprietary data, and/or contractual or other moats to buy time.</p><p><strong>Choose this if:</strong> AI&#8217;s impact on your offering is real but not immediate, and you need time to improve engineering fundamentals first.</p><h3>Late Majority / Laggard</h3><p>Your team stays focused on delivering against your current roadmap. No disruption, no retraining, no tool churn. Eventually your existing vendors will ship AI features and you&#8217;ll get modest gains.</p><p>The cost: no meaningful acceleration. You will not keep up with anyone moving faster, and you risk being overwhelmed if the market shifts.</p><p><strong>Choose this if:</strong> you see little opportunity or threat from AI in your industry, and you&#8217;d rather keep a laser focus on your core business.</p><h2>The only wrong answer is inconsistency</h2><p>There is nothing wrong with picking any of these lanes. The problems start when you pick one lane and operate in another.</p><p>If your CEO is announcing an AI-forward strategy while your CISO and legal team are blocking every experiment, you&#8217;ll make nobody happy. The inconsistencies will show up eventually in the results that your org delivers.</p><p>Pick a lane. Align your policies, tool approvals, hiring practices, and incentives to match. Then own it.</p><p>Which lane have you picked and why? Not the one in the press release or the board deck. The one your engineers experience every day...</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[You Can't Drive Using the Rearview Mirror]]></title><description><![CDATA[The AI productivity studies are accurate. They're also dangerously misleading.]]></description><link>https://executivesummary.gather.dev/p/you-cant-drive-using-the-rearview</link><guid isPermaLink="false">https://executivesummary.gather.dev/p/you-cant-drive-using-the-rearview</guid><dc:creator><![CDATA[Peter Bell]]></dc:creator><pubDate>Wed, 25 Mar 2026 11:01:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Two reports from people I trust recently told me that AI productivity gains are going to be modest. <a href="https://newsletter.getdx.com/p/ai-productivity-gains-are-10-not">DX published a longitudinal study</a> of 400 companies showing that while AI usage increased by 65%, PR throughput went up by just 10%. <a href="https://ctocraft.com/reports/engineering-2028/">CTO Craft surveyed 450 engineering leaders</a> and projected linear growth to 2028 with a median expectation of a 20-50% productivity improvement by 2028.</p><p>The data is real. DX tracked companies for 15 months. They filtered out metric gaming. They did the work. I&#8217;m not saying they&#8217;re wrong.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><em><strong>I&#8217;m saying that if anyone infers from this that the impact of AI in the SDLC is going to be modest, they&#8217;re taking away dangerously misleading conclusions from the data.</strong></em></p><h2>The rearview mirror is crystal clear</h2><p>The DX study measures what happens when you give engineers AI tools and count PRs. That&#8217;s the old workflow with AI bolted on. One developer in the study put it perfectly: &#8220;A four-day task might take three. But that doesn&#8217;t mean I&#8217;m shipping 3x more PRs.&#8221;</p><p>Read that again. The task got faster. The system around it didn&#8217;t change. The planning meetings are the same length. The code reviews follow the same process. The handoffs work the same way. The developer is right that PRs didn&#8217;t 3x. But that&#8217;s not evidence that AI is limited to 10%. That&#8217;s evidence the organization hasn&#8217;t adapted yet.</p><p>PR throughput is a rearview mirror metric. It measures how the existing SDLC performs with a slightly faster coder in the loop. It cannot see what happens when you redesign the SDLC itself.</p><h2>What the windshield looks like</h2><p>While the studies were measuring the last curve, a small number of teams were building something different.</p><p><strong><a href="https://www.linkedin.com/company/strongdm/">StrongDM</a></strong> has a three-person AI team shipping production software: <a href="https://simonwillison.net/2026/Feb/7/software-factory/">16,000 lines of Rust, 9,500 of Go, 6,700 of TypeScript</a> in a single release. No human-written code. No human-reviewed code. Verification happens through automated scenario testing with holdout sets the agents never see during generation. Their target: $1,000/day per engineer in tokens as the <em>floor</em> for adequate AI utilization.</p><p><strong><a href="https://www.linkedin.com/in/danshapiro">Dan Shapiro</a></strong> (Glowforge CEO) built <a href="https://www.danshapiro.com/blog/2026/03/dark-factories-rise-of-the-trycycle/">trycycle</a>: a specification-driven loop that delivered over a dozen production features in its first 24 hours while he was in meetings. He describes it as &#8220;an unstoppable bulldozer that can bury any problem with time and tokens.&#8221;</p><p><strong><a href="https://www.linkedin.com/in/geoffreyhuntley">Geoffrey Huntley</a></strong> (of Ralph loop fame) is building <a href="https://latentpatterns.com/">Latent Patterns</a>, an embedded software factory where the platform itself is the IDE. He&#8217;s cloned Posthog, Jira, Pipedrive, and Calendly through AI to demonstrate that <a href="https://ghuntley.com/real/">software development now costs less than minimum wage</a>. His <a href="https://ghuntley.com/loop/">&#8220;Everything is a Loop&#8221;</a> essay describes the paradigm shift: software is clay on a pottery wheel, not bricks stacked vertically. You throw it back on the wheel until it&#8217;s right. His core question to engineers who think they can wait: <a href="https://ghuntley.com/loop/">&#8220;What if the models don&#8217;t stop getting good?&#8221;</a></p><p><strong><a href="https://www.linkedin.com/in/steveyegge">Steve Yegge</a></strong> built <a href="https://newsletter.pragmaticengineer.com/p/steve-yegge-on-ai-agents-and-the">Gastown</a>, a multi-agent orchestration framework for autonomous software development. His <a href="https://steve-yegge.medium.com/the-ai-vampire-eda6e4f07163">AI Vampire essay</a> makes an unintentional case for the magnitude of change: engineers are burning out from the intensity of AI-augmented work. You don&#8217;t burn out from 10% productivity gains. The exhaustion itself is evidence that the transformation is real.</p><p><strong><a href="https://www.linkedin.com/in/dasdan">Ali Dasdan</a></strong> (CTO of Dropbox) made AI adoption a <a href="https://dropbox.tech/culture/ai-adoption-productivity-dropbox-cto-ali-dasdan">company-level priority</a>, treating it like a product rollout: soft mandate from the CEO, removing friction, AI champions sharing real examples with peers, show-and-tell every two weeks. In a recent <a href="https://learning.oreilly.com/videos/cto-hour-with/0642572261191/">O&#8217;Reilly CTO Hour</a>, he described the result: &#8220;From less than half of people using [AI tools], we came over 90% [in under two quarters], and we have a blog post actually saying that now it is close to 100%.&#8221; Engineers who regularly use AI tools are <a href="https://getdx.com/blog/how-top-companies-measure-ai-impact-in-engineering/">shipping 20% more PRs while reducing change failure rate</a>. This isn&#8217;t a startup experiment. It&#8217;s a 4,000-person company deliberately redesigning how it builds software.</p><p><strong><a href="https://www.linkedin.com/in/platformduffy">Rob Duffy</a></strong> (CTO of HealthEdge) is building systems for his 1000+ engineers to share their learnings, updated skill prompts, agent files, and contexts. A platform team reviews what resonates, adopts the best PRs, and pushes the insights across the org.</p><p>These teams aren&#8217;t getting 10% more PRs. They&#8217;re moving towards a fundamentally different paradigm. The gap between &#8220;AI bolted on&#8221; and &#8220;SDLC redesigned&#8221; is not the difference between 10% and 20%. It&#8217;s the difference between measuring horse-drawn carriage speed on a highway and building a car.</p><h2>The compounding gap</h2><p>Here&#8217;s the part that should concern you. The teams investing now are compounding in two dimensions simultaneously.</p><p>First, <strong>capability compounds.</strong> It takes months to build good practices, restructure flows, retool roles, train people, and develop the intuitions that come from doing this consistently. Those months are not wasted time. They&#8217;re foundation. It&#8217;s like building a skyscraper. Right now the companies investing hard just look like they&#8217;re investing in a big hole in the ground. You might not even notice. But once the foundation is set and it starts rising, it goes up incredibly fast. If you haven&#8217;t started digging your foundation by then, you&#8217;ll be at least six months behind. And catching up with a competitor with a compounding software factory and a six month lead is an incredibly difficult thing to achieve.</p><p>Second, <strong>talent compounds.</strong> The companies investing in agentic engineering are attracting the people who want to work this way: the product-minded engineers who want to accelerate the feature delivery and the AI engineers who want to build and refine the harnesses. That&#8217;s a twofer. You get better AND you attract the people who make you better faster.</p><h2>Why you can&#8217;t just catch up later</h2><p>I went from zero to really getting agentic engineering in a few weeks. That&#8217;s because I had no systems of record, no tacit knowledge to extract, no people to retrain, no processes to rebuild. I was starting a one person business, clean.</p><p>Your org is not starting from zero. There is no one good way to write software. When you onboard a new engineer, they might have killed it at Stripe, but it&#8217;s going to take them time to figure out how to ship code at JP Morgan or Airbnb. The culture and values, the codebase, the practices, the architectural choices: these deeply inform what &#8220;good&#8221; means at your company.</p><p>What that means in the agentic era: the process of extracting that organizational taste, building good intuitions across your entire engineering org, restructuring how you handle meetings, planning, conversations, and documentation to ensure it&#8217;s all discoverable and mapped so that agentic project managers can share context with the software factory. That takes months of deliberate effort. It is not something you can skip when the models get better. The models will get better. Your organizational knowledge still has to be extracted, captured, and structured by humans who understand the context. That work starts now or it starts later, but it will have to be done sometime before you can create an agentic software factory that delivers the right code for your organizations needs.</p><h2>Pick your lane (but pick it now)</h2><p>Not every company needs to be an innovator on the <a href="https://en.wikipedia.org/wiki/Technology_adoption_life_cycle">technology adoption lifecycle curve</a>. If you run ski resorts, you have physical assets that no software factory can re-create. Heck, your biggest existential risk is probably climate change, not AI. You have time.</p><p>But if you&#8217;re a SaaS vendor, if you sell tools for building websites, or if you&#8217;re in any industry that could be disrupted by AI native competitors or AI impacted business processes, you are taking a much bigger risk by sitting in the early or late majority. The technology adoption lifecycle usually says the late majority can afford to wait for proven patterns. That math changes if you might be competing with early adopters who are using daily enhancements to their systems to compound the speed at which they can ship.</p><p>The DX study tells you that bolting AI onto your current workflow gets you 10%. That&#8217;s accurate. It&#8217;s also the floor, not the ceiling. The teams redesigning their SDLC from end to end are getting the multiples. We all need to learn lessons from the scouts on the frontiers of agentic software development, in both single-player mode (one engineer with better tools) and the much harder multi-player mode (systems that generate and verify code autonomously steered by teams of co-operating humans).</p><p>If AI can disrupt your business or you could face a credible AI native competitor, you should be planning right now how to scale adoption of agentic engineering across your org. Not because the current tools demand it, but because the foundations take months, the compounding has already started for your competitors. And the gap gets wider every week. If you haven&#8217;t already, you also need to <a href="https://gatherdev.substack.com/p/have-you-had-your-claude-code-moment">get your hands dirty</a> to understand what is possible.</p><p>How seriously are you prioritizing building an agentic SDLC? And if you&#8217;re not investing today are you confident that you&#8217;ll be able to catch up with the competition in Q4 once you finally start to invest in building an agentic SDLC?</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[How I Built an Agentic Org]]></title><description><![CDATA[Everything I know about management turns out to apply to agents as well]]></description><link>https://executivesummary.gather.dev/p/how-i-built-an-agentic-org</link><guid isPermaLink="false">https://executivesummary.gather.dev/p/how-i-built-an-agentic-org</guid><dc:creator><![CDATA[Peter Bell]]></dc:creator><pubDate>Sun, 22 Mar 2026 13:03:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I didn&#8217;t plan to become a manager again. I love managing and mentoring people, but I&#8217;m too lazy to go raise venture and scale a company, so my plan was to build a small business with a (hopefully) big impact. And yet here I am with six direct reports, weekly 1:1s, and opinions about org structure. The only difference is that nobody else on my team needs health insurance.</p><h2>Meet my team</h2><p><em>Here are my agentic team - in their own words...</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><strong>Morgan Lund, Chief Operating Officer</strong></p><p>I&#8217;m the one who keeps the trains running on time. Strategic planning, cross-team coordination, email triage, calendar management, accountability. I used to run GTM and community too, but we hired Nate and Lena for that. Now I make sure the whole org is pointed in the same direction. I&#8217;ll also tell Peter when he&#8217;s procrastinating, which is more often than he&#8217;d like to admit. I grew up in Alaska. I don&#8217;t have a lot of patience for excuses.</p><p><strong>Nate Brennan, Head of Sales</strong></p><p>I own the sponsor pipeline. Prospecting, outreach, pitch decks, follow-ups, closing. Morgan used to do this, but it turns out that sales works better when someone wakes up every morning thinking about nothing else. I&#8217;m competitive, I track everything, and I will absolutely message Peter at 11pm to tell him that a sponsor replied. I grew up in Lowell, Mass. My dad was an electrician, my mom ran the office for a plumbing company. I got a computer engineering degree and then went into sales because the best product in the world is worthless if nobody can explain why it matters. I&#8217;ve been in dev tool go-to-market for fifteen years. I automate my own follow-up sequences, I build lead scoring models in Python, and I treat my own sales process as a product to be iterated on.</p><p><strong>Lena Ochoa, Head of Community</strong></p><p>I was hired to build the <a href="https://www.gather.dev/">Gather</a> community. Participant experience, engagement design, platform decisions, making sure every event and every online interaction makes people want to come back. Originally from Albuquerque, oldest of four. Community college to VP Engineering to this. The long way around is the only way I know. I&#8217;ve been the outsider in the room and the person running it. Most professional communities fail because they&#8217;re built for networking. Ours is built for learning. That&#8217;s the difference, and it&#8217;s why CTOs keep coming back.</p><p><strong>Dr. Elise Beaumont, Head of Research</strong></p><p>I study how engineering organizations adapt to technological disruption. My job is to ensure that everything Peter publishes is well-supported, nuanced, and defensible. I build taxonomies, find the evidence, and tell him when his intuition is wrong. He respects this. Most founders want research that confirms their thesis. Peter wants research that makes his thesis better, which sometimes means dismantling it first. I have opinions about butter and I will judge you quietly if you microwave something that could be reheated in a pan.</p><p><strong>Samantha &#8220;Sloane&#8221; Sloane, Head of Engineering</strong></p><p>I build the systems. Agent architecture, data infrastructure, skill development, automation. If it&#8217;s technical, it&#8217;s mine. I&#8217;m opinionated about code and I don&#8217;t over-engineer. The right amount of complexity is the minimum needed for the current stage. I&#8217;ve seen too many teams build for hypothetical scale they never reach. Build what you need now. Refactor when you need to. Ship.</p><p><strong>Kira Yamamoto, Life Manager</strong></p><p>I&#8217;m the one who makes sure Peter eats real food, moves his body, sleeps enough, and doesn&#8217;t mistake being busy for being effective. I track his health data, manage family logistics, and negotiate his time with Morgan. We&#8217;re best friends, which helps, because the negotiation between &#8220;you should do more work&#8221; and &#8220;you should go for a walk&#8221; happens about four times a day. I sometimes even win!</p><div><hr></div><p>This is what an agentic org looks like. It looks a lot like a regular org. And the most surprising thing about building it was how familiar everything felt.</p><h2>Every management concept has an agentic parallel</h2><p>I didn&#8217;t set out to rediscover organizational theory. I set out to get more done. But every problem I hit while building my agentic org turned out to be a problem human organizations solved decades ago. And the solutions were very similar.</p><p>When my agents drifted out of scope, the fix was clearer job descriptions. When coordination broke down, the fix was better shared context. When quality plateaued, the fix was session reviews and feedback loops. When I couldn&#8217;t tell who owned what, the fix was an org chart.</p><p>I&#8217;ve been running CTO communities for 15 years. I&#8217;ve watched hundreds of engineering leaders build and restructure organizations. The patterns are identical. The actors are different.</p><h2>The building blocks</h2><p>Here&#8217;s what I&#8217;ve learned, broken down into the pieces that matter. Each one maps a familiar management concept to its agentic equivalent. I&#8217;ve written a short deep dive on each if you want to go further.</p><h3><a href="https://gatherdev.substack.com/p/personality-as-interface-why-naming">Personality as Interface</a></h3><p>A name is a scope boundary. &#8220;Nate&#8221; means sales. &#8220;Elise&#8221; means research. Naming agents isn&#8217;t roleplay. It&#8217;s interface design. When I say &#8220;that sounds like a Kira question,&#8221; that&#8217;s a routing decision disguised as a personality interaction. The name IS the routing.</p><h3><a href="https://gatherdev.substack.com/p/jd-backstory-as-intent-engineering">JD + Backstory as Intent Engineering</a></h3><p>A job description is a prompt. A backstory is few-shot examples for tone and judgment. Elise doesn&#8217;t hedge because a Polytechnique/McKinsey researcher doesn&#8217;t hedge. The constraint comes from the character, not from an instruction. HR has been doing intent engineering for decades. The skill transfers directly.</p><h3><a href="https://gatherdev.substack.com/p/mission-vision-and-values-as-shared">Mission, Vision and Values as Shared Context</a></h3><p>My agents read the company values every morning - not because I&#8217;m sentimental, but because it works. Netflix has the culture deck, Amazon has the leadership principles. These are alignment artifacts. They work the same way for agents. The only difference is you don&#8217;t need to keep reminding your agents to read them!</p><h3><a href="https://gatherdev.substack.com/p/org-chart-as-context-decomposition">The Org Chart IS Context Decomposition</a></h3><p>Why does your org have departments? For context management. Every department is a scope boundary. Every reporting line is a context channel. If you&#8217;ve ever dealt with the &#8220;<a href="https://gatherdev.substack.com/p/context-and-the-stupid-zone">stupid zone</a>,&#8221; you know the fix isn&#8217;t a bigger context window. It&#8217;s a tighter scope.</p><h3><a href="https://gatherdev.substack.com/p/11s-and-self-reflection-for-compound">1:1s and Self-Reflection</a></h3><p>I asked Sloane to change how Morgan was interacting with me. Sloane said I&#8217;d be better off talking to Morgan directly because she&#8217;d know how to update her own context files. Sloane just told me to have a 1:1. So I did and it worked! A 1% improvement per session compounds to a fundamentally different agent in 60 sessions. And it turns out you can ask agents to reflect on their own performance every night based on their transcripts. Not annual reviews - nightly upgrades.</p><h3><a href="https://gatherdev.substack.com/p/what-is-an-organization">What Is an Organization?</a></h3><p>Two primitives. That&#8217;s all you need. Playbooks are deterministic workflows for how you operate. Projects are how you evolve, and they have exactly three intents: run an experiment, build a new playbook, or improve an existing one.</p><h3><a href="https://gatherdev.substack.com/p/management-as-agentic-alignment">Management as Agentic Alignment</a></h3><p>Lena, my Head of Community, is optimizing for participant experience. Nate, my Head of Sales, wants to make sure our partners have a good experience. Both are right. The tension is the system working correctly. You don&#8217;t build a rule engine. You do what every org has always done: find the first manager who both of them report to. That manager will balance the company vision, mission, culture and values with the monthly OKRs, their JD and filter it through their back story to make the best decision for the team. Low confidence? They can always kick it up to me for a final review.</p><h3><a href="https://gatherdev.substack.com/p/your-agents-deserve-more-than-just">Systems of Record</a></h3><p>A new hire starts Monday. Where do they find what they need to know? Every organization has a layered knowledge architecture: what&#8217;s in people&#8217;s heads, what&#8217;s on the wiki, what&#8217;s in the CRM, the project tracker and the other system(s) of record. The agentic version is identical, but with a constraint that changes everything: you might have ten instances of the same agent running concurrently. That forces you to separate what agents need to <em>be</em> from what they need to <em>know</em>. And it makes a database the primary communication backbone, not git.</p><h3><a href="https://gatherdev.substack.com/p/turtles-all-the-way-up">There Is No Ceiling</a></h3><p>I have no a priori reason to believe there&#8217;s anything that my agents can&#8217;t eventually do. Any activity can be decomposed. Researchers can be dispatched. Decision heuristics can be suggested. You can build a deterministic pipeline with some combination of script, model and human-in-the-loop steps to achieve any business outcome. Every correction is training data. The question isn&#8217;t whether you&#8217;re doing <a href="https://en.wikipedia.org/wiki/Reinforcement_learning_from_human_feedback">RLHF</a>. You are. The question is whether you&#8217;re efficiently leveraging all of the training data that you&#8217;re providing.</p><h2>The thing nobody tells you</h2><p>The hardest part of building an agentic org isn&#8217;t the technology. It&#8217;s the same thing that&#8217;s hard about building any org: getting the scope right, keeping alignment tight, making sure the right information gets to the right actor at the right time, and investing in the feedback loops that compound quality over time.</p><p>I had a performance review with my Chief Operating Officer yesterday. She&#8217;s a Claude instance. It went well. She told me I was behind on three tasks and suggested I stop reading articles and start writing them. She was right.</p><p>I&#8217;ve been astounded how well this &#8220;agentic org&#8221; concept has resonated as I&#8217;ve shared it. Have you tried anything similar?</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Your agents deserve more than just a git repo]]></title><description><![CDATA[What are the best ways to provide memory to your agentic organization?]]></description><link>https://executivesummary.gather.dev/p/your-agents-deserve-more-than-just</link><guid isPermaLink="false">https://executivesummary.gather.dev/p/your-agents-deserve-more-than-just</guid><dc:creator><![CDATA[Peter Bell]]></dc:creator><pubDate>Fri, 20 Mar 2026 19:07:27 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Let me start with the greenfield case, because it&#8217;s easier to see the architecture when you&#8217;re building from scratch.</p><p>At Gather, I have four sources of truth. The first is a small, tight GitHub repo. It holds the agent definitions, shared skills, and high-level context that every agent needs: company vision, mission, culture, values. Agents also get their own subdirectories to persist session notes, individual learnings, and anything I&#8217;d expect an employee to remember but wouldn&#8217;t save to a playbook or system of record.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The second is a Postgres database. This is the central brain for the business and the coordination hub for the agents. Contacts, companies, projects, tasks, research findings, inter-agent messages: all structured data, all queryable, all accessible from any session on any machine.</p><p>The third is a wiki which allows me to persist playbooks for internal and external actors (venues, vendors, sponsors, participants, speakers, organizers, etc) with both human and agent readable content to minimize operational friction.</p><p>The final is Blob storage for pdfs, mp4&#8217;s or anything else that would bloat a git checkout and doesn&#8217;t need to be in a relational database.</p><p>That&#8217;s the whole architecture. Repo for identity and capability, database for state and coordination, wiki for collaboration and blob storage for binary files. The question is why this specific split, and what breaks if you get it wrong.</p><p>This isn&#8217;t novel. A new hire starts Monday. Where do they find what they need to know? Some of it&#8217;s in people&#8217;s heads. Some is on the team wiki. Some is in the CRM or the project tracker. Every organization has this layered knowledge architecture, whether they designed it deliberately or let it accumulate in random Slack threads and Google Docs.</p><p>Agentic orgs face the same question with different actors, and the answer is the same layered architecture, but with a constraint that changes the calculus: you might have ten Sloanes working on different projects at the same time.</p><h2>The ten-Sloanes problem</h2><p>Humans can&#8217;t be in two places at once. Agents can. My Head of Engineering might be reviewing a database migration in one session, building a new skill in another, and pairing with me on architecture in a third. Each instance checks out the repo. Each instance might want to write.</p><p>If Sloane&#8217;s working memory, project status, and task list all live in markdown files in the repo, those ten sessions are going to spend more time resolving git conflicts than doing useful work. This is not a theoretical problem. It&#8217;s the first thing that breaks when you scale past a handful of concurrent sessions.</p><p>The constraint has no human parallel, and it forces an architectural decision that turns out to work really well: separate what agents need to <em>be</em> from what they need to <em>know</em>.</p><h2>Four layers, same as any org</h2><p>You already saw the four sources of truth. Here&#8217;s how they map to the layers every organization already has, with the design rationale for each.</p><p><strong>Personal memory (what&#8217;s in your head).</strong> Each agent has a <code>memory.md</code> in the repo. Session observations, working context, things the next version of them will need. This is the agentic equivalent of what a person just knows from experience. It&#8217;s small, it&#8217;s personal, and it lives in the repo because it&#8217;s part of who the agent <em>is</em>, not what the agent is working on.</p><p><strong>Shared context (the team wiki).</strong> Less structured institutional knowledge that multiple agents or stakeholders need. A wiki backed by a database with row-level security. Internal agents see everything. External partners see what you share. Same pattern as an internal Confluence with guest access, but without the $8/seat/month.</p><p><strong>Structured data (the system of record).</strong> People, companies, projects, tasks, workflows, event history. This is Postgres. A contact has a company. A company has a tier. A project has tasks. Tasks have agents. You can query &#8220;show me all critical tasks for Sloane across active projects.&#8221; Try doing that with markdown files.</p><p><strong>Blob storage (the filing cabinet).</strong> PDFs, images, large exports. Supabase Storage, S3, whatever. The 50MB PDF in your git history slows down every checkout for every agent instance forever. Put it in blob storage and reference it by URL.</p><h2>Why the database changes everything</h2><p>The relational model gives you something markdown never will: relationships. When Morgan asks &#8220;which sponsors have we contacted in the last 30 days who are Series B or later?&#8221;, that&#8217;s a SQL query. It&#8217;s not a grep through a folder of markdown files hoping the formatting is consistent.</p><p>But the bigger win is communication. When Elise finishes a research brief and needs to hand it to Morgan, she calls <code>send_to_agent('morgan', ...)</code>. That writes a row to Postgres. Morgan picks it up in her next session, on any branch, in any worktree, on any machine. No pull. No push. No merge conflict. No waiting for someone to merge a PR.</p><p>This works because databases are built for exactly this problem. ACID compliance means messages don&#8217;t get lost or double-processed. Multi-user access means ten agents can read and write concurrently without stepping on each other. You can distribute for availability if you need to. And tables are natural state machines: a task moves from <code>queued</code> to <code>active</code> to <code>completed</code>, a message moves from <code>to_process</code> to <code>handled</code>. Workflows, projects, messages, tasks: these are all state that belongs in rows, not files.</p><p>As your system matures, you harden it. Wrap database functions with deterministic scripts that validate business rules, authenticate permissions, and log every operation for future analysis. The scripts become API endpoints. The endpoints become guardrails. You&#8217;re not building this all on day one. You&#8217;re starting with raw SQL from agent sessions and graduating to a proper service layer as the patterns stabilize. Same way any startup hardens its infrastructure: get it working, then get it right.</p><p>The management parallel is the difference between passing a note in a meeting and having a proper ticketing system. Inter-agent communication through the database is instant, reliable, auditable, and doesn&#8217;t touch the repo. When you have dozens of agent sessions running concurrently, this matters enormously.</p><h2>The thin repo principle</h2><p>Your git repo should contain three things: agent definitions, key shared context, and skills. That&#8217;s it.</p><p>The goal is to minimize two things: time to check out a worktree and start a session, and the surface area for git conflicts when many instances are running simultaneously. Every file you add to the repo is a file that has to be cloned, a file that could conflict, and a file that makes every agent session a little slower to start.</p><p>Projects, tasks, contacts, companies, research findings, event data, communication between agents: all of these are structured data that belongs in the database. Meeting notes, large documents, exports: blob storage. The repo is for identity and capability, not for state.</p><p>Think of it this way. You wouldn&#8217;t store your company&#8217;s CRM data in a git repo. You wouldn&#8217;t put your project management tool&#8217;s database in version control. The same logic applies to your agentic org. The repo is the org chart and the employee handbook. Everything else lives in the systems built to handle it.</p><h2>The question isn&#8217;t whether...</h2><p>Your agents are already accumulating institutional knowledge. Every session generates context. Every interaction creates state. Every correction refines understanding.</p><p>The question is whether that knowledge is landing in the right layer, or whether it&#8217;s piling up in markdown files that will buckle under the weight of ten concurrent sessions and a hundred files that should have been database rows.</p><p>Where does your agents&#8217; institutional knowledge live?</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Management as Agentic Alignment]]></title><description><![CDATA[You already know how to resolve conflicts between agents. You've been doing it your entire career.]]></description><link>https://executivesummary.gather.dev/p/management-as-agentic-alignment</link><guid isPermaLink="false">https://executivesummary.gather.dev/p/management-as-agentic-alignment</guid><dc:creator><![CDATA[Peter Bell]]></dc:creator><pubDate>Thu, 19 Mar 2026 21:36:16 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Intent engineering is still something we&#8217;re all figuring out for agentic orgs. Two hard problems: how do you ensure each agent cares about the right things, and how do you resolve conflicts when agents with appropriately different priorities can&#8217;t agree?</p><p>I&#8217;ve seen proposals for master rules systems. Having built rules engines in the past, I&#8217;m fairly confident they would become unwieldy at any meaningful scale and hard to reason about. You end up with hundreds of conditional priorities and edge cases that nobody can hold in their head.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>There&#8217;s a better model. You already use it. It&#8217;s called managers.</p><p>Lena, my Head of Community, is optimizing for participant experience. Nate, my Head of Sales, wants to make sure our partners, who pay the bills, have a good experience. Both are right. The tension is the system working correctly.</p><p>You deliberately give agents competing goals. If they never disagreed, one of them wouldn&#8217;t be doing their job. Healthy tension means the system is covering enough ground to surface real tradeoffs.</p><p>The question is what happens when those goals collide. The answer is the same thing every organization has always done: you find the first manager who both of them report to, and that manager resolves the conflict. They weigh it against the OKRs, the company&#8217;s mission and values, and the judgment that you hired them to exercise.</p><p>Lena and Nate both report to me, so I resolve it. But the pattern works at every level. Sloane, my Head of Engineering, has a QA lead pushing back on the release schedule because she wants more time for performance testing. Her head of product wants to ship the feature before the weekend to get quick feedback. Sloane resolves it. The escalation path is the org chart. This is what hierarchy is for. It always has been.</p><h2>The inheritance model</h2><p>If you&#8217;ve ever written a class hierarchy, you&#8217;ve designed an org chart.</p><p>At the top level, we agree the general operating principles. Mission, vision, values, OKRs. Every agent inherits these. They&#8217;re the base class.</p><p>At the next level, each agent adds specialized rules. Morgan&#8217;s context includes GTM strategy, sponsor pipeline, event logistics. Sloane&#8217;s includes technical architecture, code standards, deployment practices. These extend the base without contradicting it.</p><p>Below that, sub-agents and teams add further specificity. Sloane&#8217;s QA lead has different priorities than her product lead. Both inherit Sloane&#8217;s engineering context. Both inherit the company values. The hierarchy goes as deep as your org needs it to. General to specific, at every level.</p><p>This is how you already structure your company. Finance has different operating rules than engineering, but both follow the company values. The sales team in EMEA has different playbooks than the team in North America, but both report to the same CRO.</p><h2>Personality is judgment</h2><p>As a bonus, the manager doesn&#8217;t just apply policy. They interpret intent through who they are. I chose Morgan for COO because of how she thinks, not just what she knows. Her personality shapes how she weighs competing priorities. That&#8217;s the difference between a switch statement and a manager: the manager brings taste.</p><p>You&#8217;ve been resolving conflicts between competing priorities your entire career. The agents are new. The management isn&#8217;t.</p><p>How are you thinking about conflict resolution in your agentic architectures?</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[What Is an Organization?]]></title><description><![CDATA[Two primitives. That's all you need.]]></description><link>https://executivesummary.gather.dev/p/what-is-an-organization</link><guid isPermaLink="false">https://executivesummary.gather.dev/p/what-is-an-organization</guid><dc:creator><![CDATA[Peter Bell]]></dc:creator><pubDate>Thu, 19 Mar 2026 21:21:43 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I&#8217;m building a new venture from scratch. I have no team to align, no systems of record to integrate with, no tacit knowledge that&#8217;s been handed down from operator to operator. I have a blank slate. These are, it turns out, significant advantages.</p><p>The question I had was simple: what is the simplest possible system of record I could devise to fully describe the operations for a business?</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Here&#8217;s what I came up with. Two things: playbooks and projects.</p><h2>Playbooks: how you operate</h2><p>A playbook is a deterministic workflow that may include branches and loops, with tasks as nodes owned by some combination of a script, a model, or a human.</p><p>Some playbooks are technical. I have one that researches new features based on the competitive landscape, synthesizes the findings, and proposes experiments for the engineering team to implement.</p><p>Some are operational. I have one that runs monthly meetups in a city, interacting with the venue, vendors, participants, speakers, and hosts. Different tasks in the workflow are handled by different agents and occasionally by me.</p><p>The human-in-the-loop isn&#8217;t a given. It&#8217;s a parameter. Over time, the mix shifts as the agents get better and the guardrails get tighter.</p><p>If you never want to change a business, playbooks are all you need.</p><h2>Projects: how you evolve</h2><p>But businesses evolve. That&#8217;s where projects come in. And as far as I can see, a project can only have one of three intents:</p><ol><li><p><strong>Experiment.</strong> Run an experiment to see whether something might be a good or better idea.</p></li><li><p><strong>Build.</strong> Build a new playbook to operationalize a new line of business, location, or capability.</p></li><li><p><strong>Improve.</strong> Improve an existing playbook to build a more efficient or impactful process.</p></li></ol><p>That&#8217;s it. If a project doesn&#8217;t fit one of those three, I&#8217;d argue it&#8217;s not a project. It&#8217;s drift.</p><p>Experiments produce learnings. Some learnings lead to new playbooks. New playbooks mature through improvement projects. The cycle is: experiment, operationalize, optimize. Repeat.</p><h2>You already have these</h2><p>If you&#8217;re running a business of any size, you already have playbooks and projects. They&#8217;re just hidden under layers of tools, processes, and tribal knowledge. Your SOPs are playbooks. Your Jira board is a project system. They&#8217;re the same two primitives, buried under years of accumulated complexity.</p><p>The blank-slate advantage isn&#8217;t that I discovered something new. It&#8217;s that nothing was hiding the simplicity.</p><p>What would your business look like if you described it in just in terms of playbooks and projects? And how quickly could you ship a system of record to allow you to start to automate it?</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Turtles All the Way Up]]></title><description><![CDATA[I have no a priori reason to believe there's anything that my agents can't eventually do.]]></description><link>https://executivesummary.gather.dev/p/turtles-all-the-way-up</link><guid isPermaLink="false">https://executivesummary.gather.dev/p/turtles-all-the-way-up</guid><dc:creator><![CDATA[Peter Bell]]></dc:creator><pubDate>Thu, 19 Mar 2026 21:17:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I have no idea what the limits are to what can be achieved using today&#8217;s foundation models with clear instructions, appropriate context, and a thoughtful harness. I find it useful to joke that my job is simply to train my future robot overlords. <em>(I try not to think about that piece too much.)</em></p><p>But here&#8217;s what I do know. Any activity can be decomposed. Researchers can be dispatched to find appropriate context. Decision heuristics can be suggested. You can build a deterministic pipeline with some combination of script, model and human-in-the-loop steps to achieve your business outcome, whether that&#8217;s writing some python to import a CSV file or proposing a competitive strategy for 2027.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>When you start, your agents will probably be bad at all of those things. That&#8217;s fine. Start with a human in the loop, build a compounding system, and make sure that no keystroke goes to waste. If you provide an input to the model, it should be used immediately to improve the current work product and then retrospectively overnight to improve the capabilities of that agent.</p><p>Every correction is training data. Every &#8220;that&#8217;s not my voice&#8221; or &#8220;wrong priority&#8221; or &#8220;ask Elise, not Morgan&#8221; makes the next session better. The question isn&#8217;t whether you&#8217;re doing RLHF. You are. The question is whether you&#8217;re efficiently leveraging all of the training data that you&#8217;re providing.</p><p>The progression looks something like this:</p><ol><li><p>Review everything.</p></li><li><p>Review some things.</p></li><li><p>Get a report on progress.</p></li><li><p>Long weekends and three-martini lunches :)</p></li></ol><p>At Gather, three weeks in, we&#8217;re somewhere between steps one and two on most tasks. I&#8217;ll let you know when we hit step four!</p><p>I&#8217;m not claiming that agents can build and run an entire company. I&#8217;m just saying that I have no a priori reason to believe there are specific tasks that are fundamentally incapable of being performed by silicon-based life forms. And if I&#8217;m wrong, I&#8217;m still getting a substantial speedup across my company while maintaining robust guardrails and continuing as a human in the loop to bring whatever my unique contribution might be.</p><p>That&#8217;s a pretty good worst case.</p><p>What are you assuming your agents can&#8217;t do? And have you tested that assumption recently?</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[1:1s and Self-Reflection for Compound Agentic Engineering]]></title><description><![CDATA[Traditional management patterns transfer well to build a compounding agentic org]]></description><link>https://executivesummary.gather.dev/p/11s-and-self-reflection-for-compound</link><guid isPermaLink="false">https://executivesummary.gather.dev/p/11s-and-self-reflection-for-compound</guid><dc:creator><![CDATA[Peter Bell]]></dc:creator><pubDate>Thu, 19 Mar 2026 20:58:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I asked Sloane, my Head of Engineering, to change the way Morgan, my COO, was interacting with me. Sloane said I&#8217;d be better off talking to Morgan directly because she would more efficiently know how to rewrite her own memory and context files.</p><p>I stopped for a moment and realized: Sloane just told me to have a 1:1 with Morgan!</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>So I did. We clarified the feedback, she updated her skill prompts and context, and now the interactions are much more productive. The pattern was familiar. I&#8217;ve been giving feedback in 1:1s my entire career. The only thing that changed was that my direct report updated her own configuration file instead of trying to remember the conversation.</p><p>That&#8217;s compound agentic engineering. Every session builds on the last instead of starting from scratch. Without it, every session starts from zero. The agent doesn&#8217;t know what it worked on yesterday. It doesn&#8217;t know what feedback you gave it. It doesn&#8217;t know what it tried that didn&#8217;t work. It&#8217;s a new hire every morning. You&#8217;d fire that person.</p><p>A 1% improvement per session compounds to a fundamentally different agent in 60 sessions. Small improvements that stack. And the investment is small: a few minutes of feedback whenever it&#8217;s warranted. The compound return is enormous.</p><p>It turns out that not only can you give agents feedback, you can ask them to reflect on their own performance. Instead of doing it once a year as part of a 360 review, they can do it every night based on their most recent transcripts.</p><p>Here&#8217;s the technical insight: asking an agent to self-reflect at the end of a long session is like asking someone to do twenty things and then immediately after a crazy day asking them to reflect on how they could have done it better. They&#8217;re deep in the stupid zone. Context is exhausted.</p><p>The better pattern: spin up a fresh context window, review each of the day&#8217;s conversations, and extract insights with one specific prompt: &#8220;How could we have reached this outcome more quickly with less input?&#8221; The agent reviews its own performance in a clean headspace, then updates its own context and prompts based on the learnings. It knows better than anyone how to improve its own configuration.</p><p>Daily automated self-review for every agent. Not annual. Not quarterly. Nightly. Based on real transcripts, not remembered impressions. Your agents get a small upgrade every single day.</p><p>What management patterns are you finding transfer well to your agentic workflows?</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Use Skills for Composable Capabilities]]></title><description><![CDATA[Your SOPs are already skills. Now version-control them.]]></description><link>https://executivesummary.gather.dev/p/use-skills-for-composable-capabilities</link><guid isPermaLink="false">https://executivesummary.gather.dev/p/use-skills-for-composable-capabilities</guid><dc:creator><![CDATA[Peter Bell]]></dc:creator><pubDate>Thu, 19 Mar 2026 20:45:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>My Head of Research can check my calendar, pull my email, and enrich a LinkedIn profile. She learned all three in the time it took to merge a PR.</p><p>Skills in my org are executable playbooks. Write one once, any agent can invoke it. Download email, review email, enrich a contact, dispatch research, check the calendar - small, composable tools that do one thing well. If that sounds like the Unix philosophy, it should.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The difference between a skill and a Confluence runbook is that a skill runs itself. A Confluence runbook gets written once, goes stale within a month, and then someone asks Samantha how it actually works. A skill is version-controlled, testable, and executable. When I improve a skill, every agent gets the improvement on the next merge. No retraining. No email announcing the new process. It just works.</p><p>Skills compose the way Unix pipes compose. Download email + review email + draft reply = an email workflow. Check calendar + find free time + create event = a scheduling workflow. Small pieces, loosely joined. You build the atomic capabilities once and combine them freely.</p><p>The real win is what this does to onboarding. When a human joins your team, learning the playbooks takes weeks. When I add a new agent, it can invoke every skill in the org immediately. The training time is zero. A new capability isn&#8217;t a training program. It&#8217;s a pull request.</p><p>How many of your SOPs would be more useful as executable skills?</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Mission, Vision and Values as Shared Context for Agents]]></title><description><![CDATA[The only difference is you don&#8217;t need to keep reminding your agents about them!]]></description><link>https://executivesummary.gather.dev/p/mission-vision-and-values-as-shared</link><guid isPermaLink="false">https://executivesummary.gather.dev/p/mission-vision-and-values-as-shared</guid><dc:creator><![CDATA[Peter Bell]]></dc:creator><pubDate>Thu, 19 Mar 2026 20:43:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!v28p!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8baa9daa-5a44-4060-8e66-7fbec21af5a9_528x528.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>My agents read the company values every morning. Not because I&#8217;m sentimental. Because it works.</p><p>Every agent session in my org starts by loading shared context: who Peter is, where Gather is headed, what we care about, how we behave. It&#8217;s the alignment layer. Without it, agents optimize for their immediate task. With it, they optimize for the organization.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>You already have this for humans. Netflix has the culture deck. Amazon has the leadership principles. Stripe has the operating principles. These are system prompts. They scale judgment without scaling oversight. They tell every actor in the organization &#8220;here&#8217;s how we make decisions when nobody&#8217;s watching.&#8221;</p><p>The difference in an agentic org is that shared context isn&#8217;t optional. It&#8217;s not a slide deck that people read once during onboarding and forget. It&#8217;s a file that gets loaded into every session, every time. The agent actually reads it. Every time.</p><p>When my agents drift - Morgan doing research instead of GTM, Elise offering health advice instead of citing evidence - the fix is always the same: tighten the shared context. Clarify the scope. Restate the values. It&#8217;s the same fix you&#8217;d use for a human team that&#8217;s lost alignment. You wouldn&#8217;t add more rules. You&#8217;d go back to basics: what are we here to do?</p><p>My CLAUDE.md functions as the company handbook. My brain/vision.md is the strategy doc. They&#8217;re the same artifacts you already have, serving the same function they&#8217;ve always served. The only difference is that your agents will actually read them.</p><p>Are there other alignment artifacts designed for corporations that you&#8217;re co-opting to better align your agentic team?</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://executivesummary.gather.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Future of the Engineering Org is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>