Why AI Coding Made Me More Productive - And Less Happy | Dillon Mulroy

Dillon Mulroy stopped writing most of his own code after Opus 4.5. He is more productive, still reads every line, and enjoys the work less.

Why AI Coding Made Me More Productive - And Less Happy | Dillon Mulroy

Episode 22 · September 10, 2026 · Dillon Mulroy

Why did AI coding make Dillon Mulroy more productive and less happy? He stopped writing most of his own code after Claude Opus 4.5. Six months later he ships more than he ever has, still reads every line, and enjoys the work less. The small implementation hits that used to create flow are gone. The day is now one hard design problem after another, and that is more exhausting.

Dillon is a principal engineer at Cloudflare. This episode answers what Cloudflare's May layoff of about 1,100 people looked like from inside, why lab-style agent loops are a poor default for most teams, and how he keeps a model on a short leash with Ghostty, Herdr, Pi, and Plannotator. Specs are TypeScript types, call stacks, and tests. Pull requests stay in the 300 to 800 line range. He uses Pi's slash tree instead of sub-agents.

Dillon is on X at @dillon_mulroy, GitHub at @dmmulroy, and Twitch at @dillon. https://dillonis.online

Key Takeaways

  • Dillon Mulroy has not written most of his own code since Claude Opus 4.5, from about mid-December 2025 or mid-January 2026. He still reads every line.
  • He is more productive and less happy because the small implementation hits that created flow are gone. The day is one hard design problem after another.
  • A single human still owns the output. Greater reach in less time means more review, not less.
  • Training juniors is the open problem: the struggle that built architecture intuition is easy to skip with a prompt.
  • Cloudflare's May layoff cut about 1,100 people, roughly 20% of the workforce. In his view it was mostly an AI-driven change in which roles exist. They opened different roles the next day.
  • Lab-style agent loops are impractical for a median developer at today's prices. He wants receipts before treating them as the default.
  • Anthropic Fable is a non-starter for a company like Cloudflare, he says, because it does not ship with zero data retention and can drop a session onto Opus 4.8.
  • His stack is Ghostty, Herdr, Pi, and Plannotator. Specs are TypeScript types, call stacks, and tests. Pull requests stay in the 300 to 800 line range. He uses slash tree instead of sub-agents.

Chapters

  1. 0:00 · I have not written my own code
  2. 1:18 · From bearish to all-in
  3. 6:55 · Tools do not define the work
  4. 10:02 · I enjoy this work less
  5. 14:16 · Macro problems, no flow state
  6. 17:38 · Feeling junior again
  7. 19:46 · Training the next generation
  8. 23:34 · Cloudflare cut 1,100 people
  9. 26:12 · Roles are compressing
  10. 34:55 · Loops are not for the median dev
  11. 41:50 · Why Fable is a non-starter
  12. 44:05 · Getting the model to write your code
  13. 46:20 · Herdr, Pi, and Plannotator
  14. 51:20 · Specs that look like code
  15. 52:44 · Slash tree instead of sub-agents
  16. 1:02:00 · A queue he can see coming

Mentioned In This Episode

Pull Quotes

I have not written my own code for the most part in the past six months. I am far more productive than I've ever been, for better or for worse.
Some days I hate that fact, because I enjoy my job and this work far less than I did a year ago.
I'm reading every single line of code that these things write.
We're just going from macro hard problem to macro hard problem. That's way more exhausting.
The biggest item of uncertainty that I have right now is how do we onboard and train the next generation of developers.
Fable is a non-starter for any real business because they don't ship it with zero data retention.
It's the design and composition of abstractions that agents suck at. That's where I spend most of my time.

Guest Bio

Dillon Mulroy, Principal engineer, Cloudflare

Dillon Mulroy is a principal engineer at Cloudflare, working on Artifacts and agent experience around the Workers platform. He previously worked on domains at Vercel and at Parcel, and he still comes from a Neovim-first setup. He co-hosts The Next Token with Sunil Pai and Rhys Sullivan. Find him on X at @dillon_mulroy, GitHub at @dmmulroy, and Twitch at @dillon.

Full Transcript

Read the full transcript

Cloudflare back in May, we had a layoff of 20% of our workforce, which was roughly 1,100 people. And AI was noted as the main driving reason behind that. You were like diehard NeoVim fan, if I remember right. I have not written my own code for the most part in the past six months. I am far more productive than I've ever been, for better or for worse.

And some days I hate that fact because I enjoy my job and this work far less than I did. A year ago, because we were just going from macro hard problem to macro hard problem to macro hard problem. That's way more exhausting. And there are statistics out there that like 80% of the value generation is usually done by like 10% of the people. Something along those lines.

That is the biggest item of uncertainty that I have right now is how do we onboard and train the next generation of developers. But you said something along the lines, I think it was like, This is the first time that I got AI to write code the way I would have written it, line for line. Tell me how you got them. Frustration, struggling, trial and error, nights of despair. So where I've ended up is...

It has been a wild journey since we last talked. And depending on... I'm sure whenever that was, it was probably pre-OBIS. It was pre-OBIS times. Four or five. And just as a little background for everyone listening. Last year at this time, I... Even last year, this time, and today's June 25th. Even up until November of 2025, I was not heavily using AI. And I was incredibly, incredibly skeptical.

And in fact, I've made public statements at that time being explicitly saying that I was bearish on the use of AI. You know, I rolled my eyes when Dario said by... I still do that, to be fair. I mean, oh yeah. The man does not do himself any favors when he talks about AI. When he talked about... AI would be writing 90 to 100% of code in 12 months. I was like, okay. Yeah, right. Roll my eyes.

And sure enough, here we are. As you alluded, I tweeted the other day that I have not written my own code for the most part in the past six months. Pretty much since mid-January, maybe mid-December. That transition's been interesting. And hard. And frustrating. And fun. And exciting. And all the feelings. I don't think there's...

I don't think there's enough discussion probably around the human element of all of this. That is a fair point, yeah. There's like a... At least for me, and a lot of the people I've talked to, I think... It's a lot of whiplash, right? Like, I've said this to my family. That like, six months ago, my job was completely different than it is today. At least my day-to-day. The outcomes that I'm producing are the same.

But the way I do my job is 100% different. And that was a hard transition and can be a hard transition. No. And so, like, how I got from NeoVim to not writing my code myself was really much like many other people, I think, in this industry. At least the bubble of tech Twitter is Opus 4.5 came out right after Thanksgiving. And over Christmas break and the holidays, I sat down and I played with it.

And it was distinctly different than the days of, like, Opus or Sonnet 3.5, 3.7. Like, up until that point, I was using clog code, like, on the side to write maybe 5 to 10% of my code. Yeah. And it's been a lot of work to get to the point where I feel proud or good about the work that I'm putting out with AI. I think everyone should go read a blog post by Ethan Neiser, who was on my team when I was at Purcell.

He put out a blog post called Don't Hold Back the Ocean or something akin to that. And he talks about that, this experience he went through over the same time of basically realizing that his job is going to change completely. And he, you know, for better or for worse, and I resonate with this article so much and it's so well written. Like, he has so much wisdom for his age.

He calls out that, like, programming and software engineering was a huge part of his personality and his identity and taking pride in being good at his craft and his work and excelling and going the extra steps among other engineers in the industry, right? Like, just being a true craftsman and peeling back the layers and understanding everything.

And he compared this kind of timeframe to the film industry back in, I don't know, sometime in, you know, the last century where it went from manually cutting film to digital cinematography. And that, you know, he imagined or extrapolated that people in that industry probably felt similar things.

And he ultimately comes to this conclusion that for better or for worse, the genie's out of the bottle and this isn't going away. Things are changing. Things are changing.

And he is choosing to lean in and, just as he's always done, become a craftsman of his tools, become an expert of his tools, still own the outcomes and the accountability and responsibility around the stuff he's putting out, and just get, you know, just get really good at the new era. And that's kind of the mentality I've also taken on.

So talk a little bit through this more, because I've taken the stance before, and you, to some extent, alluded on that, but then I think there was a certain path of divergence. So I always looked at, and I still do this as AI as a tool for my job, how I do my job and the tools I use also change throughout the 10, 15 years that I'm in this industry. Yep.

So in that sense, I get what you're saying from, like, the craftsmanship perspective. I very much do. But at the same time, I try to distance myself a little bit from the tools I use, which is a strong statement coming, working for a tooling company. But I don't think the tools I use define my work. So from that aspect, talk me a little bit through that, how that is different for you.

Because I get what you're saying with the way I work is different now than the way I used to work six months ago. Yeah. But at the same time, if I look at it abstractly enough, the things that where I generate value to my company in a project have never been the code that I put out.

So therefore, I, to some extent, I don't want to say struggle, because I can very much empathize with the way, like, looking at the craft perspective of like coding. But talk me a little bit through that, how that is for you. Yeah, so I think I agree with your framing and put the box or the categorization of LLM as a tool or that, like, that categorization around LLMs, because it's, it is still a tool. Yeah.

And I am very much, at least today, and I've tried to back off a little bit on making predictions around this technology, because my track record is pretty poor.

But I am not in the camp or one of those people that is like running cloud agents on loops like AFK and letting things get committed to the code base. Like I'm reading every single line of code that these things write. And I'm building software at a, like, kind of what you're saying, at a holistic level, the same way I always have. Like I'm instilling my principles, my experience, my past mistakes, my past successes.

Yep.

That's right. And I think it's maybe not a fair categorization to just call it simply a tool because it's, it is just unlike anything that, that is, that is fair. like it it's both and wielded by the right person it is a incredible productivity game like i am far more productive than i've ever been for better for worse and some days i hate that fact because like i mean some like genuinely like i think as of today

i would be comfortable saying that like i enjoy my job and this work far less than i did a year ago like it brings me i can see the personal satisfaction and less joy than it did i see a light at the end of the tunnel and as i've learned how to build with these things i'm starting to get into the flow of things but like you know sunil uh well we'll get into uh the podcast i'm starting with

sunil reese in a little bit but um i guess yeah i ultimately it's similar to like what ethan wrote about in his argo article and i've tweeted about this a bit like at the end of the day the human driving the llm is like the one that should be accountable and responsible and that's the same person that's always been accountable and responsible for their outcomes it's just like that that middle part of how you get from

the start of something to the end of something is very very different than what it is and i think to get good at that and to get the same quality that we were doing previously takes a lot of work i think and part of that based on what you're saying i'm kind of trying to draw some parallels here because i i think the part where you draw joy from all like satisfaction as you describe it i think

are very different for everyone yeah that's fair so and i think that is probably why there are very different perspective on this whole topic so a couple of years ago my wife went from being a developer on a team to an engineering manager and effectively didn't write any code anymore so whereas like as an engineer you still you have like little success moment or you had

little success moments every time you solve the problem or you um had like this like oh this is a very elegant solution to a problem whatever right as an engineer you had these ups and downs constantly as an engineering manager she didn't have that it was mostly just like making sure the team is productive and that so her she needed to adjust her value system of like what brings her joy in her

job very much she's i'm not sure if i hope she listens to this or doesn't but i do think for what it's worth she's a far better engineering manager than she's an engineer but like this jump was definitely there and i think in a similar way that is now the case too if you're like taking pride before in the code you wrote how you solve those problems then that is a very difficult and hard transition

yeah for sure i i think there's like a couple different angles i want to attack this from like the first is like i said i can see the light at the end of the tunnel i i had that on my mental to-do list yeah so i can i am starting to feel and i can see where i can and will hopefully start getting those kind of moments of like joy or uh just like excite not that i'm not excited about my work i love

building great products and getting out to people don't get me wrong but there was something about writing code and being so just yep enveloped in that that is different um for sure so the way i work in the way i develop systems and code i can start to see where those moments of fulfillment will come back and i do periodically get them on the other hand um i don't think i've gotten into like a flow state

in the past six months where i'm just like so supplicating into it and just in the zone and just uninterrupted um and um so we talk about this uh i think in our first episode on the podcast that sunil pi risoldan and i are putting out called the next token and like sunil i think kind of drove the conversation here um where

he talks about being tired and i agree with this a lot like i am far more tired on a day-to-day basis after working than i was before and i think part of that is that the hard things we're solving now and that require a lot of effort are really the only things we're solving and doing right now and what i mean by this is uh let's say you're starting a you know we need to ship a new feature right and this

might be just like something that's uh scales from you know like the ui down through a new endpoint down through new business logic down to figure out new indexing and storage patterns of the database right like kind of a full stack thing let's say you have to build that before we we'd sit down and like we if you've been working on the code base you kind of understand the the system and the architecture and

all that already and you might write out a tech spec or you might just jump in and start programming and stuff and uh just develop how you want to build this bigger feature and architecturally and the abstractions that you want and how data should flow uh and like the types and stuff and once you have like the big level picture of what this thing needs to be how it needs to be implemented

like that's the kind of the macro hard part that is the part that like will make or break a system

uh and then you kind of got to take a break from that and then you get to go implement it and when you're implementing it you're like solving these tiny hard problems like how should i write this class what api should i use what library should i use oh this data isn't flowing how i want it or there's this error path i didn't think about how should we handle this so like as you're implementing

this you're getting like all these tiny constant like tiny problems and hard problems to solve and like for me those were like tiny dopamine hits it's like oh i got and that's what got me into the flow state and that piece is largely missing right now so i think because we're just going from macro hard problem to macro hard problem to macro hard problem i think that's way more

exhausting and like i i i would want to continue to hammer this home that like the output from lms and ai i believe still ultimately needs to fall at least to a single human always and forever for accountability and responsibility and your the impact of your decisions and output it can be much further reaching in much shorter time periods now um so greater care has to be given to

like the macro big level hard problems um so i i think that's certainly part of like what makes this so exhausting i think there's other parts which we can dive into like it's like being new again like i i've kind of realized in the past couple weeks like some of these feelings and frustrations i'm having like remind me of when i was a very junior dev

and struggling with things like like i remember distinctly being back at one of my first jobs and like i was so frustrated building this service that i couldn't like figure out how to like get like the perfect abstraction of data flowing from like every act app down to a go lang in the orm and like i just couldn't understand how the other senior devs on the team were just so easily putting these things

together and like that like consumed me and like it was very frustrating not knowing how to get the results i wanted and like i'm kind of feeling that same thing now how so i talked on this podcast before i feel like there there's a it requires i i'm 100 with you that llm code needs to be reviewed for various person accountability responsibility being the highest on the list code quality being a close second

um the where did i wanted to go with this um i think it requires a lot of care and diligence to go with that because and that is also i think one of the important things creating that much signal to noise on online interaction around this topic because you if you're not diligent you're just completely going over the step not having these painful moments this experience making you feel like a junior

again because you're kind of like oh well but it works you can get something from nothing to working with llm in a matter of minutes um and that is what we see a lot of conversation around where it's like okay but working is for most business level production application not good enough for various reasons liability being a very big one um so that is a little bit the point where i'm like still

scared of how that is going to develop particular now for newcomers into the industry by no means and i've said this before i don't think the software development as a career is going away i think darius full of shit on that i i do think engineers are probably even more valuable right now than they were before because like this critical assessment this is good code this is bad code um

is very much needed howard junior developers that now though i'm super scared of because the way that i used to learn and you kind of alluded on that too was by going through the motions struggling with hard problems and right now you can throw in an easy advil that fixes that problem in the form of a prompt and cloud code will will work out something for you yeah this that

that that is the thing that is the biggest item of uncertainty that i have right now is how do we on board and train the next generation of developers uh for those exact those exact reasons i i i don't i don't have the answers like i have some ideas and like there's some more junior faults like cloudflare that i've seen uh you know some success with um but like holistically like i don't know

if it turns into something that's more akin to like uh apprenticeships or um

yeah i i don't know because the way i learned is like you said struggling and making mistakes and taking down prod so many times over the years like that's that's how you learn like that that's why i think one of the most important skills going forward is like system design and architecture skills and the only reason that i am as good at those things as i am is because i've iterated across

thousands of patterns and architectures for various different domains and apps over the years and i've seen what works and i've seen what fails and uh i've seen what is just hype and i don't know how you build that intuition without going through and maybe it's just faster like you're just putting out crappier software and watching it crash and burn and then having to rebuild it again i i which i don't

hope is because i think we do have a we're in an era of kind of worse software i i don't think that's i think many would agree uh so i hope it's not that route i so i said this on this podcast before i think this year is going to be like the value of despair or something of like shitty software and we we see all these layoffs where the ceos are like well for ai you very there's always a flavor of ai that

requires layoffs i i think this is utter nonsense the most that i was really pissed about was i think click up or something talking about 10x and 100x level engineer i you talked about jeff um seen the productivity increase and with dex horthy on this podcast we talked about like a two to three times if you know how to wield these tools is somewhat realistic yeah whereas like above that i think that

is just marketing and hype talking i i do probably have some disagreement with you on maybe a little on ai layoffs so cloudflare we did layoffs and oh that's true yes that's true i yes and uh i'm happy to talk about my perspective on that um i at cloudflare i don't think i would have necessarily expected this of a company the size of cloudflare um so for context to listeners cloudflare back in may

we had a layoff of 20 of our workforce which was roughly 1100 people um and ai was noted as the main driving reason behind that and um i i actually do think in cloudflare's case that that is to a degree true i i think there's probably other reasons there that probably weren't acknowledged and this is me taking off my cloudflare hat and if my employee's listening you know i'm i'm so sorry but like i think

there is probably a degree of and this is probably i think true of what won any company doing layoffs and especially right now i think there's certainly a degree of overhiring that happened in the past couple years that being said i do think the overwhelming truth of cloudflare's layoffs um is that it was ai driven um i mean i can only obviously speak about cloudflare but cloudflare did not

cut to just slim down we cut and the next day we had hundreds and hundreds of open roles we're hiring for they just happened to be very different roles than we were hiring for before that um we've done an incredibly good job at cloudflare and like to matthew prince's credit and dane next credit like they saw this trend not the trend of layoffs but like ai actually enabling people yeah yeah let's

ai enabling the workforce and they doubled down and went all in and they did it in what i consider to be a pretty diligent way an effective way at the scale of cloudflare it's just like that middle category of like do we really need scrum masters and uh three layers of directors and you know roles that are kind of just responsible from you know the busy work uh i i'm glad you touched on that because

we had uh a couple months back very shortly after we scheduled this podcast uh a conversation on twitter where i was kind of saying i don't see so i don't remember the right context i would need i was trying to find the tweet um but we were basically talking about engineering managers or team leads being able to manage more people with ai while also being productive in various capacity like shipping production code

um and we had a slight disagreement on that if i recall correctly so one thing that that kind of touches on though is that these strict separation of rules is kind of going to disappear from your perspective um and you just touched on that too so do you think ai will enable an engineer just as a concrete example to be qa and product manager in a single role certainly qa and it's my opinion that

and i've kind of always thought this and like i think both my time at versell and now at cloudflare has really shaped this but i think the best engineers are very product-minded engineers i think it's more important than ever important than ever for engineers to be product-minded and have deep understanding and empathy for their users and just live and breathe their pain and experiences and like there's certainly

some exceptions right like cloudflare we have people working on like embedded like linux kernels like okay there are some like very deeply technical discipline yes uh yes sub uh i don't want to say the word substrate that's uh uh lol as i'm leaking in there uh you know like sub-specializations of engineers and product people in any role hr legal that like you cannot like yeah you can't talk about like two percent

if even right so so i i do and like even in my own day-to-day work like i have like i go push a pr i have an agent that sees that i push to pr it can go out and it can go write its own playwright test against what i did record its screen doing it and just give this back to me and i can look at that and like that's like what in some larger enterprise corporations that have qa in us like that that's their whole job

like and i'm not this is not me saying oh we need to get rid of qa i don't really have a hot take that whether that's a role that should exist or not but like just as an example of like that would have been something that would have taken if i didn't have a qa analyst like i would have had to gone in qa and that's something that would have taken half a day of time that now just is done for me in 30

seconds um i i get where you're coming from and i'm a little bit like i i don't have my thoughts clearly stated because on the one hand i agree with you that the best engineers are um i i fundamentally think that the best engineers are very t-shaped with like can do engineer can do good engineering work but can do qa work but can do product work can talk to customer can empathize with customers

um i had several fun interactions with legal being able to talk to them and like have conversations around certain things definitely is makes me as an engineer more valuable instead of me if i'm just taking my well but i'm a coder i don't talk to legal right even a year ago i would have said that's the difference between a senior engineer and a staffer principal engineer

it's at a certain point there's like diminishing returns on technical expertise and the things that really move the needles is catalytic or soft skills and being able to work with people build influence build political capital and companies and just work across orgs and teams and likewise i think there's like there's like a forcing function that like is compressing that because of ai because

now there is more overlap in roles and responsibilities i would agree with that day-to-day basis so like those skills i think are more important than ever because you're going to be running into the situations where you need to be skillful and diligent and empathetic of other people other departments other teams uh how they operate than you ever had to be before like it's so much easier now to like

i have a problem that might be in another team's code base that's impacting me i can just send an agent out to go explore that figure it out and i i can work with them much more faster uh or much more fast than i used to be and i think that applies again i was just going to say i think that's happening across like all labor categories at least at cloudflare because like i i generally do think we have

put our money where our method mouth is in like enabling our entire workforce to operate with ai i i think with this way of looking at things i very much agree with you i'm not quite sold on the idea that i i don't have a good name for it but that in the future we'll just hire product engineers or something that i kind of expected to do all of these things at the same time i think i i don't

see that because because i if i look at it from a different perspective right like a um product managers at jetbrain some of them have very very technical knowledge so partially even sometimes deeper of the intelligent code base than i do uh but also some of them have no idea like know the product but have no idea from the code base would i be totally okay with them like vibe coding a

pr and be like hey this is a starting point this might be a good point of discussion let's iterate on that and bring that in the right structure for production absolutely i think that would be great um talking on something as tangible as source code is fantastic maybe compiling it and having a look at it do i think there will be the case where i could write a proper spec on a production like on a product

on a prd level that were um that the designers can take that qa can take that an engineer can take uh that marketing can take etc etc i don't know right so i think there's always like this level of specification specialization in that job i i think the lines between jobs are getting a little bit more blurry and there are like certain things that ai can facilitate to make that handoff or like that

transition easier more seamless also faster also faster yeah actually i i don't disagree with anything you just said and like i i don't want to i want to make it clear that i'm not saying that like any one kind of position should go away or anything like that because like i i don't think we're at the point where this i hardly know how to do my day-to-day job with this right right like i that being said

related to product managers and engineers i do actually have quite a bit of thought here that is fairly aligned with what you're saying in that product managers need to move away from this like uh workflow of like we're gonna take two weeks to write out this real long formal prd that's you know 4 000 words and chop it around and go bike shed it across all the other pms and organizations and teams

and the expectation is shifting as you said that like pms now need to take a step further in the direction toward engineering because they're unable to do so does that mean they're shipping you know pr straight to prod absolutely not but like we can suddenly they can put an mvp to their idea with a much more minimal prd spec and we can talk about that go back and forth and so you see a

cloudflare like our product managers are starting to contribute more are spending time producing mvps and prototypes to hand off to engineers that communicate the ideas in a better way and that can you know get taken and polish the rest of the direction and i think the last thing i'd say on this one is that like i definitely think there is different skill sets between even a product engineer and a

product manager and i think a very very good product manager is worth their weight and gold and platinum and every day that's an interesting perspective okay how

so you you touch on that very early um with like you don't use uh a million loops and agents in parallel i'm wondering because the people that are promoting this way of working like um boros from claude um peter seinberger from open ai i have nothing but sheer respect for the way they're the what they're doing in their day-to-day so i'm constantly wondering okay how much of this is just

the partially enabled by working on an ai lab and having like effectively infinite tokens um effectively infinite resources in that regards so how much of that is a look in the future of this is might be the way or the direction we're heading or more of like okay these are very different constraints these people are working with making this more of so um uh david kramer from century he put

report on this podcast like a science experiment how much where are you if this would be a scale of like okay this is the way everyone is going to work in a future arbitrary number or this is this works for them maybe very well uh which is also debatable but um this works for them and might work for a couple others but it's not broadly applying to the industry i think uh i think i would follow somewhere smack dab

in the middle it's very clear that like today's models and today's constraints like for the median dev even not even average the median dev like it's pretty impractical and it doesn't it doesn't work like i'm sorry like it doesn't work unless maybe you're at one of the ai labs and genuinely have infinite tokens to spend maybe then these can work and uh so that kind of like leads into my opinion it's like

i don't love the way the labs and these guys talk about this stuff and in particular uh the cloud codex team like

given the cost of the models today it is so far out of reach for the median developer using these toolings whether you're on a subscription plan or even at cloudflare we're on api billing like most enterprises are with these places like it is just prohibitively expensive now to use these features like dynamic workflows for every

for daily use we we talked about like if if model costs don't come down like does this turn into a thing where uh the principal engineers and product people in org are submitting rfcs and then debating and coming to consensus over the pieces of work that get to go into these types of loops and tooling right uh so like the entire organization is aligned on this is where we want

to spend tokens on solving a problem or figuring out a solution or looking for security vulnerabilities because like man it's just it's just impractical and i think to that end like i don't think the models are good enough i'm i want to see more receipts than we have from coming out of the labs that these workflows work i hear them talking about it and obviously we see products like codex and cloud code which

you know you their quality can be debated like i i think uh from one release to another like quality can drastically change it some days it feels like a polished beautiful app and system and the next day it's like dude this is like broken is all help and um like there is something about iterating fast but like the consistency needs to be there too to justify the cost and the the rate at which they're shipping

these features into the harnesses while at the same time and increasingly more people are getting probably priced out of these like if there feels like a misalignment there like um um and this can also this also leads into like a tangent or a conversation about like open weight models like gln 5.2 is getting these open weight models are getting pretty good and the day that they're

actually good enough then suddenly these capabilities do make more sense but are we're not you can see a world where that is the case and i think this is a lot of ai critics perspective is like it's it's always six months from now it's always in the future with ai um but and i think this is kind of what armin wrestle with in his article it's like i can imagine i can see the world i can tangibly sometimes get

results in my own experiments using these kind of workflows but it's not it's not mainstream yet will it be in six months maybe maybe gpt6 is incredible and they curve costs or token efficiency um but i think today june 25th of 2026 i think uh i think it's just a it's really annoying to be blunt hearing this kind of uh rhetoric coming from from the labs um to be honest yeah i i think if

quality and if quality goes dramatically up and prices need to come down otherwise we're going still even this year to have a very dramatic roi conversation within companies because like so far and this is slowly starting to happen already so far that for ai users there was pretty much this infinite amount of gold that was just coming from somewhere at companies to justify the

expenses to justify learning and getting exposure for this hopeful promise of like 10x 100x productivity um uber was the first that announced okay we're introducing very hard limits now after they spent through the entire air budget for the year within four months so and i think a lot of companies will do the same very soon if there is not an anti-cycle of like quality up and prices down table was the first

where we saw where i at least noticed a dramatic increase in quality at the same time it was also twice as expensive and slow as fuck yeah so and like i have a rant on fable but two problems with fable real quick uh just to throw some hot spicy takes in one yeah i'll clip that and put that on your youtube on your twitter account do it so fable

is a non-starter for any real business because they don't ship it with zero data retention so that means they will retain all your prompts all your code you send in and train on it that is not true of any other model with these companies everything else has zero data retention so immediately day one fable non-starter for a company like cloudflare yep uh the other problem i have with fable is

the fact that it takes it upon itself to downgrade to opus 48 yes two problems i have with that first problem is i do not take thinking away from me that's something i feel very important about and i know we haven't gotten too deep into like my workflow but like i hand handhold and i steer these agents and i i know what i want from them and what i expect of them and i don't want more

dynamism out of them in that regard the other problem with dropping to opus 48 is cost let's say you're running a session for your you build up 200 000 tokens in your context window and suddenly fable decides to drop to 48 the whole time you're working with fable you're likely getting cashets and thus yes you drop down to 48 suddenly you're gonna have to repay that entire cost on inference

it's just that's crazy to me like i'm sorry okay that's my rando table so my rand on fable is that it's a self-made problem that it's not available right now oh like 100 the whole rhetoric around it was just like oh it's so dangerous oh no we cannot do it and then they're like ah oh no they blocked it no i was like yeah dude you fucked around and found out like what are you expecting exactly

yep aligned so you had this tweet a while ago where you said like this is the first time that i got ai to write code the way i would have written it line for line yep yep tell me how you got there

frustration frustration struggling trial and error nights of despair and existential uh dread um yeah i mean it took a long time to get there and it's like i said at the beginning of the show it's i still do not love my day-to-day as much as i did before no that being said um if you're gonna choose to kind of go all in and master if you think agentic engineering is probably the future and you're

someone that cares about your craft and your output and being very very good at what you do and how you do it um spending time building with these and iterating and figuring out what works i think is important and to do that i think you almost need to completely ignore everything on twitter uh because it's just like a lot of the times like you're a lot of the people i see talking about ai

workflows are young yc founders and i have no doubt that they are talented and they're producing good products but they just it's like i'm sorry but yes you just don't have the experience to know what makes good production systems that run for years and years and years and years and they're maintainable right so immediately everything you read with a grain of salt also people trying to sell you something

and um yeah it was just a lot of trial and error and frustration i mean just like just like i got like i talked earlier how i learned to be a good software engineer by failing and creating slop that produced outages over the years it was kind of the same thing compressed into this time frame um but where i've ended up sorry go ahead no no that's literally where i wanted to transition to so where

i've ended up is uh let me outline my tooling to start so i kind of have three or four major tools that i use one i'm working in ghosty with uh a terminal multiplexer called herder which is um on my checkout list looks very interesting so i've started using that in the past two or three weeks and it's basically just a modern version of tmux built on top of lib ghosty so things like kitty image protocol

work uh it has like um you know like agent tracking and indicators and stuff and uh the cli is just much more even just human friendly but also agent friendly for like controlling like uh like dev servers and monitoring stuff like that so herder is my terminal motiplexer how i organize projects and workspaces and then i use pi uh and maybe this is like the the neo vim coming out of me uh i like i like the pi's

minimal i like that it's very customizable um that being said i've i don't customize it heavily and i think where i've landed on like harnesses is that i just want my harness doing as little as possible and i want it changing as infrequently as possible uh one frustration i had with like open code and clog code was that like new features would ship week to week tool definitions would change the system prompt

would change and that deviates behavior with your model and so it's harder to get a baseline uh i was saying the exact same thing here on the podcast when we talked about pi so totally get that yeah so like having this tiny system prompt only a handful of tools and just actually minimizing what gets injected into your context uh i've i've learned to value tremendously um and then um the one other

tool that's been incredibly important to me is a tool called plan notator uh you find that plan tentator.ai free tool and basically it works with all the harnesses pi open code codex cloud code um it has two big features i'd say the first is local code review and this is something i was tweeting about at the beginning of the year super heavily it was very clear to me up front in the beginning of

the year that our code review processes were broken like across the board from many different angles but one of them was i didn't want to be pushing ai changes to prs even publicly not publicly but like even for my team to see because i haven't had a chance to fully holistically review it and we at the time didn't have good tooling to really review those changes locally planetator amazing uh uh tool where you

can just do like slash planetator review it launches a web app that looks like your github pr review you can go file by file it's built on top of the pierre disk tooling so it's very fast very at great ux and ui very fast and you can go through add comments just like you would a normal pr review and when you're done you click submit and it injects it back into your current session with your agent

so like that feedback loop of working and reviewing code becomes much faster um and another like important thing i would say up front is that the way i deliver in design work hasn't changed at all i'm still focusing on tiny changes like incrementally stacked prs and i try very very hard to limit a individual atomic piece of work or a pr to be around 300 to 800 lines of code that way it's manageable like it's

scoped it's atomic i can ship it i can review it without my eyes glazing over right i can still fully understand everything in my code base if i can review it in those small chunks yep so what do you mean with stack prs just to make sure we're on the same page uh so you do like a tiny pr and then you would branch off of that pr or that branch on a new branch and then uh do another you know 300 800

change and rather than let's assume that bottom one hasn't merged into main yet you would open a pr merging this one into the first branch and it kind of creates this stack and there's a bunch of tooling um that makes this really really good things like graphite that got acquired by cursor it's really good uh jj kind of has so it doesn't necessarily do it out of the box um git lab and github are both

shipping stack pr support um by default soon uh it's just a way of i don't i've been using stack prs for years now um uh so then the other way the other part of planetator that i use is uh kind of two features that do the same thing the first one is slash planetator annotate which you can give it a file path and it will pull up that file and you can basically highlight things and add comments to it

and submit it back to the harness the same way so that could be like you know i could tell the agent go write a spec for adding a new or orm abstraction and then i you know write it to a markdown file then i can annotate that and feed my feedback back to the agent all in the same loop very like tight the other one is uh slash planetator last which takes the uh agent's last message and loads that into the

the annotation tool and then i can go edit that so more often than not i'm like building up this shared understanding of like a a plan for an implementation going back and forth in messages with the agent rather than like a markdown file and i'm using slash planetator last and over time this builds up to look more and more like a spec and we finally have when i finally trust that we have a shared

understanding of what needs to be intimate implemented clearly that's when i'll have it right to a markdown file and implement um which leads to like the next part like how do i actually do this so pi has a feature called slash tree and slash tree i would also say that pi notably does not have sub-legents uh so what slash tree lets you do is you can start having a conversation with an agent let's say i'm uh

i need to build like a new endpoint into like artifacts which is a product i work on i will often start my work by asking the agent questions i already know the answer to and this is generally around uh particular pieces of my tech stack uh or libraries or frameworks or existing um abstractions in the code base so i'll be like you know what is hono how does hono work how do you

configure it what are the apis blah blah blah and then it will go off and do research and like i know all the answers to this right it will go off do all the research and i might go back and forth a couple times until i i feel good about its understanding of hono and then i will use slash tree to jump back to the very root of the conversation and that you can choose when you do slash tree to like have it

summarize almost like compaction back to there and oftentimes i'll either just like let that message that tree sit or i'll like write that last bit to like a markdown file and i'll jump back to the beginning of the conversation where there's no context and they'll be like how do you use durable objects for git store or for uh sqlite storage what are the apis so i basically build up its

understanding of the technologies and the patterns i want it using and then i have these tiny concise summarizations of the things it worked itself to and i bring those back to the main line and then i start iterating with planetator to build up like a tech spec which um is another interesting thing i think i do maybe a little bit differently than most people my tech specs don't look like prds my tech

specs look like code like it is largely typescript types uh typescript types uh interfaces where are the bounds and seam uh boundaries and seams what adapters and implementations of those interfaces do we need and another thing i started doing is making the agent outline the call stack that it's going to implement is it editing an existing call stack is it creating a new code path entirely outline the

the call stack and show me the input and output types and the errors that can occur in each step and then the final bit of my tech spec is i make it write uh the tests that it's going to write using rdr or red green refactor tdd or whatever i use that poke oxy for this so then my tech spec is like i don't know 650 lines maybe a thousand but it's like here's all the types

here's the code path that you have to go implement here's all the interfaces like i trust agents to implement if i give it a function i trust agents to implement that pretty correctly most of the time it's the design and composition of abstractions that agents suck ass at like they are bad and that's where i spend most of my had been spending most of my time fighting it and in this way you come to

alignment on the the shape of what you're building with actual code and you can leave to it the implementation which it's already good at is it perfect no i still have to it still does the stupid things exhausting things annoying things that llms do um but i'm solely working away chipping down on that stuff too i i might steal some of these things because a lot of the things you describe

tackled the issues that i have with spec-driven development fundamentally yep um the a i'm too explorative usually that like me sitting down and writing like a whole ass prd that is like flawless for an llm to implement this i think unrealistic i think it's unrealistic for for anyone see people seem to have success with it so i want to leave room there i i have receipts

exactly i want to see the receipts uh the idea of writing the types uh i like i have done something similar with a red green approach where i like to like either implement the test myself or like use an llm to have like very solid tests and then have the llm implement based on that or already gets that 10 times better the results and be like oh write tests for this existing function they're all

trash all useless yep um i i like the plan notator thing that is really cool i didn't know about that tool i definitely need to check that out um so talk a little bit through that treat concept a little more because there you lost me for a minute so how is it how do you get back from so i assume you have like kind of like a couple tree or like branches and leaves at some point then where it's like okay

this is how hona works this is how sqlite works this is how i don't know we do testing in the project right like isolated uh very isolated context and shared understanding in that space how do you then drag that back is it through compaction that these things then get in form of a summary back into the root context or how is that way back yeah so um when you use slash tree to jump somewhere else in the

conversation you have three options one is like do nothing and just go back to an earlier point in the conversation and it's important to note that this does not roll back like git history it's only context and conversation yeah um so you have the come you can either do nothing you can have it auto summarize or you can give it a prompt to summarize itself and

i have a tendency to like do nothing and like i get that like last message in a branch to a point where like these are the things you need to know about and i'll either use slash copy and take that back to another point in the conversation or i'll have it write that message to a markdown file so it's captured and i can reference it from another point in the conversation

and like my i will have sessions in pi that have hundreds of messages back and forth like i might start building this new you know from the front end down to the database in one single session because like i'll start playing out the feature like we just talked through get it implemented and then i'll start code reviewing with planetator we'll go back and forth a couple times and then i might jump back

to you know message zero with no context all right let's now write end-to-end tests and create a new harness uh test harness around this and that will be another couple hundred messages that have their own branches down and like you could label spots and you can tell the agent go look at this spot with this label and it knows how to refer to um yeah there's

i don't think i could work without tree again it's just like it's almost like you are your own sub-agent like you're doing the work of sub-agents yourself and i i actually had a tweet before i really got bought in on tree whereas like i think i don't love sub-agents because like they have a tendency to like they're very good at like exploring and they're convenient but like it's all at the discretion

and judgment of the core agent and a lot of the times the summaries that i would see these sub-agents bring back or context that it would bring back aren't the things i actually care about getting into the main context window and it can also like inhibit your judgment and your thinking um so yeah i i don't use sub-agents and i kind of do it myself because i want i i've built up a good

enough intuition of how these agents and models respond to context and the impact context has on their performance that like i know exactly you know i don't want to say exactly right it's still very hand-wavy non-deterministic but i have a relative instinct now of like what i need in the agent's context window to get better results and like i at this point i need to be the one driving

that because agents aren't good enough i i think sub-agents the only place where i come to appreciate them if you know okay i have this massive chunk of work and this is going to explode the context window in this way to kind of steer compaction a little bit more graceful that you give it chunks of work that are like more holistically together i don't i usually don't work in like these big

chunks anyway so i come to the same conclusion i'm not a big fan of uh sub-agents i think this original narrative like oh yeah i have a designer sub-agent i have an architect i have a review sub-agent utter trash utter trash it is interesting though and like uh to just recall back to the loops conversation and armin's piece on like how you could see this in the future like i'm kind of at this point too

because like i can see and i have like this like this workflow is exhausting like it is a lot of work i can work on like two or three projects or maybe two or three features at one time or debugging like a century issue at one time and i am certainly more productive and i can get very close if not the same quality that i would do handwriting but it is exhausting and i feel tired and it's annoying

but i can see a world where i have like a queue where i say like here's the feature that i want and like some agent picks that up off the queue and in like one pane it starts the session of us talking back and forth of exploring the technologies and building up that context and then when we're done with that it takes its first pass at the tech spec and then we iterate on that like you know

like it puts that work into the queue the next one's like okay i know what dylan likes in a tech spec i know he likes input pipes like output types interfaces call stacks like i see a world where i can have these kind of like queues driving some of this work that takes some of the manual parts that i'm currently doing out of it and like even like with review like i've been working a lot the past few works to like

capture the things i find valuable in software design and architecture and review and trying to put those into skills and probably maybe like a flu agent which flu framework is like something the asher team put out um to like take that first pass at code review uh applying my taste and principles even if like and i do not anticipate this being anywhere near perfect or even good up front but

like if it can take five minutes off of my review and like i don't have to correct it on things like is record or passing unknown the hallway down and constantly revalidating stuff like if it can take those things away that are just so annoying to correct every time that's a that's a that's a win for me right um so that's like i'm trending in that direction but like it's still a lot of work to do that

thank you so much yeah man this was awesome i love this conversation love to hear that thank you