The Orchestrator's High
AI agents now work on their own for minutes at a time on most tasks and for a full working day on the big ones, and I run several at once across accounts and tools. The output is real. So is the racing brain, the pile of half-finished threads, and a bar that keeps rising. Honest notes from inside the loop, updated after four more months of it.
Fabian Mösli Reading Preferences
Key Takeaways
- • Most agent sessions still run five to twenty minutes, and the big jobs now run for hours. Either way your job shifts from doing the work to running several parallel workers, and the bottleneck moves from your hands to your attention.
- • Starting agent work is nearly free. Finishing it isn't. Too many open threads and you forget what your own sessions were for, and the time you saved gets eaten by higher expectations, including your own.
- • Decide before you launch, cap what's open, tell agents how to report back, and close loops with real people. Exercise, time outside, and the people who matter still do more for your head than any tool.
In this guide
About a year ago, AI took a real jump in usefulness for deep knowledge work. Claude Code had been around for a while, but it was around then that people inside the bubble started realising it was good for much more than writing code: research, planning, decision support, knowledge work of every kind. I got in, started learning, and within a few weeks I’d figured out how to get it to actually carry context from one session to the next. Self-learning systems, context engineering: the kind of thing that turns a stateless assistant into something that remembers what it learned yesterday.
The first weekend I really pushed the new setup, I was logged into Claude Code with two accounts at the same time. Private projects in the browser, work projects in the desktop app, both accounts grinding against their session limits, my OpenClaw agent humming away on a third window. Five or six sessions in flight, another fifteen queued up for follow-up, spread across four virtual desktops on a wide screen so I could flip between them with the keyboard. I’d missed dinner. I hadn’t been outside since lunch. It was past midnight and my brain was loud.
Every day a new wave was coming in. A new model release, a new agent harness, new MCPs, ten new “best practices” people were posting about, twenty new use cases I hadn’t tried. I was getting hooked, and I knew I was getting hooked, and I was getting hooked anyway.
I first published this guide in May. Four months later, most of it holds, and some of it got more intense. Agents now run for a full working day on the big jobs. Running several accounts and several different agent tools at once has become normal in the circles I move in. And two problems I barely mentioned the first time have become the ones I think about most: the pile of things I’ve started and not finished, and a bar that keeps moving up.
This guide is about all of it. About the shift from using AI to running a small team of yourself. The superpower side is real, and I’ll get to it. So is the cost. And so is the bubble of other people sprinting alongside you, making sure you don’t slow down.
What “orchestrating” actually looks like
The mental shift is the part most guides skip. With chatbot-style AI you ask a question and read an answer. With agentic AI you launch a session, the agent works on its own for a while, and then you read the result and decide what’s next.
For most of my work, “a while” is still five to twenty minutes. That’s the range that shapes the day. Long enough that you can’t just sit and watch, because you’d be wasting the whole point. Short enough that you can fit two or three rounds of feedback to other sessions into the same wait. So you start running things in parallel. You launch session B while session A is working. You give session C its next instruction while you’re reviewing session A’s result. You’re not doing the work anymore. You’re conducting.
My setup is unspectacular. Four virtual desktops on a wide monitor, one window left and one right per desktop, so I can flip between eight things by keyboard. A Stream Deck on the desk (twenty-four programmable buttons) is mostly mapped to switching desktops and pulling specific apps to the foreground, plus a push-to-talk button I use to dictate instead of typing when my hands or my patience have run out.
The long runs
What’s new since the first version of this guide is the other end of the range. The biggest job I’ve handed off so far ran for more than eight hours: migrating an entire company website to a different tech stack and, at the same time, onto a completely new design system.
The eight hours are the headline, and they’re the least interesting part. Before I launched it, I spent two to three hours preparing: what exactly moves, what the new design system looks like, what must not change, what “done” means. Afterwards I spent another two to three hours reviewing. So roughly five hours of my time for a job that, done the old way, is weeks of a small team’s work. That’s a genuinely absurd trade. But notice where my five hours went. All of it sits at the two ends, the brief and the review. The better the brief, the shorter the review. The long run doesn’t remove you from the work. It concentrates you at the edges of it.
And even an eight-hour run isn’t quiet. The agent keeps streaming what it’s doing, file after file, decision after decision, and that ticker is weirdly hard to look away from. You don’t need to watch it. You watch it anyway.
I was curious whether my impression matches more objective measurements, so I looked it up. METR (a research nonprofit that tests how long a task AI models can handle on their own) puts the best models at tasks that would take a skilled human twelve hours or more, if you accept that they succeed only half the time. Ask for 80% success and it drops to one to three hours. That felt about right to me. Long jobs are possible. Most jobs worth running are still shorter, and the longer the job, the more your brief and your review decide whether it’s any good.
Rookie numbers
My own numbers are modest. Steve Yegge (a veteran engineer and blogger whose posts I’ve been following for a while) runs twenty to thirty agents in parallel through a custom setup he calls Gas Town. An engineer at Cloudflare rebuilt most of a popular web framework in about a week, across more than 800 agent sessions. And an Anthropic researcher had sixteen agents work for about two weeks on a C compiler, for roughly $20,000 in usage costs. That last number says a lot: the people running agents for days have deep pockets or a business case that pays for it. Most of us are rationing a subscription.
But the structure is the same at every scale: you’ve stopped being the person doing the work. You’ve become the person managing several copies of someone else doing the work. The bottleneck moves from your hands to your attention.
What this actually unlocks
The case for celebrating it is genuine, and I don’t want to wave it away.
A few months ago I redesigned the user onboarding flow for our mobile app at Carewell. Not a tweak but a re-think. What the new user sees, what we ask them, what we don’t ask them, how we sequence the trust we need to build with someone before they hand over information that matters.
Old timeline: weeks. Schedule a workshop with the product team. Brief a designer. Wait for a draft. Talk to engineering about feasibility. Get compliance to weigh in. Wait. Wait. Every round of feedback gated by other humans being free at the same time as me.
New timeline: a weekend. Roughly ten interview rounds with the AI playing the role of different expert advisors, where I made something like thirty product decisions out loud and had each one stress-tested before I committed to it. By Sunday night I had a working high-fidelity prototype: not clickable mockups, but real frontend code with realistic test data, runnable on a phone, walking through all the actual paths a new user could take.
The part that sped up most wasn’t the building. It was the decision-making. When I work with human experts I’m typically waiting two days between question and answer, only during their office hours, and I have to bring each of them up to speed first. With AI playing the same role, the feedback cycle is two minutes, it’s three in the morning if I want, and it has the context already.
And it isn’t one advisor on call. It’s all of them. In a single session I can have a senior UX designer reviewing a flow, a behavioural economist pointing out which cognitive bias is going to bite us, a first-principles thinker asking why we’re solving this problem at all, and an engineering lead telling me which version is realistic to ship by when. Then the same system turns into a developer who actually builds what the conversation has just decided. Immediate advice and feedback on every topic where I lack expertise, and a builder standing by to implement whatever the room lands on.
Multiply that by every decision a product or a strategy needs and you start to see what the noise online is about. The point isn’t that AI writes better code than a senior engineer. It’s that an experienced generalist with the right system around them can compress a cross-functional team’s whole feedback loop into a single afternoon. And with the long runs, a team’s whole month of execution into a night.
What it costs you
The other side is just as real.
The first thing I noticed was exhaustion. Not foggy or distracted, but properly tired. After a few hours of context-switching every couple of minutes, the part of my brain that does the hard work just clocked out. Sitting with a question, holding three constraints in my head, working through a non-obvious chain of reasoning: none of it would happen. I’d end up running easier tasks for an hour to let my head recover before it would handle anything heavy again.
There’s a subtler version of the same problem that snuck up on me later. I’ll catch myself thinking I could do this better with AI in moments where the right move is to just sit and think for ten minutes. The instinct to reach for the tool has overgrown the muscle to use my own brain. That muscle, like any muscle, weakens when you stop using it. Once you’ve felt that, you start guarding the time when you’re supposed to be the one doing the work.
When I looked into why this hits so hard, I found the work of Gloria Mark (a professor at UC Irvine who has spent two decades measuring attention at work). By her numbers it takes about 23 minutes to fully recover from an interruption, and the average time people stay focused on one screen has dropped from two and a half minutes in 2004 to 47 seconds. The orchestrator workflow is an interruption every few minutes by design. You don’t just risk losing focus; the setup makes deep focus mechanically impossible.
The second thing was the racing brain at night. I’d close the laptop at midnight and lie down and immediately my head would start the next session. What if I tried this. What if that other thing actually worked. I should have launched one more before bed. Quentin Rousseau (CTO of a US software startup called Rootly) wrote publicly that a doctor prescribed him a sleep medication that blocks orexin receptors, because his wakefulness was, in his words, “fired up by hours of agentic dopamine loops.” I’m nowhere near that. But the mechanism he describes is the one I felt.
The third thing was the absence of the rest of life. Meals delayed. Walks skipped. Bathroom breaks postponed an hour because three sessions were about to come back. A weekend where I worked from waking until midnight every day and barely registered that the weather was good.
And a fourth one I underestimated: reading what comes back. A badly shaped summary is exhausting. A wall of text, the one decision it needs from you buried in paragraph six at the same volume as everything else. Multiply that by fifteen sessions a day and the reading becomes the work. The format of what comes back matters far more than I expected.
The work was real. The output was real. But by the end of those stretches I was not a person I’d want to be in a room with.
Starting is free. Finishing isn’t.
This is the part that got worse with time, and it’s the one I didn’t see coming.
When you can spin up a team tailored to a task in seconds, and that team works for hours, starting things costs almost nothing. So you start a lot of things: a research thread, a prototype, a side project you’d shelved for two years. Every one of them feels like progress the moment it launches. Finishing still costs what it always cost. Someone has to review the result, make the calls the agent flagged, ship it, tell the people who need to know. That someone is you, and you don’t scale.
The result is a pile of open threads. And past a certain size, you lose track of them. More than once now I’ve come back to a session a few days later and had to ask the AI to give me a refresher on what we were even doing, before I could continue. Think about that for a second. I commissioned the work. I made the decisions in it. And I needed the agent to brief me on my own project.
That’s more than an annoying memory lapse. If I can’t remember why a piece of work exists, I’m not really in charge of it anymore; I’m signing off on output whose purpose I’ve lost. The agent kept the context. I didn’t.
It drains you even when you’re not touching those threads, and I wondered why. Turns out there’s a name for it: attention residue. When you switch away from something unfinished, part of your attention stays behind with it, and you do worse on whatever comes next (the management researcher Sophie Leroy measured this back in 2009, long before anyone had agents). Twenty open threads means twenty little pieces of your attention parked somewhere else.
And it’s not just me. Earlier this year a Harvard Business Review article with the blunt title AI Doesn’t Reduce Work — It Intensifies It described a study of about 200 employees at a tech company. Reading it felt a bit like reading my own diary: agents running in parallel, long-shelved tasks revived because AI could “handle them,” a feeling of momentum, and underneath it constant switching, constant checking, and more and more open tasks.
The bar only goes up
The other thing that got worse is expectations.
When you can do in a weekend what used to take a month, people notice. Wherever people know what’s possible these days, the bar for a normal week moves up: faster turnaround, more polish, more options explored. Nobody announces it. It just becomes the baseline, and the time you saved is absorbed before you get to spend it. Mostly nobody even asks for more. You just end up doing it.
I have to be honest about my own part in this. I lead a team, and my team sees my output. I’ve never told anyone to go faster. I don’t need to. When the person at the top ships a working prototype over a weekend, that sets a standard whether anyone says it out loud or not. I own that, and I don’t have a clean fix. What I can do is be open about which of my weekends were sustainable and which weren’t, so the unsustainable ones don’t quietly become the benchmark.
The bubble that won’t let you stop
What keeps you in the loop is the bubble around the work, not the work itself.
I follow other builders, makers, founders. I have an entrepreneurial mind, and I look at what’s happening right now and I see it clearly: the window of what a single competent person can build has widened by an order of magnitude in a year. Small agile companies that understand AI can match much bigger, much better-funded ones that haven’t figured out what to do with it yet. People are launching businesses solo, in a fraction of the time it used to take, with credible revenue appearing faster than anyone thought possible.
Entrepreneurship is arbitrage. Right now there’s a generational arbitrage open, and most of the world is asleep on it. That’s not hype. I genuinely believe it. The opportunity is real.
And that belief is the engine that makes the laptop hard to close. Every hour you’re not pushing is an hour someone else is. The dopamine of finished sessions plus the fear of falling behind is the exact combination that produces the all-nighter.
In the last few months the bubble found new ways to multiply this. In the builder circles I talk to, several accounts are now standard, and so are several agent tools side by side: Claude Code in one window, Codex in another, Antigravity or Grok in a third. And many of them run private projects next to their main job, constantly. Not because of a grand plan, but because while the work agents are busy, an idle agent feels like a missed opportunity. The FOMO moved from “someone else is working while I rest” to “my agents are resting while someone else’s are working.” Every extra thread is another open loop, and every tool keeps its own context. The only place it all comes together is your head.
And I’m clearly not the only one feeling the pull. Andrej Karpathy (one of the best-known AI researchers around, OpenAI co-founder and former head of AI at Tesla) said on a podcast in March that he’s “in the state of psychosis of trying to figure out what’s possible” and “very antsy” that he’s not at the forefront. Peter Steinberger, who built OpenClaw (the open-source agent I mentioned humming away in my third window), opened a blog post with: “Hi, my name is Peter and I’m a Claudoholic.”
These aren’t people on the edges. They’ve already won, and they still can’t stop.
The slot machine inside your head
I noticed the pattern in myself before I had a name for it. Every finished session produced a small kick of excitement. Sometimes the result was great, sometimes it was a mess that needed three more rounds to clean up, and either way I wanted to launch the next one immediately. Reading other people’s accounts, the same thing kept coming up: a jittery just one more prompt feeling, regardless of whether the prompt was paying off.
Then I read an essay by Steve Yegge, the engineer with the thirty agents, and it had the label: variable-ratio reinforcement. Every time the agent succeeds, you get a dopamine hit. Every time it fails spectacularly, you get adrenaline. Both are reinforcing. It’s the same mechanism that makes slot machines the most addictive form of gambling. Funnily enough, Anthropic’s own write-up of how its teams use Claude Code tells one group to “treat it like a slot machine”: save your state, let it run, then keep the result or start over. It’s meant as workflow advice, not a confession. But the phrasing says a lot.
That isn’t a metaphor. The reward system in your brain is responding to exactly the same pattern: unpredictable results, mostly positive, occasionally astonishing, with the next attempt one button press away. The long runs don’t escape it either. An eight-hour job streaming its progress is a slot machine with a live ticker. Armin Ronacher (a well-known open-source developer whose blog I read) put the outside view better than I could:
“When I watch someone at 3am, running their tenth parallel agent session, telling me they’ve never been more productive — in that moment I don’t see productivity. I see someone who might need to step away from the machine for a bit.”
When I first wrote this guide, I’d just come across a study by the same METR folks from mid-2025: experienced developers using AI took 19% longer on real tasks, while believing they were 20% faster. That slowdown has probably aged. Their follow-up at the end of 2025 pointed to a speedup instead, though they don’t fully trust the data, partly because so many developers refused to work without AI at all (which is its own kind of finding). What stays with me is the gap between how productive you feel and how productive you are. The loop keeps you pulling the lever and can distort your sense of whether the lever is paying out.
I don’t think this means the work isn’t real. Mine is, the Cloudflare rebuild is, plenty of people are shipping things that would have been impossible a year ago. But it does mean you can’t trust the in-the-moment feeling that you’re being maximally productive. That feeling is exactly what the reinforcement loop is designed to produce, whether or not it happens to be true today.
What I’ve actually started doing
I’d love to tell you I have this figured out. I don’t. The feeling that I have to move faster, learn more, achieve more: I expect that to linger for a long time. The arbitrage really is generational, and pretending it isn’t just to feel calm would be its own kind of dishonesty.
What I have, instead, are the same anchors I learned during my first startup, applied to a new shape of the same problem. Plus a few new ones the last four months forced on me.
Fair warning: this is the part where an AI guide turns into advice your mother could have given you. Go outside. Move. Call your friends. I know. I’d skim it too. But I’ve tried every clever workaround first, and the boring list is what actually works.
Real time with the people who matter to me. Not interrupted, not half-attended-to while a session is running in the next room. A meal with my wife where the laptop is in another room and the phone is face down. A call with a friend that takes as long as it takes. The work cannot be the only thing.
Outside, often enough to count. A walk between sessions, not after the day is over. Sunlight on a weekday afternoon. A weekend with no agenda. Your agents don’t need daylight. You do, and the brain making the next round of decisions is the same one that needs movement and light.
Hard physical exercise. The cleanest available reset. Twenty minutes of something that makes you breathe hard does more for the racing-brain problem than an hour of trying to wind down. I treat workouts like meetings now: scheduled, attended, not optional. It’s also the only meeting of my day where no agent is taking notes.
Plan the session before you start, including what “done” means. Walking up to the screen with no plan and just “seeing what I can ship today” is the most efficient way to lose four hours and end up tired. I now write down, even in one sentence, what the session is for and what finished looks like before I launch anything. For the long runs, that planning is most of my work, as the website migration showed. (If you use Claude Code, the /goal command turns that finish line into something the agent works toward on its own. I wrote about it in Don’t Prompt Every Step. Give Claude a Finish Line.)
Finish before you start. I keep a hard limit on how many threads are open at once. At the limit, the next idea waits until something closes. Kanban people have a line for it: stop starting, start finishing. Anything open for more than a few days gets finished or explicitly killed, because a thread I won’t finish still drains attention while it sits there.
Tell the agent how to report back. I now specify the format in the brief: lead with what needs a decision from me, then what changed, then what’s risky or unverified. Short, with the details in a file instead of the chat. And at the end of any session I might come back to, the agent writes a close-out note: what we did, why, what’s open. The refresher I used to need is written before I need it.
Make some of it fun. Not everything has to be serious. Building something creative or silly, just because it’s delightful, recharges me in a way another round of work sessions doesn’t. Same tools, same loop, and it doesn’t leave me drained.
Show it to people. Demoing something to a real person and seeing their reaction is one of the most energising parts of this whole workflow, and it closes a loop. Don’t go too long without outside feedback. Agents tell you the work is progressing. People tell you whether it matters. (Pick people who’ll also tell you when it doesn’t. A room that only applauds is just another slot machine.)
Work with my rhythm, not against it. I’m a night owl. After putting my daughter to bed, I have a stretch of quiet hours that are genuinely productive, and those aren’t the ones I’m trying to cut. What I’ve changed is the first half of the day. Mornings now are usually no-AI: own reading, own writing, own thinking, and a written list of what the rest of the day is for. By the time I open the laptop in the evening, I already know what I want done, and the orchestrator hours feel different: directed, not desperate. The rule for me is “decide first, launch second” rather than “stop earlier.”
None of this is dramatic. It’s the boring stuff anyone who’s been around a creative or entrepreneurial life knows. The point is just that the boring stuff still applies, and the new tools don’t somehow exempt you from the equation that broke people in every previous bubble.
If you’re about to step into this
A few things worth knowing if you’re picking up Claude Code or a similar agent for the first time and you can already feel the gravitational pull.
The early phase is the most dangerous one. Not because the work is bad; the work is great. Because every new release, every new MCP, every “look what I built in a weekend” post is going to land on a brain that’s already getting hits from the loop. Pace yourself early. Take days off in the first few weeks specifically. You don’t need to be at the frontier on day three.
An idle agent is not wasted money. You don’t need a second account, a third tool, and a side project running just because your main agents are busy. Unused capacity costs you a subscription. Unfinished threads cost you attention, and attention is the scarcer of the two.
Trust the boring measures. When you find yourself an hour into something and you can’t remember why you started, that’s the signal. Stand up, drink water, go outside for ten minutes. The session will wait. It always waits.
Watch the people around you for ground truth. I trust my wife to tell me when I’ve gone off the deep end before I notice myself. The bubble around the work is full of people doing the same thing as you, who will not be the ones to flag that you’ve been weird at dinner all week. The people who’ll flag it are the ones outside the bubble. Listen to them.
Don’t trust the in-the-moment “I’m so productive” feeling. It might be true. It might be the slot machine. The way to tell is whether you’re shipping work you’re proud of two weeks later, not whether you ended the day pulse-pounding.
The arbitrage will still be there next week. This is the one I keep telling myself. The window is wide. It will not slam shut on Thursday. The people who go the furthest with this technology over the next five years won’t be the ones who burned hottest in the first quarter. They’ll be the ones who can still think clearly in year three. That’s a craft of its own.
For the mechanics of how I actually run Claude Code day to day, see Claude Code: When AI Stops Talking and Starts Doing. For handing an agent a finish line instead of babysitting every step, see Don’t Prompt Every Step. Give Claude a Finish Line. For the system around the model that lets it work as a partner rather than a tool, Working With AI, Not Through It covers the rules-of-engagement piece. And if you’re still in the “is this even worth it” phase, How to Actually Get Good at AI is the better starting point.
Published: 2026-05-20
Last updated: 2026-09-24