I priced Instinct's proactive agents. Here's how to build your own and not burn money doing it
An agent that wakes up every 30 minutes to check for work costs about $46 per user per month. One that wakes on events costs $2. Here is the math, and four fixes.
Instinct raised $1B yesterday. It's completely free, every user gets their own computer, and they're growing ~10% a day.
if you're thinking of building an agent like instinct, this post is about how to not lose money doing it.
the expensive part of a proactive agent isn't doing the work, it's checking whether there's any work to do. an agent that wakes up every 30 min to check costs about $46 per user per month. one that only wakes up when something happens costs approximately $2.
why this matters: if your product is free or flat-priced, every dollar per user comes out of your pocket. at 1M users that's $557M a year vs $28M. and the $46 version is the default in most agent frameworks right now.
The reason this hasn't been priced yet
A few people have already calculated the costs (cpus, power, ram) and the general consensus seems to be that the model ends up costing way more than the machine, and that you can optimize things like compute cost (cpu) by putting the computer to sleep when nobody's using it.
that's true, but it assumes the agent only works when you text it - i.e. it's reactive. the whole point of agents like instinct, muse and openclaw is that the agent does stuff without you having to ask it. it follows up on things like a refund request, watches a price, checks if a fb marketplace seller replied.
Noah Shinn talked about this on Patrick O'Shaughnessy's podcast yesterday. Instinct might wake up at 6am because it knows you get up at 7, look over your day, decide not to bother you, then wake up again at 4pm because something came up. one thing he mentioned was that the compute this needs is "orders of magnitude more" than people expected. i haven't seen someone look at the numbers behind this yet, so i did.
btw for some background i run maritime (yc f26). i've spent a lot of time on this exact problem and ended up building a company around it: we give every agent its own machine that sleeps and wakes. all the math is below and the spreadsheet is at the end, check it yourself.
Some relevant numbers w.r.t. Instinct (and muse)
How a proactive agent checks for work
Say you want your agent to do something when something new appears on your calendar today. the agent doesn't know anything changed unless it looks.
the simplest way to do that (and the default in openclaw) is a heartbeat: every 30 min the agent wakes up, runs a turn, and asks itself "does anything need doing?" (docs). most of the time the answer is no.
that ends up being 48 wake-ups a day, each of which costs machine time and tokens.
Cost 1: machine time
When an agent finishes a task it doesn't go to sleep right away. it waits out an idle timeout in case the user replies. so:
share of time awake = (work time + idle timeout) / check-in intervalsay each check-in is 1 min of work, every 30 min, on instinct's box (2 vCPU, 2 GB) at e2b's list price of $0.133/hr:
| idle timeout | awake | machine cost per user per month |
|---|---|---|
| 15 min | 53% | $52 |
| 5 min | 20% | $19 |
| 45 sec | 6% | $6 |
| never sleeps | 100% | $97 |
with a 15-min timeout, a machine that "sleeps when idle" costs more than half of one that never sleeps. and none of that time is the user actually talking to it.
Cost 2: tokens
every check-in is a model call, and what it costs depends on how much context the model has to read:
| what each check-in reads | tokens per check-in | tokens per month (48/day) | cost per user per month |
|---|---|---|---|
| the full session (openclaw default) | up to ~100k | 146M | $193 |
| instinct-sized profile + context | ~14k | 20M | $27 |
| a trimmed check-in (openclaw isolatedSession) | ~3k | 4.4M | $6 |
| a script checks, model only runs when something changed (2x/day) | ~14k × 2 | 0.85M | $1 |
$1.32 per 1M input tokens (deepseek v4 pro on together), no caching. output left out since most check-ins just return "nothing to do".
a user who texts their agent 10 times a day at 14k tokens a message costs ~$6/month. so with instinct-sized context, checking for work costs 4x more than the user actually using the product.
The total
The naive approach: 30-min check-in, 5-min idle timeout, instinct-sized context every time.
The event-driven approach: wakes only when something happens (~10x a day), 45-sec timeout, a script decides whether to call the model.
| per user per month | 1M users per year | |
|---|---|---|
| naive | $19 machine + $27 tokens = $46 | $557M |
| event-driven | $1 machine + $1 tokens = $2 | $28M |
that's before the agent does anything useful. i'm not saying this is how instinct does it, but if you're building your own, it's the easy default. avoid it.
Getting your cost down from $46 to $2
1. wake on events, not a timer. your calendar can tell the agent when a meeting gets added. so can your inbox. a lot of what agents check for already has a signal like that, so use it to wake the machine. polling every 30 min is 48 wakes to catch maybe 2 events.
2. don't use the model to decide whether to use the model. "did the price drop?" is a comparison, not a reasoning problem. run a script or a small classifier first and only call the model on a hit. hermes can run scheduled jobs as plain scripts with no model at all. a paper from may found a small non-LLM trigger model was more accurate than asking an LLM when to act, and 4-83x faster. HINT: you can build your harness such that it writes these scripts itself.
3. use a short idle timeout when the agent woke itself. if a person texted, keep the machine up in case they reply. if the agent woke up to check a list, nobody is going to reply. on maritime a scheduled wake gets 45 sec and a human wake gets 5 min.
4. batch what can wait. dhravya found instinct's memory takes ~23 hours to update. my guess is that's on purpose: batch inference is up to 50% cheaper if you can wait a day, and a lot of background work can.
The deets
here's the spreadsheet, you can plug in your numbers (and pls send me what you get, im curious).
and if you don't want to build the sleep + wake-on-event stuff yourself, that's what we do at maritime!
thanks to rohan adwankar for the look inside the box, dhravya shah for the memory teardown, and the people who did the infra math this week (here and here).