- Judge agent runtimes by what happens on the hundredth run, not the demo run. Cost and reliability diverge sharply at scale between compiled and LLM-per-step architectures.
- Repeatability, meaning same input and same output, is a buying criterion you can test directly by asking whether the vendor re-reasons every run or executes compiled code.
- Auth and session handling, plus observability with a full audit trail and video replay, are production requirements, not add-ons. Ask how credentials are stored and how failures get diagnosed.
- Compile versus LLM-per-step is the architectural fork underneath every other criterion in this guide.
- Self-healing runs are a useful signal that a vendor has separated build-time reasoning from runtime execution.
A vendor walks you through an agent that signs into a partner portal, pulls a weekly report, and posts a summary to Slack. The live demo works. You send the recording to leadership and approve a pilot. Then the failures start: an expired session, a CAPTCHA out of nowhere, and logs that don't say what broke when the layout changes.
The demo isn't the hard part; the hundredth run is. Most teams buying agent builders see the same arc: an impressive first run, then cost, auth trouble, and unannounced failures on later runs.
This guide is for evaluating vendors before you buy, not for building an agent from scratch. If you want a build tutorial, start with How do you make agents deterministic?. The sections below cover the questions to ask about repeatability, sessions, observability, cost at scale, and self-healing before you sign.
Why the demo is the wrong test
Every agent runtime looks good once. A model reasons through a login form and extracts the data. It's impressive, but it isn't the question you're buying for.
The value of letting a model reason through a task is highest the first time you do it, and lowest the hundredth time. By run 100, the task is the same task. The ambiguity is gone. Paying an LLM to re-derive the plan every time adds cost without adding judgment.
A successful demo login is only the first check. What matters more for buying is cost and failure rate on later runs, including whether authentication and sessions still hold when the job runs unattended. That's the standard this guide holds a vendor to.
Repeatability: same input, same output
Ask a vendor whether the agent re-reasons through the task on every run, or executes compiled code. Those are very different products wearing the same demo.
An LLM-per-step loop reads the page, decides what to click, and generates the next action, every single time, for every run. It's flexible, but nondeterministic. Research cited on Airtop's agency spectrum page says Computer Use models show unexpected or incorrect behavior between 20% and 60% of the time, depending on how complex the task is. At that rate, a run can succeed one day and silently fail the next, with no code change in between.
The alternative is a runtime that treats reasoning as a build-time cost, not a runtime cost. Airtop's approach is to reason once, at build time, compile the result into code, and run that code afterward, calling a model only where a step genuinely requires judgment. The result behaves like software: same input, same output, every run.
When you evaluate a runtime, ask the vendor to list which steps become reusable code, such as navigation, clicking, typing, and control flow, and which steps still call a model, such as reading a page or judging whether a condition is met. A vague answer often means most of the run is still model-driven.
Auth and sessions as a buying criterion
Sessions expire mid-run. CAPTCHAs appear without warning. Two-factor authentication asks for a code nobody is watching for. None of that shows up in a five-minute demo, and all of it shows up by week two.
Ask a vendor how credentials are stored, and whether they can appear in agent code or logs. If they can, raise that in security review before rollout, especially for customer or regulated data. Ask whether the system can sign in once, stay signed in, maintaining a live session across runs instead of re-authenticating from scratch every time.
Ask how the runtime handles CAPTCHA solving and proxy rotation, and separately, how it handles single sign-on. These aren't edge cases for an operator running agents against real vendor portals and gated dashboards. They're the default terrain. A runtime that treats browser automation as the fallback for APIs when they exist, the browser when they don't is acknowledging something true: most of the web that matters to a GTM workflow doesn't have a clean API. API-only automation covers roughly 4 percent of the web, according to Airtop's Claude Code integration page, which puts a number on why login handling belongs in the buying decision, not the appendix.
Observability: what debugging should look like
When a run fails, the person on call needs an answer fast, not a guessing game against a log line. Ask a vendor what a failed run looks like from the inside.
The standard to hold a vendor to: a full audit trail of every action the agent took, the data it touched, and video replay of the session. Debugging should look like watching a recording back. Airtop states this as its own claim: 100 percent traceable runs, according to the Claude Code integration page. Treat that as the standard to hold any vendor to when you evaluate their observability, whether or not they meet it the same way Airtop describes.
Observability isn't a nice-to-have for a compliance review either. If the runtime touches customer data, health data, or anything behind a login your legal team cares about, ask about certifications directly. Airtop reports SOC 2 Type II and HIPAA compliance on its Claude Code page. Confirm certifications before a security review blocks a rollout.
Cost on the hundredth run
The difference between architectures shows up on the invoice, not in a design review. LLM-per-step cost scales linearly: more steps and more runs mean more tokens, every time. Compiled cost is dominated by infrastructure, not reasoning, because the reasoning already happened once.
Airtop publishes a direct comparison on the same multi-step task. A compiled Airtop agent finished in 1 minute 21 seconds for $0.063. Claude Code running Opus 4.7 took 7 minutes 58 seconds and cost $6.26 for the same task, per Airtop's how-it-works page. The same task run against Claude with Sonnet took 17 minutes 21 seconds and cost $3.01, which the same page describes as 48 times more expensive and 13 times slower than the compiled version. Codex ran the task in 7 minutes 45 seconds for $5.23.
That cost difference widens, not narrows, as run volume grows. At scale, Airtop reports up to 6 times faster runs and 1 percent the cost, according to its Agent Builder page, with per-run costs under $1 and 10 to 100 times cheaper than LLM-per-step agents, per the Claude Code page. One worked example: a coding-agent workflow that mines LinkedIn engagement data runs as an agent that runs the same way every time from the terminal, instead of re-planning the scrape on every invocation. Ask a vendor to show you their version of this comparison, on your actual task, not a curated one.
Compile vs LLM-per-step: the fork underneath everything
Every criterion above traces back to one architectural decision: where does the model belong, and where should code run instead? Airtop's position, laid out in its code-first vs LLM-first agents comparison, is that planning is judgment and belongs with a reasoning model, while execution is repetition and belongs in compiled code. The site's own framing is build once, run forever.
One way to evaluate a vendor: ask what happens when a run breaks, say a button moves or a page layout shifts. If the answer involves an engineer editing a selector, the model and the execution layer probably haven't been separated. If compiled agents heal themselves when a run breaks, that's a signal the vendor has drawn the line between build time and runtime, rather than just describing one. That self-healing is described narrowly, for a moved button or a changed layout, not as a fix for a full redesign, so it's worth asking a vendor exactly how far their self-healing extends.
Airtop reports up to 100 simultaneous sessions and up to 100 times greater efficiency than uncompiled LLM agents, figures shown on its homepage and Agent Builder page. Whichever runtime you land on, that self-healing question is worth asking directly, because it exposes the architecture the sales deck won't.
A short buyer's checklist
Before signing anything, run the vendor through six questions:
- Repeatability: does the agent re-reason every run, or execute compiled code with the same input producing the same output?
- Auth and sessions: how are credentials stored, and does the runtime handle CAPTCHA, proxies, and two-factor or SSO flows without manual intervention?
- Observability: is there a full audit trail with video replay for every run, or just log lines?
- Cost at scale: what does the hundredth run cost, not the first one, and can the vendor show a real benchmark?
- Self-healing: what happens automatically when a page changes, without an engineer editing code?
- Compliance: does the vendor carry SOC 2 Type II and HIPAA, if your data requires it?
This buying decision usually sits with the GTM engineer or operations owner who keeps growth systems running day to day. For that role context, see What do GTM engineers do?. Use these six questions during the vendor evaluation, while you can still compare answers across options and before you commit to a contract.
FAQs
Is this guide about buying an agent runtime or building one?
Buying. This guide gives you criteria for evaluating a browser-agent vendor or platform before you commit, not a tutorial for building your own reliable agent. If you're building, the compile vs LLM-per-step distinction described here still applies, but the questions shift from "what do I ask a vendor" to "what do I architect myself."
What's the practical difference between a compiled agent and an LLM-per-step agent?
A compiled agent reasons through a task once, at build time, then runs the resulting code the same way every time, calling a model only where judgment is genuinely required. An LLM-per-step agent re-reasons through the task on every single run, so every run carries the cost and variability of a fresh model call. Ask a vendor which one you're actually buying.
How should a vendor store and handle login credentials?
Credentials shouldn't live in agent code or show up in plaintext logs. Ask specifically how sessions persist across runs, since a runtime built to sign in once and stay signed in avoids re-authenticating from scratch on every execution. If a vendor can't answer clearly, treat that as a red flag before a security review does.
Why does per-run cost change so much as volume grows?
LLM-per-step architectures scale cost linearly with steps and runs, since every step is a fresh model call. Compiled architectures front-load the reasoning at build time, so cost at scale is dominated by infrastructure rather than tokens. The cost difference between a compiled agent and an LLM-per-step agent widens as usage grows.
What counts as real observability, versus a marketing checkbox?
A full audit trail of every action and every piece of data touched, plus video replay of the run, is the standard. If debugging a failure means guessing from a single log line instead of watching a recording of what happened, observability isn't there yet. Ask to see an actual failed-run trace before you buy, since a slide about it isn't the same thing.
How do CAPTCHA and proxy handling factor into a reliability decision?
Most of the web that matters for GTM work sits behind logins, bot protection, or legacy vendor portals, not clean APIs. A runtime that can't handle CAPTCHA solving and proxy rotation will fail where the work actually is, not in some edge case. Confirm this coverage as core infrastructure, not an optional add-on.
What does self-healing indicate about a vendor's architecture?
Self-healing means a broken run, caused by a moved button or a changed layout, gets fixed automatically instead of requiring an engineer to edit a selector. That only works if the vendor has genuinely separated build-time reasoning from runtime execution. It's a narrower guarantee than it might sound: it addresses small drift, not a full redesign of the target page. If a vendor can't describe how self-healing works, the separation probably doesn't exist yet.
Does compliance actually matter for an agent runtime, or is it boilerplate?
If the runtime touches customer data, health data, or anything behind a login your legal team cares about, SOC 2 Type II and HIPAA aren't boilerplate. Confirm certifications before a rollout, not after a security review stalls it. Ask the vendor directly rather than assuming coverage from the sales deck.
Try Airtop for free
Most of the criteria in this guide come down to one thing: watching how a runtime behaves on a repeat run, not a demo run. Spin up your first agent in five minutes and see whether it holds up on run two, then run 100. Try it for free.






