Measure and improve your agents.
Connect any agent and we run it at scale, on our eval datasets or your own. Analyze the results, make it better, and repeat.
Screenshot: product overview
How good is your agent? How do you make it better?
ArcDrive helps you answer both.
- 01
Define
Connect your agent and pick what to measure it on: eval datasets from our library, or your own.
Agentacme/swe-agentDatasetsLibrary · 40+- swe-bench-liteCoding144 tasks
- tau-bench-retailCustomer support115 tasks
- webarenaWeb browsing200 tasks
- gaiaResearch165 tasks
- osworldComputer use120 tasks
- regressions.jsonlYour own40 tasks
- 02
Run
We run your agent in sandboxes, at scale.
swe-bench-lite144 tasks · 8 sandboxesRunning 0/1440 passed0 failed0%pass rate - 03
Analyze
Learn how your agent performs, where it succeeds and fails, and what to change.
swe-bench-literun 5Score78%Tokens / task43kCost / task$0.64Token usage- Input9k
- Input (cached)28k
- Output5k
TaskResultTokensTurnsTime- django-11099Pass38k122:18
- astropy-12907Fail71k356:43
- scikit-learn-13142Fail64k295:34
- +141 more tasks
- 04
Improve
Make the changes and run again to measure the gain.
swe-bench-lite6 runs
Let an agent run the loop.
The built-in agent runs the whole loop for you. Or run it from your own coding agent.
Screenshot: built-in agent running the loop
Screenshot: the loop from your coding agent
Get early access.
Leave your email, or contact us directly.