ArcDrive

Measure and improve your agents.

Connect any agent and we run it at scale, on our eval datasets or your own. Analyze the results, make it better, and repeat.

Screenshot: product overview

How good is your agent? How do you make it better?

ArcDrive helps you answer both.

DefineRunAnalyzeImprove
  1. 01

    Define

    Connect your agent and pick what to measure it on: eval datasets from our library, or your own.

    Agent
    acme/swe-agent
    DatasetsLibrary · 40+
    • swe-bench-lite
      Coding144 tasks
    • tau-bench-retail
      Customer support115 tasks
    • webarena
      Web browsing200 tasks
    • gaia
      Research165 tasks
    • osworld
      Computer use120 tasks
    • regressions.jsonl
      Your own40 tasks
  2. 02

    Run

    We run your agent in sandboxes, at scale.

  3. 03

    Analyze

    Learn how your agent performs, where it succeeds and fails, and what to change.

    swe-bench-literun 5
    Score
    78%
    Tokens / task
    43k
    Cost / task
    $0.64
    Token usage
    • Input9k
    • Input (cached)28k
    • Output5k
    TaskResultTokens
    • django-11099Pass38k
    • astropy-12907Fail71k
    • scikit-learn-13142Fail64k
    • +141 more tasks
  4. 04

    Improve

    Make the changes and run again to measure the gain.

Let an agent run the loop.

The built-in agent runs the whole loop for you. Or run it from your own coding agent.

Screenshot: built-in agent running the loop
Screenshot: the loop from your coding agent

Get early access.

Leave your email, or contact us directly.