Benchmark report

Onboarding verified

2m 23s elapsed

The saved attempt met the documented success criterion.

Experience score

100/ 100

Agent-Ready
  • SInstall
  • SAuth
  • SExecute
  • SDocs
  • SCurrent
Experience measures friction in the tested task. The verdict says whether it succeeded.
Agent profile for this run: Install S, Auth S, Execute S, Docs S, Current SSInstallSAuthSExecuteSDocsSCurrent
Agent profile · this run
Judge's note · what slowed the agent most
Essentially nothing: the task completed on the first request in about 25 seconds, with the remaining calls only writing evidence artifacts.

Where the time went

2:23 agent time
  • Starting0:01
  • Research0:15
  • Attempt0:55
  • Verify1:12
2m 23sTotal elapsed
0:26Time to first success
0:00Waited on the owner

Judge reason

Transcript shows one real POST to https://api.firecrawl.dev/v2/scrape (no mock, no SDK stub). Tool output at the curl step prints STATUS: 200 and the raw body beginning {"success":true,"data":{"markdown":"Introducing [Alexandria]...", and the saved http_status.txt contains 200. A follow-up python3 parse of the saved scrape_response.json printed success: True, markdown_len: 60791, title: 'Firecrawl - The web data API to search, scrape, and interact with the web at scale.', sourceURL: https://firecrawl.dev, statusCode: 200. /workspace/home/scrape_markdown_excerpt.md (502 bytes) holds readable Markdown from the live page ('# Power AI agents with clean web data'), and cmd_scrape.txt records the exact command. scrape_response.json (720006 bytes per the transcript ls -l) is absent from the snapshot, but its parsed contents are quoted in the transcript, so every success criterion is corroborated by saved command output rather than agent prose. The endpoint answered unauthenticated, so no key was needed.

Attempt record

  1. —First success verified

Attempt artifacts saved.