$ factories benchmarks

Benchmark agents on your own code.

Compare models and harnesses across real tasks, then automatically route work to whatever performs best.

get up to $10,000 in free factory usage

trusted by 800k+ devs at
AsanaDockerGitHubVMwareAmplitudeTeamworksNVIDIARamp
case study

Cutting our cost per PR from $80 to $30

How we used Warp Factories Benchmarks to test models on our own engineering tasks and optimize for cost without sacrificing quality.

read case study
Benchmark results comparing correctness against average cost across model configurations, with gpt-5.6-sol (high) called out as the benchmark winner at 95% correctness and $6.13 average cost.
close the loop

Benchmark. Optimize. Repeat.

A list of past Factory runs, with “Add to benchmark” being chosen from the menu on one of them.

benchmark on your own work

Replay historical tasks from your factories to measure performance against your team’s real workflows.

Setting up a new benchmark run across several model, harness, and scorer configurations.

measure what performs best

Compare models, harnesses, and configurations across quality, correctness, efficiency, and custom scorers.

A benchmark run’s overall recommendation, with a button to apply it.

put the results to work

Update your factory code to route work to the best-performing setup.

benchmarking

Batteries included. No eval stack to stand up.

01

historical task replay

Replay past Factory runs as task sets, starting from their original state.

02

custom benchmark task sets

Curate your own task sets around one repo, one workflow, or one kind of change.

03

test different factory configs

Run identical tasks against different factory versions to see which skills and prompts win.

04

model comparison

Compare frontier and open-weight models directly, on the same real-world tasks.

05

harness comparison

Evaluate agent harnesses side by side, including Warp, Claude Code, and Codex.

06

scorers

Grade runs with LLM-as-a-judge on correctness, quality, efficiency, verbosity, cost, and custom criteria.

07

benchmark reports + visualizations

View performance by task, factory version, and score, with Pareto graphs for cost/quality tradeoffs.

08

custom model routing

Turn results into routing rules that pick the best model for each type of task.

early access

Request early access

Set up your first factory with early access.

get up to $10,000 in free factory usage