Benchmark agents on your own code.
Compare models and harnesses across real tasks, then automatically route work to whatever performs best.
get up to $10,000 in free factory usage
Cutting our cost per PR from $80 to $30
How we used Warp Factories Benchmarks to test models on our own engineering tasks and optimize for cost without sacrificing quality.
read case study
Benchmark. Optimize. Repeat.

benchmark on your own work
Replay historical tasks from your factories to measure performance against your team’s real workflows.

measure what performs best
Compare models, harnesses, and configurations across quality, correctness, efficiency, and custom scorers.

put the results to work
Update your factory code to route work to the best-performing setup.
Batteries included. No eval stack to stand up.
historical task replay
Replay past Factory runs as task sets, starting from their original state.
custom benchmark task sets
Curate your own task sets around one repo, one workflow, or one kind of change.
test different factory configs
Run identical tasks against different factory versions to see which skills and prompts win.
model comparison
Compare frontier and open-weight models directly, on the same real-world tasks.
harness comparison
Evaluate agent harnesses side by side, including Warp, Claude Code, and Codex.
scorers
Grade runs with LLM-as-a-judge on correctness, quality, efficiency, verbosity, cost, and custom criteria.
benchmark reports + visualizations
View performance by task, factory version, and score, with Pareto graphs for cost/quality tradeoffs.
custom model routing
Turn results into routing rules that pick the best model for each type of task.
Request early access
Set up your first factory with early access.
get up to $10,000 in free factory usage