Improve your agent skills

Run /skill-doctor on your own setup to score past agent conversations and get real improvements to your skills. Supports Claude Code, Codex, and Warp.

Powered by the self-improvement system behind Warp Factories.

Try Warp Factories
Find on GitHub
process

How it works

/skill-doctor runs against your personal coding agent setup to review past conversations and propose merge-ready improvements to your skills.

>_[ fig. 1 · skill doctor · run loop ]
  1. 01

    Aggregate

    Your agent aggregates past conversation transcripts for review (Claude Code, Codex, or Warp).

  2. 02

    Score

    Subagents score these conversations against tested rubrics for efficiency, code quality, and skill coverage. See the scorers

  3. 03

    Improve

    Your agent reviews these scores to propose improvements to your skills.

"warp factories drove our cost per agent pr down by 30%."
— vp engineering, series c infrastructure company
quality loop

Use Warp Factories for automatic self-improvement

evals, benchmarks, and self-improvement loops drive measurable gains.

>_[ fig. 4 · quality loop · sample run ]

evals on your own work

pass96% 
100%50%0%apr 01may 01jun 01jul 01
8/13 07:03fail0.4$0.39
8/13 07:02pass0.8$0.53
8/13 07:02pass0.8$0.46

benchmarks across models

best1.00x 
1.0x0.5x0xapr 01may 01jun 01jul 01
Claude Fable 5pass1.00x$0.53
GPT-5.6 Solpass0.94x$0.41
Gemini 3.6 Flashfail0.89x$0.32

self-improvement loops

auto-fix+3.7 
+4+20apr 01may 01jun 01jul 01
memory updatedpass+0.2#4021
prompts tunedpass+3.2%#4022
regression caughtfail-0.4#4023
early access

Request early access

Set up your first factory with early access.

get up to $10,000 in free factory usage