Legacy Bench
Legacy-Bench results and methodology for AI agents working on legacy software.
Benchmark from Factory measuring AI agent performance on legacy engineering tasks across COBOL, Java 7, BASIC, C89, Fortran, and Assembly.
Results — Overall Pass Rate
1Factory Droid (GPT-5.3 Codex)
42.5%
2Factory Droid (GPT-5.4)
40.7%
3Codex CLI (GPT-5.4)
40.1%
4Codex CLI (GPT-5.3 Codex)
39.4%
5Factory Droid (Gemini 3.1 Pro)
38.7%
6Gemini CLI (Gemini 3.1 Pro)
36%
7Factory Droid (Claude Opus 4.6)
34.6%
8Claude Code (Claude Opus 4.6)
33%
9Factory Droid (GLM-5)
32.2%
10Cursor (Composer 2)
31%
11Factory Droid (Gemini 3 Flash)
27.2%
12Factory Droid (Kimi K2.5)
16.9%
Last updated: April 2026
Methodology
| Stage | Description |
|---|---|
| Task set | Hundreds of tasks across six legacy language families, with ten representative open samples |
| Task format | Natural language instruction, containerized source environment, reference solution, and hidden verification tests |
| Task types | Bug fixing, implementation, migration, and other legacy engineering work |
| Evaluation | Harbor-compatible tasks requiring agents to understand the specification, produce working code, and pass verification |
| Scoring | Pass rate across hidden tests for 12 model-agent combinations |
Benchmark Mix
| Language | Share | Example domains |
|---|---|---|
| COBOL | 46% | Financial settlement, payroll processing, insurance claims, telecom billing, VSAM file handling |
| Java 7 | 32% | Enterprise middleware, CDR processing, warehouse logistics, binary parsing, EJB patterns |
| BASIC | 6% | Business applications, accounting, data processing |
| C89 | 5% | Systems programming, low-level debugging, protocol implementation |
| Fortran | 5% | Scientific computing, numerical methods, physics simulation |
| Assembly | 5% | x86 firmware parsing, protocol decoding, hardware simulation |
Agents score highest on Java 7 bug fixing, where compiler and runtime feedback expose errors. COBOL remains hardest: 31 of the 44 tasks no model solved are COBOL.
Read the writeup
Legacy-Bench: Can AI Agents Maintain the World's Most Critical Software?
Read the writeup