Skip to content

Real-SWE

Benchmarking frontier AI models on private, real-world, enterprise codebases.

Read the blog

Leaderboard

September 2026
  1. 1
    Fable 5.1
    Claude Code
    Resolution rate: 38.8%
  2. 2
    GPT-6 Astra
    Codex CLI
    Resolution rate: 33.8%
  3. 3
    Gemini 3.8 Flash
    Gemini CLI
    Resolution rate: 31.2%
  4. 4
    GLM 5.3
    Claude Code
    Resolution rate: 28.8%
  5. =5
    Grok 4.6
    Grok Build
    Resolution rate: 23.8%
  6. =5
    Muse Spark 1.3
    Muse Code
    Resolution rate: 23.8%
  7. 7
    Kimi K3
    Kimi Code
    Resolution rate: 18.8%
  8. 8
    GPT-5.6 Sol
    Codex CLI
    Resolution rate: 16.2%
Resolution rate is equivalent to pass@1, averaged over eight independent runs per task. 95% confidence intervals are shown.

We use native harnesses to reflect how enterprise engineers work in practice, evaluating model-and-harness combinations rather than models in isolation.

Private codebases

Agents must navigate proprietary systems whose code and solutions aren’t available on the public internet.

Work with business consequences

Getting billing right, calculating taxes, migrating customers. Changes that affect how a business runs, often across multiple services.

Company-specific complexity

Every company has its own rules and ways of writing code. Agents have to understand those conventions and make changes that work with what’s already there.

Task examples

View analysis
Billing

Tax jurisdiction

Route each invoice through the business's tax mode, including destination pricing via the tax authority provider.

Cloud & DevOps

Multi-region sweep

Discover every enabled AWS region and sweep block-storage inventory without one bad region collapsing the rest.

Data migration

Customer identity migration

Move ownership of offerings and usage from services to customers without breaking existing reads.

Evaluation setup

Each agent was run in an isolated sandbox. All tasks are in Harbor format, and verifiers are injected at grading time. The verifiers are inspired by existing test suites in the codebase or use those tests verbatim.

Read the full analysis

Request sample access

Request access