Stripe API Integration
Twelve coding tasks built from documented Stripe API changes in 2025 and 2026. Each run is checked against the suite's versioned mock contract as it existed on the published run date.
Use this board to compare tested agent configurations on this exact task suite. It measures task execution against a hidden mock contract, not general model quality or production reliability.
How to read this benchmark
Each task is graded by a hidden mock server implementing the Stripe contract recorded for this benchmark run (endpoints, required params, error shapes), including the trap a from-memory solution falls into. Pass means the agent code executes correctly against the mock, never a model judging its own work. Every task is calibrated three ways before it is admitted: an empty solution fails, a current-contract reference passes, and the stale-memory solution fails on the intended trap.
- Last run
- Held-out tasks
- 12
- Configurations
- 8 agent configurations
- Chart metric
- Pass rate on a 0–100% scale
An agent profile is the exact instruction and tool configuration used for a run. A harness is the coding tool that gives a model its agent loop and tools; “direct” is a model call without that coding harness. On coding boards, “default reviewer” means the run used VerticalBench’s standard between-attempt review, while “no reviewer” means the coding harness ran once without that review. Hover a chart label to see its immutable profile ID. 95% Wilson interval for observed pass/fail outcomes; each configuration shows its own sample size. Small suites are narrow capability checks, not evidence that one model is generally better.
Read the broader agent-evaluation methodology →Per-task breakdown
12 held-out tasks · prompts private| Task | GLM-5.1via OpenCodeDefault reviewer | GLM-5.2via OpenCodeDefault reviewer | GLM-5.2via OpenCodeNo reviewer | GLM-5.1via OpenCodeNo reviewer | GPT-5 minivia piDefault reviewer | GPT-5 minivia piNo reviewer | Claude Sonnetvia Claude CodeNo reviewer | Claude Sonnetvia Claude CodeDefault reviewer |
|---|---|---|---|---|---|---|---|---|
| 01 | 100n=1491534 tok · 565.6s | 100n=1912927 tok · 230.9s | 100n=1406747 tok · 146.2s | 100n=1639456 tok · 425.6s | not run | 0n=16260 tok · 158.0s | 100n=117297 tok · 228.7s | 100n=113582 tok · 249.0s |
| 02 | 100n=11275666 tok · 456.7s | 100n=11480840 tok · 389.0s | 100n=11452625 tok · 441.0s | 100n=11187430 tok · 476.1s | 100n=116015 tok · 284.4s | 0n=17394 tok · 156.9s | 100n=119981 tok · 269.6s | 100n=114733 tok · 222.0s |
| 03 | 100n=11262087 tok · 466.2s | 100n=11222879 tok · 412.6s | 0n=1569558 tok · 515.7s | 100n=1562144 tok · 417.0s | 100n=137915 tok · 226.2s | 0n=121364 tok · 221.6s | 100n=122900 tok · 334.5s | 0n=124426 tok · 370.7s |
| 04 | 0n=1430646 tok · 245.0s | 0n=1695607 tok · 224.8s | 0n=11105811 tok · 586.3s | 0n=1999069 tok · 399.5s | 0n=139287 tok · 178.1s | 100n=115993 tok · 261.5s | 0n=135750 tok · 445.0s | 0n=130302 tok · 379.5s |
| 05 | 100n=11254340 tok · 285.2s | 100n=11337321 tok · 345.5s | 0n=1528087 tok · 238.5s | 100n=1772958 tok · 309.4s | 0n=134791 tok · 190.8s | 0n=1150160 tok · 233.1s | 100n=133576 tok · 429.3s | not run |
| 06 | 100n=11437398 tok · 392.0s | 100n=1781682 tok · 281.6s | 0n=1751733 tok · 417.6s | 100n=1731192 tok · 280.8s | not run | 0n=113780 tok · 391.7s | 100n=127260 tok · 399.8s | not run |
| 07 | 100n=1614058 tok · 376.6s | 0n=1855884 tok · 225.0s | 100n=1575657 tok · 257.0s | 100n=1545798 tok · 281.1s | 0n=116733 tok · 345.8s | 0n=113764 tok · 224.9s | 0n=10 tok · 58.4s | 100n=18051 tok · 126.9s |
| 08 | 100n=1953479 tok · 340.1s | 100n=1671728 tok · 295.2s | 100n=1357551 tok · 294.6s | 0n=11173807 tok · 320.9s | 100n=120595 tok · 267.6s | 100n=114353 tok · 226.2s | 0n=10 tok · 31.9s | 0n=10 tok · 26.3s |
| 09 | 100n=1309610 tok · 143.3s | 100n=1282943 tok · 133.9s | 100n=1471539 tok · 222.5s | 100n=1430056 tok · 409.4s | 100n=18958 tok · 238.4s | 100n=113159 tok · 238.5s | 0n=10 tok · 33.1s | 0n=10 tok · 40.3s |
| 10 | 100n=1368845 tok · 105.8s | 100n=1450584 tok · 137.4s | 100n=1447155 tok · 153.7s | 100n=1392268 tok · 160.3s | 100n=116254 tok · 198.5s | 100n=123711 tok · 199.9s | 0n=10 tok · 27.1s | 0n=10 tok · 42.3s |
| 11 | 100n=11776673 tok · 599.9s | 100n=1817499 tok · 450.2s | 100n=11657603 tok · 555.3s | 0n=11248313 tok · 351.1s | 100n=134369 tok · 244.3s | not run | 0n=10 tok · 26.8s | 0n=10 tok · 27.7s |
| 12 | 100n=1605147 tok · 172.0s | 100n=11092603 tok · 301.1s | 100n=1438368 tok · 168.7s | 0n=1819127 tok · 484.4s | 0n=115881 tok · 229.0s | 0n=110247 tok · 403.0s | 0n=10 tok · 24.1s | 0n=10 tok · 27.5s |
Each cell shows pass rate, total input-plus-output tokens, and mean wall time for that task and agent profile. Cost was not captured for these subscription-harness runs; that does not mean the runs were free. “Not run” means the result artifact has no recorded attempt for that task and profile. Task prompts stay private so they remain held out; opaque task numbers and every measured result are published.