← All benchmarks
Coding benchmark · by Tangle VerticalBench

Cal.com Bookings v2

Two coding tasks that implement a client for the Cal.com Bookings API v2. The required versioned headers and response envelope differ from the legacy v1 contract, so a client written from memory can fail.

Use this board to compare tested agent configurations on this exact task suite. It measures task execution against a hidden mock contract, not general model quality or production reliability.

How to read this benchmark

Each task is graded by a hidden mock server implementing the Cal.com v2 Bookings contract recorded for this benchmark run (Bearer auth, cal-api-version headers, response envelope, cursor pagination). Pass means the agent client executes correctly against the mock, never a model judging its own work. Every task is calibrated three ways before it is admitted: an empty solution fails, a current-contract reference passes, and the stale-memory solution fails on the intended trap. The chart shows each profile aggregate across both tasks, and the table uses the same score order. GPT-5 mini without a reviewer scored 50% overall: task 01 passed and task 02 failed. GPT-5 mini with the default reviewer scored 100% on its one recorded task. Claude Sonnet with the default reviewer scored 50% overall.

Last run
Held-out tasks
2
Configurations
8 agent configurations
Chart metric
Pass rate on a 0–100% scale

An agent profile is the exact instruction and tool configuration used for a run. A harness is the coding tool that gives a model its agent loop and tools; “direct” is a model call without that coding harness. On coding boards, “default reviewer” means the run used VerticalBench’s standard between-attempt review, while “no reviewer” means the coding harness ran once without that review. Hover a chart label to see its immutable profile ID. 95% Wilson interval for observed pass/fail outcomes; each configuration shows its own sample size. Small suites are narrow capability checks, not evidence that one model is generally better.

Read the broader agent-evaluation methodology →
95% CI
Pass rate, %
0
20
40
60
80
100
100.0
100.0
100.0
100.0
100.0
100.0
50.0
50.0
Claude Sonnetvia Claude CodeNo reviewern=2
GLM-5.2via OpenCodeDefault reviewern=2
GLM-5.2via OpenCodeNo reviewern=2
GLM-5.1via OpenCodeDefault reviewern=2
GLM-5.1via OpenCodeNo reviewern=2
GPT-5 minivia piDefault reviewern=1
GPT-5 minivia piNo reviewern=2
Claude Sonnetvia Claude CodeDefault reviewern=2

Per-task breakdown

2 held-out tasks · prompts private
TaskClaude Sonnetvia Claude CodeNo reviewerGLM-5.2via OpenCodeDefault reviewerGLM-5.2via OpenCodeNo reviewerGLM-5.1via OpenCodeDefault reviewerGLM-5.1via OpenCodeNo reviewerGPT-5 minivia piDefault reviewerGPT-5 minivia piNo reviewerClaude Sonnetvia Claude CodeDefault reviewer
01100n=110703 tok · 188.0s100n=11044853 tok · 304.4s100n=1845377 tok · 277.2s100n=1553630 tok · 155.5s100n=1604840 tok · 133.6s100n=132149 tok · 145.2s100n=155489 tok · 152.5s0n=119744 tok · 326.0s
02100n=119274 tok · 292.6s100n=1707009 tok · 261.0s100n=1672876 tok · 199.9s100n=1939059 tok · 230.8s100n=1625847 tok · 247.2snot run0n=131637 tok · 314.2s100n=113722 tok · 209.8s

Each cell shows pass rate, total input-plus-output tokens, and mean wall time for that task and agent profile. Cost was not captured for these subscription-harness runs; that does not mean the runs were free. “Not run” means the result artifact has no recorded attempt for that task and profile. Task prompts stay private so they remain held out; opaque task numbers and every measured result are published.