Terminal-Bench 2.1 · Controlled comparison · August 2026

Babysitter adds 4.5 points to Gemini Flash

We ran the full 89-task benchmark twice under identical conditions — same model, same prompts, same time budget. The only difference: one pass ran under Babysitter's orchestration. That pass solved 57 tasks (64.0%) against 53 (59.6%) without it, and cost less.

With Babysitter
64.0%
57/89 solved · $97–292
Without
59.6%
53/89 solved · $168–637
Lift
+4.5 pts
more solved, lower cost
Orchestration verified
87/87
trials ran a validated plan

A flash-priced model near pro-grade results

Every Gemini entry on the official leaderboard is a Gemini 3 or 3.1 Pro, scoring 65.6–73.9%. With Babysitter, Gemini 3.6 Flash reaches 64.0% — within 1.6 points of the published 3.1 Pro rows at a fraction of the model price. Board entries get five attempts per task; ours get one, so our rows are placed for context, not rank.

Accuracy, % of tasks solved Flash + Babysitter Flash without published entries

0255075100

Where the points come from

Both passes solved a common core of 41 tasks and both missed 20. Babysitter alone solved 16; the bare pass alone solved 12 — a net gain of 4 tasks. Hover the segments for the task names.

41both
16Babysitter only
12bare only
20neither

The gain is orchestration, not prompt wording. In every orchestrated trial the model first wrote a step-by-step plan, a mechanical checker accepted it, and the model then executed exactly that plan — 1,506 planned steps across 87 trials, median 15 per task. In the pass without Babysitter, a sweep of every transcript found no trace of it: no plans, no orchestration commands, nothing.

Read the margin honestly

One pass per configuration cannot settle a 4.5-point gap beyond doubt: the confidence intervals overlap (53.7–73.2 vs 49.2–69.1) and 28 tasks flipped one way or the other. What stands: under identical conditions the orchestrated pass solved more and spent less, and the orchestration verifiably happened.

One variable

Both passes: Gemini 3.6 Flash driven by Claude Code (pinned version), the same 89 task images, one attempt per task at a 10× time budget, the same 8 hosts and task split, and every trial scored — nothing re-run in either pass. The passes differ only in the rows below.

With BabysitterWithout
Plugininstalled from the pinned copynever installed — the model's environment contains no trace of it
Promptingturn 1: write a plan · turn 2: execute itthe task text, unchanged
Finishingthe trial only ends once the planned run provably completedthe model stops when it decides it is done; any Babysitter activity would void the trial
Result57/89 = 64.0%53/89 = 59.6%
Cost$97–292$168–637
Agent time27.4 h · median 9.8 min/task24.5 h · median 5.7 min/task

The "without" pass is verified, not assumed: every trial's own records show no Babysitter run was ever created, sampled prompts contain the task text and nothing else, and fleet-wide transcript scans found zero orchestration activity. Costs are ranges because the gateway does not disclose its cached-token pricing; the ranges barely overlap, in Babysitter's favour.

Not leaderboard rows

The board runs five attempts per task at standard time limits; we run one attempt at a 10× budget, which is not submission-eligible. The leaderboard chart positions our results for context and claims no official rank.