Terminal-Bench 2.1 · Controlled comparison · August 2026
Babysitter adds 4.5 points to Gemini Flash
We ran the full 89-task benchmark twice under identical conditions — same model, same prompts, same time budget. The only difference: one pass ran under Babysitter's orchestration. That pass solved 57 tasks (64.0%) against 53 (59.6%) without it, and cost less.
A flash-priced model near pro-grade results
Every Gemini entry on the official leaderboard is a Gemini 3 or 3.1 Pro, scoring 65.6–73.9%. With Babysitter, Gemini 3.6 Flash reaches 64.0% — within 1.6 points of the published 3.1 Pro rows at a fraction of the model price. Board entries get five attempts per task; ours get one, so our rows are placed for context, not rank.
Accuracy, % of tasks solved Flash + Babysitter Flash without published entries
Where the points come from
Both passes solved a common core of 41 tasks and both missed 20. Babysitter alone solved 16; the bare pass alone solved 12 — a net gain of 4 tasks. Hover the segments for the task names.
The gain is orchestration, not prompt wording. In every orchestrated trial the model first wrote a step-by-step plan, a mechanical checker accepted it, and the model then executed exactly that plan — 1,506 planned steps across 87 trials, median 15 per task. In the pass without Babysitter, a sweep of every transcript found no trace of it: no plans, no orchestration commands, nothing.
One pass per configuration cannot settle a 4.5-point gap beyond doubt: the confidence intervals overlap (53.7–73.2 vs 49.2–69.1) and 28 tasks flipped one way or the other. What stands: under identical conditions the orchestrated pass solved more and spent less, and the orchestration verifiably happened.
One variable
Both passes: Gemini 3.6 Flash driven by Claude Code (pinned version), the same 89 task images, one attempt per task at a 10× time budget, the same 8 hosts and task split, and every trial scored — nothing re-run in either pass. The passes differ only in the rows below.
| With Babysitter | Without | |
|---|---|---|
| Plugin | installed from the pinned copy | never installed — the model's environment contains no trace of it |
| Prompting | turn 1: write a plan · turn 2: execute it | the task text, unchanged |
| Finishing | the trial only ends once the planned run provably completed | the model stops when it decides it is done; any Babysitter activity would void the trial |
| Result | 57/89 = 64.0% | 53/89 = 59.6% |
| Cost | $97–292 | $168–637 |
| Agent time | 27.4 h · median 9.8 min/task | 24.5 h · median 5.7 min/task |
The "without" pass is verified, not assumed: every trial's own records show no Babysitter run was ever created, sampled prompts contain the task text and nothing else, and fleet-wide transcript scans found zero orchestration activity. Costs are ranges because the gateway does not disclose its cached-token pricing; the ranges barely overlap, in Babysitter's favour.
The board runs five attempts per task at standard time limits; we run one attempt at a 10× budget, which is not submission-eligible. The leaderboard chart positions our results for context and claims no official rank.