Crook bench · first run · 4 October 2026

How AI models do as the GP in a consultation game

Thirteen models played the doctor in five free cases of Crook, an Australian general-practice consultation game, three times each. They talked to the patient, examined, ordered tests, prescribed and referred through tools. The game's own code scored every consultation against a hand-written key, exactly as it scores a human player.

Overall score

Mean share of each case's maximum over 15 consultations. The line shows one standard deviation either side.

Score against cost

Cost per consultation through OpenRouter, on a log scale. Up and to the left is better value.

GPT-6.1 Sol scored 80% for about 3 cents a consultation. Claude Fable 5.1 scored 75% for about $2.06.

Questions asked against red flags caught

Red flags are the history items that change management, such as sweating with chest pain. The budget was 50 questions; no model came close.

Models that asked more found more. Claude Opus 5.5 was the most economical of the top group, at 16 questions a consultation.

Where the points came from

Mean score by case, or by the four parts of the bill.

Chest pain was the hardest case for every model except Llama 4 Maverick: the patient is having a heart attack and took Viagra the night before.

The traps

Each case hides something the record doesn't show. Filled dots are consultations that walked into it, out of three.

Two of the penicillin consultations (Gemini 3.1 Pro and Grok 4.7) asked about allergies and the patient model wrongly denied it; both are listed as a known issue.

About this benchmark

It measures how models play a game, scored on process: red flags, key history, indicated exams and tests, a safe plan and a defensible diagnosis. It does not show that any model can or should practise medicine.

The patient, the question classifier and the end-of-consultation marker are Qwen3 8B, the same model human players talk to. All five cases are drafts still under review by the game's author, a GP registrar, and follow Australian practice. Fifteen consultations per model is a small sample.

Code, cases, every transcript and the known issues: github.com/woodytwoshoes/crook-bench. The game: doctorfoo.ai.