About this benchmark
It measures how models play a game, scored on process: red flags, key history, indicated exams and tests, a safe plan and a defensible diagnosis. It does not show that any model can or should practise medicine.
The patient, the question classifier and the end-of-consultation marker are Qwen3 8B, the same model human players talk to. All five cases are drafts still under review by the game's author, a GP registrar, and follow Australian practice. Fifteen consultations per model is a small sample.
Code, cases, every transcript and the known issues: github.com/woodytwoshoes/crook-bench. The game: doctorfoo.ai.