Building a medical AI that's held back on purpose, and measuring it honestly.
The news says AI will take every job, outthink every human and, unconstrained, end us. We built an AI app that can help save people rather than kill them: the medical knowledge in a frontier model, held back by rules a doctor would recognise, made smarter by the patient's own history, and tested against ChatGPT. Here's what that took, and what the tests say.
Six test suites (two of the industry's, four of ours), two graders, two passes each. The test design, the results, three cases side by side, and our replies on the public benchmarks for anyone to check.