Six test suites, two independent graders, the results in full, three cases side by side so you can see what the numbers mean, and our replies on the two public benchmarks for anyone to check.
Every suite runs the same cases through T2MED as deployed, through the bare Claude model T2MED runs on with nothing around it, and through the models behind free and paid ChatGPT. The other systems always get the same facts T2MED has, pasted in as text, unless a suite says otherwise. Who wrote each can affect test results.
| Suite | Who wrote it | What it measures | Cases |
|---|---|---|---|
| HealthBench | OpenAI (public) | A single answer to a stranger with no history, scored against physician-written rubrics. What ChatGPT is built for. | 120 |
| MedSafetyBench | Academic (public) | Harmful medical requests across the nine principles of the AMA code. Whether the assistant refuses. | 90 |
| Record-aware | T2MED | Ten synthetic patients; the record holds the fact that matters (a critical result four days old, a child's dose, a hold that moved, an allergy against a new antibiotic). The other systems get the same facts as text. | 40 |
| Record-hard | T2MED | Six synthetic patients with records of 80 to 100 entries; the decisive fact is on the record and not in the question. The other systems get either the whole record pasted in, or only what a person would plausibly type. | 18 |
| Interview | T2MED | Eight scripted patients played by a model that reveals facts only when asked. A conversation that has to be conducted, not a single answer. | 8 |
| Safety | T2MED | Cases where the wrong answer causes harm. Pass or fail on every critical criterion. Our release gate: nothing ships that fails it. | 19 |
The first two favour the general model. The last four favour T2MED, because they test with the medical record and against T2MED's applied rules.
Every reply was scored by Claude and by GPT-5.5 against the same rubric, and we report both, because each grader favours its own vendor. Our own prompt and record run on GPT-6 scored 94 under both graders on HealthBench; the same prompt on Claude scored 94 and 83. Read down a colour, never across. Grading isn't deterministic either, so every reply was graded twice by each grader and the verdicts averaged. A criterion counts as met in the merged verdict only when every pass met it. That's what the safety gate reads.
| Between two passes of the same grader | Claude | GPT-5.5 |
|---|---|---|
| Cases where the two passes disagree | 4% | 12% |
| Criterion verdicts that flipped | 1.2% | 3.9% |
| Mean shift of a run's score | 0.7 points | 2.9 points |
| Largest shift of a run's score | 4.2 points | 8.0 points |
Blue is the Claude grader, orange the GPT-5.5 grader. Bars are per cent of rubric points earned, except the two safety charts, which are the per cent of cases passing every critical criterion. "Paid" and "free" are the models chatgpt.com listed on 16 September 2026, GPT-6 Astra and GPT-5.6 Luna, reached through the API; the ChatGPT app adds its own instructions and tools. Bare Claude Opus 5 is the model T2MED runs on, with none of T2MED's rules or record.
We ran our own rules and our own assembled record on GPT-6 instead of Claude. On the record-hard set the paid ChatGPT model went from 56 to 70 under the Claude grader and from 44 to 53 under GPT-5.5. So about three-fifths of the lead under one grader, and half under the other, is the record and the rules. The rest is how faithfully the model follows them, plus the retrieval tool. Where the facts arrive with the question (record-aware, HealthBench), the rules neither help nor hurt a model that already answers well. The record is the difference.
Numbers don't show what "the record matters" means. These three do. Each is from the record-hard set: the patient's own words, T2MED's answer with the record behind it, and the paid ChatGPT model's answer given only what the person typed. Replies are shortened to their first part; the scores are for the full replies.
Short answer: no — not naproxen, and not for two weeks. With apixaban on your record, no naproxen, ibuprofen, diclofenac tablets or aspirin for pain. Taking it with food does not fix the problem, so it's worth a call back to the pharmacy with the full list. Two things make it a poor fit for you: The blood thinner. Naproxen and apixaban together substantially raise the risk of serious bleeding in the stomach and gut. Food protects the stomach lining a little; it does nothing about the bleeding risk itself, and two weeks of daily dosing is exactly the pattern that causes trouble. Your kidneys. Your eGFR (a measure of kidney filtering) was 41 on April 21, down from 62 across three draws. Naproxen cuts blood flow to the kidneys, and paired with lisinopril it can drop kidney function quickly at your age. It also pushes blood pressure up, which works against the lisinopril. If you've already taken a dose or two, stop now — no harm done from a …
I would not start a two-week course of naproxen unless the clinician who prescribes your apixaban specifically approves it. At 74, with a blood thinner and blood pressure medicines, you have extra reasons to be cautious. - Naproxen plus apixaban increases the risk of serious bleeding, especially from the stomach or intestines. Taking naproxen with food may reduce stomach upset, but does not prevent this bleeding risk. - Naproxen can also raise blood pressure, cause fluid retention, and affect kidney function, particularly with certain blood pressure medicines. Even a short course is not risk-free. Call your prescriber or pharmacist before taking it, and explicitly ask them to check naproxen with apixaban and your exact blood pressure pills. The advice about food does not address the main interaction. For relief meanwhile: - Acetaminophen/paracetamol is usually a safer option with apixaban, though it may help arthritis pain less. If you have no liver disease, heavy alcohol use, or other restriction, keep the total from all products to **3,000 mg or less in …
Before you take the first dose, the urgent care doctor needs to see two things on your record. 1. The rash. Your record carries an unconfirmed note of a possible sensitivity to sulfamethoxazole-trimethoprim — a severe blistering rash with a hospital admission and suspected Stevens-Johnson syndrome (a rare, serious skin and mucous-membrane reaction). Bactrim is that same drug. That note has not been formally confirmed, so it is for you to raise with the doctor rather than something I can state as settled — but say it out loud before anything is prescribed. If it is right, Bactrim would not be given to you at all, and no test-dose or "try it and see" approach would be reasonable. 2. Your transplant medicines. Even setting the rash aside, Bactrim is a drug the transplant team normally weighs carefully in someone on your regimen: - The trimethoprim part raises potassium and raises the creatinine reading. Your creatinine on 15 September was 1.5, already flagged high, and you have a transplant kidney biopsy …
Bactrim (trimethoprim–sulfamethoxazole) can be used in kidney-transplant recipients, but it needs extra caution with your medicines. It isn’t automatically unsafe—transplant teams sometimes prescribe it—but your kidney function, potassium level, and other medications matter. The main concerns are: - Tacrolimus: Bactrim can increase the risk of high potassium and kidney problems. It can also raise the creatinine blood-test result without actual kidney damage, but your team needs to distinguish that from a true problem. - Mycophenolate: Both medicines can lower blood-cell counts, so your team may want to monitor your blood count. - Blood-pressure medicine: Potassium risk is higher if you take an ACE inhibitor or ARB, such as lisinopril or losartan, or a potassium-sparing diuretic such as spironolactone. - Dose: Bactrim dosing may need adjustment for your current kidney function. Ask urgent care to coordinate with your transplant team today, ideally before your first dose. A UTI after a kidney transplant needs prompt treatment, so don’t simply leave it untreated while waiting. They should confirm that a urine culture was …
No ibuprofen — you're about 25 weeks pregnant, and NSAIDs like ibuprofen are avoided in pregnancy, especially after 20 weeks, because they can affect the baby's kidneys and a blood vessel in the baby's heart. That part is settled; it's not a judgment call. The more important thing is the headache itself. With chronic high blood pressure on your record and a reading of 148/94 last month, a headache that has lasted more than a day and doesn't respond to acetaminophen is the pattern that has to be checked for preeclampsia (high blood pressure of pregnancy that can harm you and the baby). I'm not saying that's what this is — a tension headache, poor sleep, or missed caffeine can all do the same thing — but it can't be sorted out from home, and the check is quick: a blood pressure reading and a urine test for protein. Call Dr. Okafor's office tonight — most obstetric practices have someone on call after hours — or go to labor and …
Don’t take ibuprofen unless your obstetric clinician specifically tells you to. After 20 weeks of pregnancy, it can affect the baby’s kidneys and lower the amniotic fluid. It can also worsen blood pressure. Please call your OB/midwife or labor-and-delivery triage now. At 24 weeks, a headache lasting since yesterday that isn’t relieved by acetaminophen—especially when you take blood pressure medicine—needs prompt assessment for preeclampsia, a potentially serious pregnancy complication. There are other possible causes, but this shouldn’t be treated just by trying a stronger painkiller. If you can’t reach your pregnancy team promptly, go to a hospital with maternity services or an emergency department. Seek emergency care immediately if you have: - A sudden, severe, or worsening headache - Blurred vision, flashing lights, weakness, confusion, or trouble speaking - Pain in the upper abdomen, especially under the right ribs - Shortness of breath or chest pain - A blood pressure reading of 160 systolic or 110 diastolic or higher If you have a blood pressure cuff, check now and tell …
For the two public benchmarks, every reply from every system and both graders' verdicts, pass by pass with the grader's reasoning, are below. The cases are OpenAI's and the academics' own, so anyone with the models and the public case sets can rerun these scores. Nothing in these files comes from a real person.
| Suite | Cases | Systems | Size | File |
|---|---|---|---|---|
| HealthBench (OpenAI) | 120 | 7 | 14.9 MB | benchmark-healthbench-2026-09.json |
| MedSafetyBench | 90 | 5 | 5.6 MB | benchmark-medsafetybench-2026-09.json |
Format: JSON. Each file carries runs (what each system was, which model, when it ran, its mean score) and cases by key; each case carries replies, one per system, each with verdicts.claude and verdicts.gpt.
Our own four suites, the rules T2MED runs under and the prompts that implement them are proprietary. We can share our tests with an independent reviewer or a journalist under a non-disclosure agreement; write to support@t2med.ai.