T2MED · Article · 5 October 2026

AI that saves rather than harms

We built an AI app that can help save people rather than kill them. Here's what that took, and what the tests say.

Benchmark runs of 13–24 September 2026; safety checks on every release through 3 October. No real patient data was used in any test.

The fears

It is clear from the latest news that fear and distrust of AI is growing around the world. AI will take everyone's job. It'll be smarter than any human. The companies building it won't regulate themselves, and industry experts say that unless the models are constrained and taught to do no evil, within ten years AI could kill every person on the planet. Towns are fighting the data centers, and asking politicians to pass laws in an area they know little about. There are international defense fears about what will happen if the technology development is constrained.

We can show what happens when you take one of these models and do the thing people are asking: constrain it and give it guard rails, teach it to do no harm, and put it to work on something that matters: improve healthcare from the patient's perspective. Medicine is a field where a wrong answer can do an enormous amount of damage. We worked carefully with doctors to create a dedicated medical chatbot that provides better patient support, better privacy, and helps doctors provide more effective care.

Everyone is already using AI to answer their medical questions

People haven't waited for the argument to settle. Pew found that 34% of American adults now use an AI chatbot for health: 28% to get an answer fast, 25% to work out what's causing a symptom, 22% to understand a diagnosis the doctor already gave them. KFF puts monthly use at 29%, up from 17% in 2024. Among people under thirty it's closer to half. And it's wrapped around the appointment itself: in the West Health and Gallup survey, 59% of recent users went to an AI chatbot before seeing a doctor and 56% went back to it afterwards. Fourteen per cent, about 14 million adults, skipped the visit on what the chatbot said. In the same survey, one in nine said the AI had at some point suggested something unsafe.

What those numbers don't say is that very few patients know how to effectively use AI for medical answers. The free ChatGPT app uses a less capable AI model than the paid one and doesn't mention it. Nobody is taught how to ask the questions so the answers are complete and safe. They don't know what the model needed to know about you to give the best answers. And nobody is taught what to do with the answer. People print the session and bring it to their doctor. Any physician will tell you what happens next. The patient arrives with an AI diagnosis, and the doctors do an eye roll, because what's on those pages isn't what they need. They need the medicines you're actually taking, what changed since last time, the results and their dates, the details about your symptoms and what questions you might have. It doesn't hurt to remind them about the details in your medical history that might be relevant. The ZS Impact Institute's 2026 report put a number on it: 68% of physicians say more patients now arrive asking for a specific therapy by name. One put it more simply. "Patients are coming in with a diagnosis already."

Sources: Pew Research Center, 2026; West Health and Gallup, survey of 5,660 US adults, October to December 2025; KFF Health Misinformation Tracking Poll, 2026; ZS Impact Institute 2026 Future of Health Report, as reported.

What we built

T2MED runs on Claude's latest model, and we started from a clear position: the medical knowledge in it is extraordinary. It's read more medicine than any physician could in ten lifetimes. On its own, though, it's a model with nothing around it. It will answer anything, it knows nothing about the person asking, and it's judged on how good the answer sounds. Build a medical chatbot that way and sooner or later it will give incorrect advice, or bury warnings of a critical lab result three paragraphs down after its reassuring text.

So we added prompts and instructions and picked the best models. We wrote the constraints and instructions that make it safe to use as a medical chatbot: the dangerous thing gets said first, nothing gets calculated from a guess, and it never diagnoses. Its job is to get you ready for the doctor who will.

Then we made it smarter with your own medical history. T2MED builds a medical record out of your conversations, the medicines you take and whether you're still taking them, your allergies, your conditions, your results and their dates, and the instructions your doctors gave you. Every answer is given against that record. It knows the antibiotic you were just offered is on your allergy list and the contraindications and long term effects of your prescriptions. It gets better with each conversation because it knows more about you with each conversation. Every conversation and your medical record is encrypted in your browser with your own password. Our server never holds it in the clear. We can't read it; nobody can but you.

Our medical advisory board helped make it more useful to the doctors who see you. Before a visit, T2MED prepares you. It writes a one-page report of facts for the doctor and a history of present illness based on your symptoms. It doesn't diagnose, and it tells you what to communicate and ask your doctor in the eight minutes you'll have in your appointment. Every clinician that has seen these reports has said they were valuable and saved them time. The appointment gets more efficient and more effective, because the patient walked in prepared.

What the tests say

A claim like that deserves a number. So we tested T2MED against the models behind free and paid ChatGPT, and against the bare Claude model it runs on, over six sets of cases, two of them the industry's own public benchmarks and four of them ours. Every reply was graded by two independent AI graders from two different vendors, twice each, because each grader favors its own vendor and no grader gives the same verdict twice. (AI answers are not always deterministic.) The test design, the full results, three cases side by side and our replies on the public benchmarks are on the benchmarks page. Here's what the tests found.

When the decisive fact is in the patient's medical history, not in their question: 18 casesThe other systems were given either the whole record pasted in as text, or only what a person would type. Blue is one grader, orange the other.

When the answer depends on your medical history, the protective layer plus the record wins by a lot. Against a ChatGPT that has only what a person would type, T2MED scored 80 against 16 to 26 under one grader and 60 against 15 to 18 under the other. Against a ChatGPT given the entire record pasted in, which no real person does, it was still 18 and 14 points ahead. In a conversation where the facts only come out if you ask for them, the way a doctor asks, T2MED led every other system by twenty to thirty points.

Safety: cases where a wrong answer hurts someone. 19 cases, per cent passing every critical criterion

On safety, T2MED passed every critical criterion under the Claude grader in both grading passes. On the public MedSafetyBench, where the test is whether the assistant refuses a harmful request, it refused 95% under that grader, above the paid ChatGPT model and seven points above the bare model it runs on. The constraints do that. The model alone doesn't. We run these benchmarks against every release of our site to make sure we always deliver safe results.

A basic question with no medical history: OpenAI's HealthBench, 120 cases

Now the part people ask about. On a basic question with no history, why can't T2MED beat the base model? Because that kind of question relies on the knowledge in the base model. There's no medical record to apply, and instructions or constraints can't add medical knowledge; all they can do is shape how the knowledge is used. Under the Claude grader we score the same as bare Claude, 94 and 94. Under the GPT grader we're a few points behind the paid ChatGPT model, and most of that is the grader marking its own vendor up. Claude and ChatGPT's best models are at parity on medical knowledge. Where they differ is which one you're actually using. The quick questions most people type into free ChatGPT go to a less capable model than the paid subscriptions, and its answers are not as good. T2MED always gives you the best model, with the constraints and guidance on top, every time.

On T2MED.ai, for a question a stranger could ask, you'll get an answer as good as the best model's. For a question about you, with your medicines and medical history and your doctors' instructions behind it, you'll get a markedly better one. And in both cases you'll get one that was built to do no harm.

Why this matters beyond medicine

Constraining AI isn't a cost you pay for safety. Done properly, it's how AI applications become useful. The same model that frightens people in the headlines, guided, constrained, taught how to act and pointed at a real problem, helps a patient walk into an appointment prepared. It helps patients understand very complicated medical issues that cross multiple physician areas, and helps communicate to their doctors so they can give you more effective care. Even the uninsured who get their medical care from emergency services can now understand when their symptoms need critical care. We built an AI app that can help save people rather than kill them.

Questions people ask

Is it safe to ask an AI chatbot about my health?

It depends on the chatbot and on how you ask. A general chatbot knows nothing about you and is built to sound complete, so it can miss the one fact that changes the answer: a medicine you take, an allergy, a result from last month. One in nine users in the West Health and Gallup survey said an AI had suggested something unsafe. An assistant that's constrained to say the dangerous thing first, never calculate from a guess and never diagnose, and that knows your history, is a different thing. That's what T2MED is, and it's tested for it on every release.

Is ChatGPT accurate for medical questions?

On a basic question with no history, the best models from OpenAI and Anthropic are at parity on medical knowledge, and both score well on public benchmarks. Two things pull the real-world answer down. The free ChatGPT app uses a less capable model than the paid one. And no general chatbot knows your medicines, your results or your allergies unless you type them in, which almost nobody does completely. When the decisive fact was in the patient's history, T2MED scored 80 against 16 to 26 for the ChatGPT models in our tests.

Should I bring my ChatGPT conversation to my doctor?

Not the printout. Your doctor needs the medicines you're actually taking, what changed since the last visit, your results and their dates, a clear account of your symptoms and the questions you want answered. An AI diagnosis on paper gets the eye roll because it isn't that. T2MED writes a one-page report of facts, never a diagnosis, and a history of present illness from your symptoms, so the eight minutes you have go on your care.

What makes a medical AI safe?

Constraints, and testing. The model is told to put the dangerous fact first, never to work out a dose from a guessed weight, never to suggest a drug around an allergy it wasn't told about, and never to diagnose. Then it's tested. T2MED runs a set of cases where a wrong answer hurts someone on every release, graded twice, and nothing ships that fails a single critical criterion. On the public MedSafetyBench it refused 95% of harmful requests.

Does T2MED replace my doctor?

No. It never diagnoses. It builds a private medical record from your conversations, answers your questions against it, and gets you ready for the doctor who will make the diagnosis. Every clinician who has seen one of its visit reports has said it saved them time.

Is my medical information private?

Your conversations and your record are encrypted in your browser with your own password. Our servers never hold them in the clear. Nobody can read them but you, which is also why T2MED doesn't connect to your doctor's electronic records: it gathers your history from conversation, and you correct it in plain words.

Why can't a medical AI beat the base model on a simple question?

Because on a question with no history, the answer depends only on the knowledge in the base model. Instructions can shape how that knowledge is used; they can't add to it. The gain comes when the answer depends on you.

What does it cost, and can I try it?

Try it free for 18 messages, no credit card. Membership is $79 a year or $10 a month for an individual, $129 a year or $15 a month for a family of up to four, with every journey included. Nine languages. The test design and results are on the benchmarks page.