Code & Cure
Decoding health in the age of AI
Hosted by an AI researcher and a medical doctor, this podcast unpacks how artificial intelligence and emerging technologies are transforming how we understand, measure, and care for our bodies and minds.
Each episode unpacks a real-world topic to ask not just what’s new, but what’s true—and what’s at stake as healthcare becomes increasingly data-driven.
If you're curious about how health tech really works—and what it means for your body, your choices, and your future—this podcast is for you.
We’re here to explore ideas—not to diagnose or treat. This podcast doesn’t provide medical advice.
Code & Cure
#52 - When "Once A Day" Becomes Eleven Pills
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
What if medical AI looks unstoppable right up until you change the language? One year into Code & Cure, we pull on an unsettling thread: a model can score around 90% on an English medical exam and then crash to about 55% on the same exam in French, with similar drops across other languages. That should give us pause—healthcare doesn't happen in a single language, and patient safety can't ride on English-only competence.
We dig into why this happens by putting human clinicians next to large language models. A doctor doesn't become "less medical" when they switch to Spanish or French; fluency shapes how smoothly they communicate, not what they know. LLMs work differently. They learn by predicting the next token from the data they see most, so when English dominates training, the patterns—and the medical "knowledge" riding inside them—are strongest in English. In lower-resource languages the patterns are thinner, and the model's apparent reasoning can fall apart even when the question contains everything it needs.
Then we take on the popular fix: machine translation. It sounds straightforward until you look at where it actually breaks—numbers, temporal qualifiers, negation, and culture-bound idioms. "Once a day" becoming "eleven times a day" is not a harmless glitch. We also unpack how common translation metrics can reward surface-level word overlap while missing exactly the meaning errors that matter most at the bedside.
For anyone building or using clinical AI, the takeaway is hard to dodge: if we want medical AI we can trust, multilingual competence can't be an afterthought. A system that's unsafe outside English shouldn't be called general medical intelligence.
References:
When medical AI fails outside English
Li et al.
BMJ Digital Health & AI (2026)
Credits:
Theme music: Nowhere Land, Kevin MacLeod (incompetech.com)
Licensed under Creative Commons: By Attribution 4.0
https://creativecommons.org/licenses/by/4.0/
A Shocking Language Gap
SPEAKER_01GPT 4 scores 90% on a medical exam in English. Same exam in French, only 55%. In contrast, a doctor who speaks both languages doesn't forget medicine in one of them. So what does the AI actually know?
SPEAKER_00Hello and welcome back to Code and Cure, the podcast where we discuss decoding health and the age of AI. And it's episode 52. Um, that's one year. My name is Vasant Sarathi. I'm an AI researcher and cognitive scientist, and I'm here with Laura Hagopian.
SPEAKER_01I'm an emergency medicine physician, and I'm excited that we've made it a full year.
SPEAKER_00Yeah, it's been fun. It's been fun. We've been exploring a lot of interesting topics. People have been um, you know, sending us ideas, which has been great. Um please keep them coming. Um, and we will continue to bring the worlds of health and AI together and find those you know gaps, find the promising ideas.
SPEAKER_01Um find the cool stories and um, you know, have a healthy dose of skepticism along the way. That's a little bit about what this episode is.
SPEAKER_00Yes, yes. This one is very skeptical for sure.
SPEAKER_01Yeah, it's interesting. So we we pulled up this article that is literally titled When Medical AI Fails Outside English, which I mean, it makes you nervous just to read the title of it, right? And when you think about it, I mean, a lot of medical information is concentrated in English in a couple other languages, it's not in every language that's out there. And when you think about how AI or and LLMs are trained, a lot of it is based on the data sets
Humans Keep Knowledge Across Languages
SPEAKER_01that they have, right? Available to them. And so if there's a limited data set in a certain language, even in French, the from the that that example, right? Then it doesn't do as well.
SPEAKER_00Yes, yes.
SPEAKER_01And so it's fast, what's fascinating to me about this is that when I think about my own medical knowledge, like and you're bilingual. Well, I mean, I speak some Spanish. I'm not sure I would I would totally label myself as bilingual, but I don't get like stupider in Spanish, right? Like my medical knowledge doesn't go away in Spanish. Maybe there are differences, like when I was doing work in Ecuador, I need to know about different things also, right? So it's not just language. Like if I'm traveling and doing medical work, I need to understand not just the different language, which is helpful and useful, but also what's happening in that culture or what are the typical diseases they see there that maybe uh I don't see here. Yes. Uh, for example, when I was in Ecuador, decent amount of altitude sickness because we were at altitude. I don't live at altitude here. Right. Uh, and so wasn't seeing that. You can imagine lots of examples like this where you you travel somewhere different and there are different disease processes like malaria in sub-Saharan Africa, right? Um, or chicken gunia in in like you know, southern southern parts of the United States. So there's different things that affect people in different areas, but there's also like a cultural context to it too, about how people talk about diseases, how people um explain what's going on with them, whether people bring a family member or not to help describe what's happening. There's a lot there that's not just language dependent.
SPEAKER_00Yes. And so there's that huge piece of sociocultural context that plays into your model of both the disease of the patient, of the types of medications available, the types of ways to deliver that medication, all of that stuff, right? And I I think that that is a whole layer of things that is hugely difficult. And, you know, one might argue, um, you know, a hard-nosed AI scientist might say, well, that's just context. You have to give the AI all of it. I mean, we all know, right? We can use Chat GPT, whatever, to upload PDFs of things and it's able to summarize it. And some might argue that, hey, why don't you just give it all the information to summarize it? And we've talked about this before. Um, and or why didn't why don't you just give it all the context it needs? Then it should be able to answer. But I think part of the problem is, of course, we, you know, we humans understand the context better. And providing all that context requires providing all the relevant information, which, you know, it is not that straightforward to do.
SPEAKER_01Well, I was gonna say it's it sounds easy, like snap of your fingers. Oh, here, like give it a better data set. Let's make sure it has a lot of context and it's highly curated, has all the information you'd want in it. Now you have to multiply that over all the languages in the world. Like that's just so much data to keep up with, to generate, to make sure is like a good data set. Like, who's maintaining all of that?
SPEAKER_00Well, yeah. So there's a whole practical issue, but I think there's even something even more fundamental, which you kind of hit at with your hook, which is these um language models did pretty well in English on these medical exams, which these medical exam uh benchmark data sets are set up so that all the right context is given. Yeah, there's no issue about context, all the right context is given. And um so we've taken that off the table, but it still did differently much worse on French than English. And to be clear, uh they have not done, they didn't do that necessarily with other even more rare languages.
SPEAKER_01Oh, they did, yeah, yeah. Oh, they actually outline it. It's like there's another example here where they use Deep Sea Gar One, which did 79.5% uh on these multiple choice questions in English, and then they they took rarer languages and it did in the 60 to 70 percent range in Zulu, Wolof, and Yoruba. Yeah, yeah.
SPEAKER_00So I mean this is a dramatic drop in performance um across these languages, and this is what context was given, where there's no reason to get it wrong, and this is where the weird piece comes in. Why is it that these AI systems are so smart in English, but completely dumb in other languages? That's kind of weird, and this brings back the issue about how they actually do this or reason about this, and and I want to stress how different this is from humans, right?
SPEAKER_01Right, because if you think about me moving to a new place and practicing, sure, I'm gonna have to learn some of the local things that are happening and some of the cultural stuff, but like my general medical knowledge isn't just like gone because I'm suddenly speaking Spanish or suddenly speaking French, like that doesn't go away. So how why is it that it like is is gone? Well, then it's worth it. Is it gone? Like what does it even mean? Like, does it lose knowledge? Like our brains don't like lose knowledge. So yeah, that's the part that I'm having trouble like wrapping my mind around is like, why does it just why does it do so much worse?
SPEAKER_00Yeah.
Why LLMs Depend On Training Data
SPEAKER_00And I think part of it is to appreciate the fact that large language models are different in their cognition from humans, very different. And and you can get an inkling of that just from the way they're trained, right? We don't learn things about the world by reading trillions of words on Reddit. That is not how we learn about the world.
SPEAKER_01I mean, I learned something, but definitely don't read trillions of words.
SPEAKER_00No, but I'm I'm I'm oversimplifying this. I know. Obviously, there is, you know, humans have evolved to this point, and knowledge has um sort of over time been hard-coded into our DNA potentially that we're leveraging and using. So we don't have to learn everything. This is a nature versus nurture thing. We don't have to learn everything from scratch as soon as we're born. We come in with some knowledge, fine, you know, and there's some quote-unquote pre-training that's happened through evolution. Um, and you know, there might be some argument to be made that that that pre-training through evolution is equivalent to maybe reading trillions of words in Reddit, although that's questionable. Um, but my point here is that the way LLMs learn is very different from the way humans learn. And I want to stress this because if you really think about the way LLMs learn, it's they're taking in words and learning to predict the next word. Um, it's actually next token, but forget that for a second. We'll come back to that in a second. But think about words, sequence of words, think about a language as being a sequence of words, and there's some pattern to it, and they're learning that pattern. So if I had two sentences, for example, if my whole data set was just these two sentences, was show me my files, please, or show me my photos, please, then every time you see the word show, you know for certain that the next word is me. And if you see the word me, you know for certain the next word is my. But the as soon as you see show me my, there's a 50% chance that it's files or photos because you've seen those two examples in your data set. Because we're simplifying this and we're saying there's only two examples in the entire data set.
SPEAKER_01Right. Okay.
SPEAKER_00So there's some uncertainty about which word to pick, right? Now imagine you had maybe you had, I don't know, 60 copies of that first sentence files and 40 copies of the second sentence photos. Then what's going to happen is that when you read show me my, it's going to think that there's a 60% chance that the next word is files and a 40% chance that the next word is photos. So slightly higher. So it's going to guess files, right?
SPEAKER_01Yeah, okay, that makes sense.
SPEAKER_00So now scale that up to every single possible sequence of words and every single sentence that's been used on the internet, right? It is learning those patterns. It's looking at the words and saying, okay, what is my best guess given I have a dictionary of all possible next words. What is the most likely one? And what is, you know, the order of the most likely ones or whatever, right? That's what it's doing. What is remarkable about this whole AI revolution is that that simple process has uh given us capabilities that we didn't expect in these systems. So this whole business of solving math problems or medical problems or whatever is a sub in some sense a surprise for researchers because it is still trained on next word prediction, but somehow the patterns of use in language also encode information and knowledge. And so now if you think about it, if you had a data set that is all English, it's very good at predicting the next word in English in English because there's data to support that you know pattern or whatever. So it knows statistically what the next word is going to be, but it's also coincidental because it's so much data, there is knowledge that's sort of implicit in all of that language use. And it's encoded that knowledge too, which is why it's doing well in these medical exams and these medical tests. Whereas if you pick up more rare language, like French or French is not that rare, but you can pick even rarer languages and there's less data and it's predicting the next word in that language.
SPEAKER_01So it's not like in English.
SPEAKER_00It doesn't have the knowledge. It it does have the knowledge when it comes to English, but that's not the pattern it's reading off of. I mean, there is a pattern in English, but there isn't a pattern in French or uh or Arabic or whatever else language that might be uh needed in that case. So it does it in some sense, you know, it doesn't have that knowledge in that language. And this is why it's weird, right? It's not like how humans think.
SPEAKER_01Right. Like my knowledge doesn't go away when I, you know, some speak in Spanish, but this the knowledge here is like language dependent, essentially, is what you're saying. Yeah.
SPEAKER_00And so then the question is, you know, yeah, sorry, go ahead.
SPEAKER_01No, I was I was just gonna say that like the data sets that are provided are what it goes off of. So it doesn't have a lot of data in a certain language. It's not like it's translating from English to sort of figure it out.
SPEAKER_00Yes. That is that like fair to say, yes. When there's a lot of examples of human usage of language, what comes with that is the particular words that the humans used, but also the underlying meaning and the usage of those terms and the underlying reasoning that went into it, which is not like in the words themselves, but they're implicit in the patterns. When there's more of those patterns, you get a better sense for the reasoning capabilities, you get a better sense for um if you're learning based on just numbers, you get a better sense for what those uh words encode. If you have lesser exam, fewer examples, it's just going to be worse off in generalizing um it's not this so-called knowledge, right?
Can Translation Fix Medical AI
SPEAKER_00So the question of what does an LLM really know is hard to say because it's possible that in theory it has the answer, correct answer, to the question being asked in French, because if perfectly translated to English, right? Maybe you could get at that knowledge.
SPEAKER_01Well, that's what I was gonna ask. And this paper sort of goes into this idea of like using machine translation as like a workaround. So, okay, like there's all this information available in English, these billions and billions of words, and maybe we don't have that for Zulu. And could we take all of the information that we have in in English and machine translate it to every other language out there, say Zulu or whatever, and then use that clinically, and now it has all the information.
SPEAKER_00Of course, that comes with its own. Well, that comes with a huge problem, which is machine translation models work differently, also. And you know, what's interesting, there's a little side fact that's really interesting, is that the world of LLMs and all of the models that came out, they came out of the world of machine translation, interestingly enough. Because machine translation has always been grappling with this challenge of given a sequence of words, predict another sequence of words. Given one sequence of symbols, predict another sequence of symbols. And the original um um the architecture of the L LLM, it's called the Transformer Network or Transformer, and that was originally introduced in 2017 as a solution for machine translation. And um many of the machine translation models today also use transformer architectures. They're not chatbots, they're not LLMs, they are machine translation-specific models, they're trained on machine uh on translation data, and they are supposedly very good. However, there's huge challenges with them because unlike an LLM, which is trained on all of Reddit or whatever, and therefore encodes some common sense reasoning, those machine translation systems don't, right? They're very focused on specific machine. The the training is input and output sentences, right? Um, not just like predict the next word, not just like large bodies of text.
SPEAKER_01And oftentimes those things, especially in a medical context, will get QA'd by like an actual human before it's released out because there can be mistakes and inaccuracies there.
SPEAKER_00Yes. And there are mistakes and accuracies, but there's also um, you know, fundamental challenges with translating from one language to another where words can be very similar,
Dangerous Translation Errors In Care
SPEAKER_00you know. Like we talked about this before the podcast started, was you know, this notion of the word once, right, in Spanish.
SPEAKER_01Yeah, well, I mean, well, so so once in Spanish, O-N-C-E, yeah, right, is a 11. Whereas O N C E once in English is one.
SPEAKER_00So there was actually a interestingly enough, there's a there was a 2010 uh study that found that, you know, they had a Spanish generated, uh sort of computer generated Spanish pharmacy labels, and they found that they included once a day, it often rendered it as 11 times a day. And there was actually a a documented case of a man taking 11 blood pressure pills in a day based on those instructions.
SPEAKER_01I mean, that could like be really dangerous. But and this is the thing is that you know, sometimes you can shrug off a little mistake, but that's not a little mistake when it comes to medications and a lot of the clinical context that you're doing this in. Yes. This can be the difference between life, life and death. I mean, that person could have ended up in the hospital because their blood pressure was so low or worse, right?
SPEAKER_00Yes, and these machine translation challenges can be super subtle looking, but massively important. So, like one class of machine translation problems is what they call temporal qualifiers, which is all about time. And, you know, you could have a sentence that says, take this for three days, uh, or you could have a sentence that says, take this in three days.
SPEAKER_01Oh, like so different. Those are different completely different things.
SPEAKER_00Or return if symptoms worsen, or return if symptoms, right? So completely different meanings, right? These they carry completely different clinical instructions, and these are the kinds of issues that happen. So that's one class of issues that can happen with machine translation. Another class of issues is that idioms sometimes don't map. So, like patients come in describing things like heat in the body or gastrising or my liver hurts. Uh, these are very common idioms used that are used, but when you literally translate them, they don't make any sense. And so people tend to drop them or translation systems might drop them. Um, and you have negation, which is another whole thing that I can talk about for hours, which is, you know, if I said hold the warfarin, right?
SPEAKER_01Yeah, that means don't take that. That means stop taking it. Stop taking the medication. You don't want someone to have you know blood that's too thin.
SPEAKER_00But it can translate sometimes in these systems as failures to keep taking, right? Which is the opposite of don't take. Like hold on to it, hold on to it, right? Instead of hold.
SPEAKER_01Yeah, interesting. Yeah. So I mean bad.
SPEAKER_00So these seem like tiny little random glitches, but the problem is they often are about negation, about numbers, about uh culture-bound idioms. Um, and they they can they can be hugely problematic.
SPEAKER_01Um in the clinical context for sure.
Scaling Risk And Bad Benchmarks
SPEAKER_01And then you think about hey, let's, you know, when we when you're talking AI, you're always talking scale. Like, oh, we can use machine translation and we can scale it. So now you're not doing it for one patient or five patients, you do it for thousands of patients. Yes. And if you have errors like this, you're gonna be causing a potentially a lot of issues.
SPEAKER_00Yeah, and these and you would think that machine translation systems are trained to catch these errors, but the issue is the most common training metric they use is something called the blue score, which is all about overlap of the sequences. So, like if I produce an answer, a translation answer, um, but you know the right answer, you compare them and you see which how much overlap there is between the words. And if it's high, then you know you've got a good translation. Now you can already see the problem with that because the examples I just gave you had like one word off. Yeah, exactly. So they're not gonna catch that, and it's gonna be incorrectly marked as you know, reasonably good enough. Right, good enough. And I think that's that's a huge problem. So, you know, on the one hand, we have LLM systems that have uh have have not enough data in the appropriate languages. We have no reliable means to translate um from one language to another, and that's still an open problem.
SPEAKER_01And and beyond that, I would say too, it's not just like this one-for-one translation. It's also about hey, what's what are the local norms or what is the cultural context that's happening here? Because that has to be part of it. A language isn't just one word for another. It's it's it's like when you're treating someone in a different location, in a different culture, in a different context, it's not just language that gets brought in. It's not just language that's different, right?
SPEAKER_00Yeah, it also seems like a lot of the biomedical literature is Western centric. I don't know if that matters, but it's a qu I guess it's a question to you, but like that's could potentially another bias that that can be introduced.
SPEAKER_01Absolutely, for sure. I mean, I think it's like we we've talked about how when you train on whatever data set, that's the data set that it's got in its back pocket, right? And if it's missing context that is culturally appropriate, or if it has biased data, then that's what it's gonna relay back.
SPEAKER_00Yeah, right?
SPEAKER_01That's the knowledge that it has. And so I guess the question then becomes like, well, what else what else can you do in this type of situation to try to improve on this? Because obviously it's like not equitable to only have English-centric LLMs that only work in one language reliably, or a couple languages that are, you know, you could see this actually worsening gaps. It it's like, well, you kind of need this more in resource poor settings, even, and they're not able to access it, or the information that it gives back sounds good, but isn't correct. That's it's actually makes it more dangerous, right?
SPEAKER_00Right, right, right, right, right. I mean, that's the this is the that's an uh that's the challenge, right? That's where you need it the most, and that's where it's the least effective. Um, and that's a huge problem. And and frankly, you know, people talk about how these models are improving and they're filling gaps, which is true.
Building Safer Multilingual Medical AI
SPEAKER_00They are getting better, they are addressing these questions, they are um challenging them, but there's some really core stuff that's you know that differentiates human reasoning from LLM reasoning, and we have just identified one of them, which is we are language invariant, right? We don't, we are not influenced.
SPEAKER_01We don't get dumber when we are in using a different language.
SPEAKER_00Right, right. We may have trouble communicating because our language fluency might not be.
SPEAKER_01Well, for sure. Yeah, but our like knowledge doesn't just like vanish.
SPEAKER_00Right, right, right.
SPEAKER_01And I think it it's gonna take more, obviously, than machine translation to get to a place where we can have better medical intelligence in other languages, right? It's like you need data sets that have local context that you know, physicians or other providers. Have worked to curate where you can get, you know, just really high quality data that is culturally appropriate, that is contextualized, um, to that location where they use terminology that's common for the local people, right? Right. You can't expect just translating from English to another language is going to do that. It's not. Right. It's like you need to have it be sort of fine-tuned and adapted to the area, to the area that you're that you're looking at, that you're looking to use it in.
SPEAKER_00Yeah. And from a scientific, sort of AI-centric perspective, um, I think we have to rethink our measurement, our way we measure things, and not sort of there's all these reports about AI acing medical exams. Well, there is a there's the other side to this, which is what we just talked about. And I think we need to be cognizant of that. And we need to think about ways to make the AI's knowledge language invariant. We need to think about ways to improve the quality of measurement so we're not using scores that are not incentivizing the right thing.
SPEAKER_01Yeah. I it's very interesting because you you almost have to ask different questions when you talk about AI. It's like in a human, you wouldn't be like, oh, you know, do you like, does your medical knowledge go away in a different language? You wouldn't even think to ask that question. Like, sure, you might not be, like you said, as fluent or your, you know, your language skills might not be as good, but like your medical knowledge doesn't just like go away. Yeah. Right. But here, so you know, these people have said, hey, we need to ask this question. And there's all this hype about how it's doing, yeah, just as well as doctors or better than doctors. Well, here it's like maybe not, especially in other languages. And I think um there's been a lot of hype on how we do we do we have general medical intelligence now. And the answer from this is like, no, no, we don't we absolutely do not, because it can't do it in other languages, right? Yeah, and and you can't have general medical intelligence if it doesn't work in French or in Zulu or in Wolof, you can't. Yeah, it's not possible. And so this sort of real world evaluation, this the the training data, the curation, um, the region specific editions, those need to be part of the next steps here if we want to head in that direction. But we're absolutely not there yet.
SPEAKER_00Yeah, I agree. And ling what I think what we're also saying is linguistic competence, which is language and understanding and so on uh in different languages, uh it can't just be an afterthought. You can't just like slap on a machine translation thing. That's not what this is about, right? So, like that's I think it needs to be considered upfront and thought about.
SPEAKER_01Absolutely. And we we could never have general medical intelligence until these types of things are sorted out. And it's not just about language, right? It's about culture, it's about local terminology, it's about practice norms, it's about different disease processes. All of those things have to be taken into account in order to move in that direction.
SPEAKER_00Yeah, I agree.
One Year In And Closing
SPEAKER_01All right. Well, thank you for joining us for our wow first year of Code and Cure.
SPEAKER_00Yeah.
SPEAKER_01And uh, we hope to see you again soon.
SPEAKER_00Thank you for joining us.