Starved by scarcity: A three scarcities framework and decision roadmap for medical named entity recognition in low-resource languages
Named entity recognition (NER) is a foundational task in medical natural language processing (NLP), enabling the extraction of clinically relevant information from unstructured text and supporting downstream applications such as information retrieval and clinical decision support. While pre-trained language models such as BioBERT, ClinicalBERT, and PubMedBERT have achieved state-of-the-art performance for English medical NER, progress in low-resource languages remains fragmented and uneven, even though healthcare systems worldwide generate vast amounts of textual data whose value depends on the availability of language technologies. Existing surveys provide valuable overviews but largely treat low-resource settings as a single category, overlooking important differences in the constraints practitioners face. To address this gap, we conducted a systematic review of 46 studies on medical NER in low-resource languages published since 2020. Based on the synthesized evidence, we introduce the three scarcities framework, which conceptualizes low-resource medical NLP as the interaction among data, model, and infrastructure scarcities. We also propose a practitioner-oriented decision roadmap that maps appropriate modeling and data augmentation strategies to different resource configurations. Our review develops a structured taxonomy of approaches, including cross-lingual transfer, in-domain pre-training, annotation projection, back-translation, and large language model-based synthetic data generation, and shows that the effectiveness of techniques depends less on the method itself than on the underlying scarcity profile. We also identify persistent weaknesses in evaluation practices, including inconsistent reporting and limited use of statistical significance testing.

- Demner-Fushman D, Chapman WW, McDonald CJ. What can natural language processing do for clinical decision support? J Biomed Inform. 2009;42(5):760-772. doi: 10.1016/j.jbi.2009.08.007
- Holzinger A, Schantl J, Schroettner M, Seifert C, Verspoor K. Biomedical text mining: state-of-the-art, open problems and future challenges. In: Lecture Notes in Computer Science. Springer Berlin Heidelberg; 2014:271-300. doi: 10.1007/978-3-662-43968-5_16
- Rivera-Zavala R, Martinez P. The impact of pretrained language models on negation and speculation detection in cross-lingual medical text: comparative study. JMIR Med Inform. 2020;8(12):e18953. doi: 10.2196/18953
- Zhou B, Yang G, Shi Z, Ma S. Natural language processing for smart healthcare. IEEE Rev Biomed Eng. 2024;17:4-18. doi: 10.1109/RBME.2022.3210270
- Kundeti SR, Vijayananda J, Mujjiga S, Kalyan M. Clinical named entity recognition: Challenges and opportunities. In: 2016 IEEE International Conference on Big Data (Big Data). IEEE; 2016:1937-1945. doi: 10.1109/bigdata.2016.7840814
- Yadav V, Bethard S. A survey on recent advances in named entity recognition from deep learning models. In: Proceedings of the 27th International Conference on Computational Linguistics. Association for Computational Linguistics; 2018:2145-2158. Accessed January 31, 2026. https://aclanthology.org/C18-1182/
- Lee J, Yoon W, Kim S, et al. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics. 2020;36(4):1234-1240. doi: 10.1093/bioinformatics/btz682
- Alsentzer E, Murphy JR, Boag W, et al. Publicly Available Clinical BERT Embeddings. arXiv. Preprint posted online 2019. doi: 10.48550/arXiv.1904.03323
- Beltagy I, Lo K, Cohan A. SciBERT: A Pretrained Language Model for Scientific Text. arXiv. Preprint posted online 2019. doi: 10.48550/arXiv.1903.10676
- Gu Y, Tinn R, Cheng H, et al. Domain-specific language model pretraining for biomedical natural language processing. ACM Trans Comput Healthc. 2022;3(1):1-23. doi: 10.1145/3458754
- Peng Y, Yan S, Lu Z. Transfer learning in biomedical natural language processing: an evaluation of BERT and ELMo on ten benchmarking datasets. In: Proceedings of the 18th BioNLP Workshop and Shared Task. Association for Computational Linguistics; 2019:58-65. doi: 10.18653/v1/W19-5006
- Durango MC, Torres-Silva EA, Orozco-Duque A. Named entity recognition in electronic health records: a methodological review. Healthc Inform Res. 2023;29(4):286-300. doi: 10.4258/hir.2023.29.4.286
- Fraile Navarro D, Ijaz K, Rezazadegan D, et al. Clinical named entity recognition and relation extraction using natural language processing of medical free text: a systematic review. Int J Med Inform. 2023;177:105122. doi: 10.1016/j.ijmedinf.2023.105122
- Jehangir B, Radhakrishnan S, Agarwal R. A survey on named entity recognition—datasets, tools, and methodologies. Nat Lang Process J. 2023;3:100017. doi: 10.1016/j.nlp.2023.100017
- Luo X, Deng Z, Yang B, Luo M. Y. Pre-trained language models in medicine: a survey. Artif Intell Med. 2024;154:102904. doi: 10.1016/j.artmed.2024.102904
- Wang X, Sanders HM, Liu Y, et al. ChatGPT: promise and challenges for deployment in low-and middle-income countries. Lancet Reg Health West Pac. 2023;41:100905. doi: 10.1016/j.lanwpc.2023.100905
- Shaitarova A, Zaghir J, Lavelli A, Krauthammer M, Rinaldi F. Exploring the latest highlights in medical natural language processing across multiple languages: a survey. Yearb Med Inform. 2023;32(01):230-243. doi: 10.1055/s-0043-1768726
- Schwartz AS, Hearst MA. A simple algorithm for identifying abbreviation definitions in biomedical text. In: Biocomputing 2003. World Scientific; 2002:451-462. doi: 10.1142/9789812776303_0042
- Chapman WW, Bridewell W, Hanbury P, Cooper GF, Buchanan BG. A simple algorithm for identifying negated findings and diseases in discharge summaries. J Biomed Inform. 2001;34(5):301-310. doi: 10.1006/jbin.2001.1029
- Leaman R, Khare R, Lu Z. Challenges in clinical natural language processing for automated disorder normalization. J Biomed Inform. 2015;57:28-37. doi: 10.1016/j.jbi.2015.07.010
- Lee EB, Heo GE, Choi CM, Song M. MLM-based typographical error correction of unstructured medical texts for named entity recognition. BMC Bioinform. 2022;23(1):486. doi: 10.1186/s12859-022-05035-9
- Boudjellal N, Zhang H, Khan A, et al. ABioNER: a BERT-based model for Arabic biomedical named-entity recognition. Complexity. 2021;2021(1):6633213. doi: 10.1155/2021/6633213
- Perera N, Dehmer M, Emmert-Streib F. Named entity recognition and relation detection for biomedical information extraction. Front Cell Dev Biol. 2020;8:673. doi: 10.3389/fcell.2020.00673
- Berg H, Henriksson A, Dalianis H. The impact of de-identification on downstream named entity recognition in clinical text. In: Proceedings of the 11th International Workshop on Health Text Mining and Information Analysis. Association for Computational Linguistics; 2020:1-11. doi: 10.18653/v1/2020.louhi-1.1
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021:n71. doi: 10.1136/bmj.n71
- Rivera-Zavala RM, Martínez P. Analyzing transfer learning impact in biomedical cross-lingual named entity recognition and normalization. BMC Bioinformatics. 2021;22(S1):601. doi: 10.1186/s12859-021-04247-9
- Amin S, Goldstein NP, Wixted MK, García-Rudolph A, Martínez-Costa C, Neumann G. Few-Shot Cross-lingual Transfer for Coarse-grained De-identification of Code-Mixed Clinical Texts. arXiv. Preprint posted online 2022. doi: 10.48550/arXiv.2204.04775
- Lancheros BS, Corpas Pastor G, Mitkov R. Data augmentation and transfer learning for cross-lingual named entity recognition in the biomedical domain. Lang Resour Eval. 2025;59(2):665-684. doi: 10.1007/s10579-024-09738-8
- Vakili T, Henriksson A, Dalianis H. Data-Constrained Synthesis of Training Data for De-Identification. arXiv. Preprint posted online 2025. doi: 10.48550/arXiv.2502.14677
- Gallego F, López-García G, Gasco-Sánchez L, Krallinger M, Veredas FJ. ClinLinker: Medical Entity Linking of Clinical Concept Mentions in Spanish. arXiv. Preprint posted online 2024. doi: 10.48550/arXiv.2404.06367
- López-García G, Jerez JM, Ribelles N, Alba E, Veredas FJ. Transformers for clinical coding in Spanish. IEEE Access. 2021;9:72387-72397. doi: 10.1109/ACCESS.2021.3080085
- Villena F, Bravo-Marquez F, Dunstan J. NLP modeling recommendations for restricted data availability in clinical settings. BMC Med Inform Decis Mak. 2025;25(1):116. doi: 10.1186/s12911-025-02948-2
- Sun C, Yang Z, Wang L, Zhang Y, Lin H, Wang J. Deep learning with language models improves named entity recognition for PharmaCoNER. BMC Bioinform. 2021;22(S1). doi: 10.1186/s12859-021-04260-y
- Gallego F, Veredas FJ. Recognition and normalization of multilingual symptom entities using in-domain-adapted BERT models and classification layers. Database. 2024;2024:baae087. doi: 10.1093/database/baae087
- Akhtyamova L, Martinez P, Verspoor K, Cardiff J. Testing contextualized word embeddings to improve NER in Spanish clinical case narratives. IEEE Access. 2020;8:164717-164726. doi: 10.1109/ACCESS.2020.3018688
- Sasikumar N, Mantri KSI. Transfer learning for low-resource clinical named entity recognition. In: Proceedings of the 5th Clinical Natural Language Processing Workshop. Association for Computational Linguistics; 2023:514-518. doi: 10.18653/v1/2023.clinicalnlp-1.53
- Fabregat H, Martinez-Romo J, Araujo L. Understanding and improving disability identification in medical documents. IEEE Access. 2020;8:155399-155408. doi: 10.1109/ACCESS.2020.3019178
- Mullick A, Mondal I, Ray S, Raghav R, Chaitanya G, Goyal P. Intent identification and entity extraction for healthcare queries in Indic languages. In: Findings of the Association for Computational Linguistics: EACL 2023. Association for Computational Linguistics; 2023:1870-1881. doi: 10.18653/v1/2023.findings-eacl.140
- Yada S, Nakamura Y, Wakamiya S, Aramaki E. Real-MedNLP: overview of REAL document-based MEDical natural language processing task. In: Proceedings of the 16th NTCIR Conference on Evaluation of Information Access Technologies. National Institute of Informatics; 2022:285-296.
- Phan U, Nguyen N. Simple semantic-based data augmentation for named entity recognition in biomedical texts. In: Proceedings of the 21st Workshop on Biomedical Language Processing. Association for Computational Linguistics; 2022:123-129. doi: 10.18653/v1/2022.bionlp-1.12
- Miftahutdinov Z, Alimova I, Tutubalina E. On biomedical named entity recognition: experiments in interlingual transfer for clinical and social media texts. In: Lecture Notes in Computer Science. Springer International Publishing; 2020:281-288. doi: 10.1007/978-3-030-45442-5_35
- Pengpun P, Tiankanon K, Chinkamol A, et al. On creating an English-Thai code-switched machine translation in medical domain. In: Findings of the Association for Computational Linguistics: EMNLP 2024. Association for Computational Linguistics; 2024:6055-6073. doi: 10.18653/v1/2024.findings-emnlp.351
- Lerner I, Paris N, Tannier X. Terminologies augmented recurrent neural network model for clinical named entity recognition. J Biomed Inform. 2020;102:103356. doi: 10.1016/j.jbi.2019.103356
- Idrissi-Yaghir A, Dada A, Schäfer H, et al. Comprehensive Study on German Language Models for Clinical and Biomedical Text Understanding. arXiv. Preprint posted online 2024. doi: 10.48550/arXiv.2404.05694
- Fontaine X, Gaschi F, Rastin P, Toussaint Y. Multilingual Clinical NER: Translation or Cross-lingual Transfer? arXiv. Preprint posted online 2023. doi: 10.48550/arXiv.2306.04384
- Frei J, Kramer F. Annotated dataset creation through large language models for non-English medical NLP. J Biomed Inform. 2023;145:104478. doi: 10.1016/j.jbi.2023.104478
- Schäfer H, Idrissi-Yaghir A, Horn P, Friedrich C. Cross-language transfer of high-quality annotations: combining neural machine translation with cross-linguistic span alignment to apply NER to clinical texts in a low-resource language. In: Proceedings of the 4th Clinical Natural Language Processing Workshop. Association for Computational Linguistics; 2022:53-62. doi: 10.18653/v1/2022.clinicalnlp-1.6
- Frei J, Kramer F. GERNERMED: an open German medical NER model. Softw Impacts. 2022;11:100212. doi: 10.1016/j.simpa.2021.100212
- Jantscher M, Gunzer F, Kern R, Hassler E, Tschauner S, Reishofer G. Information extraction from German radiological reports for general clinical text and language understanding. Sci Rep. 2023;13(1). doi: 10.1038/s41598-023-29323-3
- Abdulnazar A, Roller R, Schulz S, Kreuzthaler M. Large language models for clinical text cleansing enhance medical concept normalization. IEEE Access. 2024;12:147981-147990. doi: 10.1109/ACCESS.2024.3472500
- Zaghir J, Bjelogrlic M, Goldman JP, Aananou S, Gaudet-Blavignac C, Lovis C. FRASIMED: a clinical French annotated resource produced through crosslingual BERT-based annotation projection. In: Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). ELRA and ICCL; 2024:7450-7460. doi: 10.63317/4z8mxey3dmxp
- Copara J, Knafou J, Naderi N, Moro C, Ruch P, Teodoro D. Contextualized French language models for biomedical named entity recognition. In: Actes de la 6e Conférence Conjointe JEP-TALN-RÉCITAL, Atelier DÉfi Fouille de Textes. ATALA et AFCP; 2020:36-48. Accessed January 31, 2026. https://aclanthology.org/2020.jeptalnrecital-deft.4/
- Grabar N, Dalloux C, Claveau V. CAS: corpus of clinical cases in French. J Biomed Semantics. 2020;11(1):7. doi: 10.1186/s13326-020-00225-x
- Minh N, Tran VH, Hoang V, Ta HD, Bui TH, Truong SQH. ViHealthBERT: pre-trained language models for Vietnamese in health text mining. In: Proceedings of the 13th Language Resources and Evaluation Conference. European Language Resources Association; 2022:328-337. Accessed January 31, 2026. https://aclanthology.org/2022.lrec-1.35/
- Byun S, Hong J, Park S, et al. Korean Bio-Medical Corpus (KBMC) for Medical Named Entity Recognition. arXiv. Published online 2024. doi: 10.48550/arXiv.2403.16158
- Wang Y, Lin M, Hu Q, Bao L, Bai S, Li Y. Large and small models for collaborative cross-lingual data augmentation in entity relationship extraction for low-resource languages. J King Saud Univ Comput Inf Sci. 2025;37(4):56. doi: 10.1007/s44443-025-00055-w
- Sazzed S. BanglaBioMed: a biomedical named-entity annotated corpus for Bangla (Bengali). In: Proceedings of the 21st Workshop on Biomedical Language Processing. Association for Computational Linguistics; 2022:323-329. doi: 10.18653/v1/2022.bionlp-1.31
- Lu F, Qi Q, Qin H. Joint extraction of Uygur medicine knowledge with edge computing. Tsinghua Sci Technol. 2025;30(2):782-795. doi: 10.26599/TST.2024.9010006
- Kim YM, Lee TH. Korean clinical entity recognition from diagnosis text using BERT. BMC Med Inform Decis Mak. 2020;20(S7):242. doi: 10.1186/s12911-020-01241-8
- Sousa H, Pasquali A, Jorge A, Santos CS, Lopes MA. A biomedical entity extraction pipeline for oncology health records in Portuguese. In: Proceedings of the 38th ACM/ SIGAPP Symposium on Applied Computing. ACM; 2023:950-956. doi: 10.1145/3555776.3578577
- Schneider ETR, de Souza JVA, Knafou J, et al. BioBERTpt -a Portuguese neural language model for clinical named entity recognition. In: Proceedings of the 3rd Clinical Natural Language Processing Workshop. Association for Computational Linguistics; 2020:65-72. doi: 10.18653/v1/2020.clinicalnlp-1.7
- Šuvalov H, Laur S, Kolde R. Information extraction from medical texts with BERT using human-in-the-loop labeling. In: Studies in Health Technology and Informatics. IOS Press; 2023. doi: 10.3233/SHTI230281
- Šuvalov H, Lepson M, Kukk V, et al. Using Synthetic Health Care Data to Leverage Large Language Models for Named Entity Recognition: Development and Validation Study. J Med Internet Res. 2025;27:e66279. doi: 10.2196/66279
- Coelho Da Silva T, Fernandes De Macêdo J, Magalhães R. Tracking the evolution of Covid-19 symptoms through clinical conversations. In: Proceedings of the 5th Clinical Natural Language Processing Workshop. Association for Computational Linguistics; 2023:41-47. doi: 10.18653/v1/2023.clinicalnlp-1.6
- Vakili T, Hansson M, Henriksson A. SweClinEval: a benchmark for Swedish clinical natural language processing. In: Proceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies (NoDaLiDa/Baltic-HLT 2025). University of Tartu Library; 2025:767-775. Accessed January 31, 2026. https://aclanthology.org/2025. nodalida-1.76/
- Catelli R, Gargiulo F, Casola V, De Pietro G, Fujita H, Esposito M. A Novel COVID-19 Data Set and an Effective Deep Learning Approach for the De-Identification of Italian Medical Records. IEEE Access. 2021;9:19097-19110. doi: 10.1109/access.2021.3054479
- Bondarenko DA, Ferrod R, Di Caro L. Combining Contrastive Learning and Knowledge Graph Embeddings to develop medical word embeddings for the Italian language. arXiv. Preprint posted online 2022. doi: 10.48550/arXiv.2211.05035
- Postiglione M. Towards an Italian healthcare knowledge graph. In: Lecture Notes in Computer Science. Springer International Publishing; 2021:387-394. doi: 10.1007/978-3-030-89657-7_29
- Vassileva S, Todorova G, Ivanova K, et al. Automatic Transformation of Clinical Narratives into Structured Format. In: Proceedings of the Student Research Workshop Associated with RANLP 2021. INCOMA Ltd. Shoumen, Bulgaria; 2021:219-227. doi: 10.26615/issn.2603-2821.2021_030
