AccScience Publishing / AIH / Online First / DOI: 10.36922/AIH026260068
Cite this article
6
Download
83
Views
Related Info Links
More by Authors Links
Journal Browser
Volume | Year
Issue
Search
News and Announcements
View All
REVIEW ARTICLE

Starved by scarcity: A three scarcities framework and decision roadmap for medical named entity recognition in low-resource languages

Elira Hoxha1 Polina Çeço1*
Show Less
1 Department of Statistics and Applied Informatics, Faculty of Economy, University of Tirana, Tirana, Albania
Received: 24 June 2026 | Revised: 10 July 2026 | Accepted: 13 July 2026 | Published online: 23 July 2026
© 2026 by the Author(s). This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution 4.0 International License ( https://creativecommons.org/licenses/by/4.0/ )
Abstract

Named entity recognition (NER) is a foundational task in medical natural language processing (NLP), enabling the extraction of clinically relevant information from unstructured text and supporting downstream applications such as information retrieval and clinical decision support. While pre-trained language models such as BioBERT, ClinicalBERT, and PubMedBERT have achieved state-of-the-art performance for English medical NER, progress in low-resource languages remains fragmented and uneven, even though healthcare systems worldwide generate vast amounts of textual data whose value depends on the availability of language technologies. Existing surveys provide valuable overviews but largely treat low-resource settings as a single category, overlooking important differences in the constraints practitioners face. To address this gap, we conducted a systematic review of 46 studies on medical NER in low-resource languages published since 2020. Based on the synthesized evidence, we introduce the three scarcities framework, which conceptualizes low-resource medical NLP as the interaction among data, model, and infrastructure scarcities. We also propose a practitioner-oriented decision roadmap that maps appropriate modeling and data augmentation strategies to different resource configurations. Our review develops a structured taxonomy of approaches, including cross-lingual transfer, in-domain pre-training, annotation projection, back-translation, and large language model-based synthetic data generation, and shows that the effectiveness of techniques depends less on the method itself than on the underlying scarcity profile. We also identify persistent weaknesses in evaluation practices, including inconsistent reporting and limited use of statistical significance testing.

Graphical abstract
Keywords
Medical natural language processing
Named entity recognition
Low-resource languages
Large language models
Transformer models
Funding
None.
Conflict of interest
The authors declare they have no competing interests.
References
  1. Demner-Fushman D, Chapman WW, McDonald CJ. What can natural language processing do for clinical decision support? J Biomed Inform. 2009;42(5):760-772. doi: 10.1016/j.jbi.2009.08.007
  2. Holzinger A, Schantl J, Schroettner M, Seifert C, Verspoor K. Biomedical text mining: state-of-the-art, open problems and future challenges. In: Lecture Notes in Computer Science. Springer Berlin Heidelberg; 2014:271-300. doi: 10.1007/978-3-662-43968-5_16
  3. Rivera-Zavala R, Martinez P. The impact of pretrained language models on negation and speculation detection in cross-lingual medical text: comparative study. JMIR Med Inform. 2020;8(12):e18953. doi: 10.2196/18953
  4. Zhou B, Yang G, Shi Z, Ma S. Natural language processing for smart healthcare. IEEE Rev Biomed Eng. 2024;17:4-18. doi: 10.1109/RBME.2022.3210270
  5. Kundeti SR, Vijayananda J, Mujjiga S, Kalyan M. Clinical named entity recognition: Challenges and opportunities. In: 2016 IEEE International Conference on Big Data (Big Data). IEEE; 2016:1937-1945. doi: 10.1109/bigdata.2016.7840814
  6. Yadav V, Bethard S. A survey on recent advances in named entity recognition from deep learning models. In: Proceedings of the 27th International Conference on Computational Linguistics. Association for Computational Linguistics; 2018:2145-2158. Accessed January 31, 2026. https://aclanthology.org/C18-1182/
  7. Lee J, Yoon W, Kim S, et al. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics. 2020;36(4):1234-1240. doi: 10.1093/bioinformatics/btz682
  8. Alsentzer E, Murphy JR, Boag W, et al. Publicly Available Clinical BERT Embeddings. arXiv. Preprint posted online 2019. doi: 10.48550/arXiv.1904.03323
  9. Beltagy I, Lo K, Cohan A. SciBERT: A Pretrained Language Model for Scientific Text. arXiv. Preprint posted online 2019. doi: 10.48550/arXiv.1903.10676
  10. Gu Y, Tinn R, Cheng H, et al. Domain-specific language model pretraining for biomedical natural language processing. ACM Trans Comput Healthc. 2022;3(1):1-23. doi: 10.1145/3458754
  11. Peng Y, Yan S, Lu Z. Transfer learning in biomedical natural language processing: an evaluation of BERT and ELMo on ten benchmarking datasets. In: Proceedings of the 18th BioNLP Workshop and Shared Task. Association for Computational Linguistics; 2019:58-65. doi: 10.18653/v1/W19-5006
  12. Durango MC, Torres-Silva EA, Orozco-Duque A. Named entity recognition in electronic health records: a methodological review. Healthc Inform Res. 2023;29(4):286-300. doi: 10.4258/hir.2023.29.4.286
  13. Fraile Navarro D, Ijaz K, Rezazadegan D, et al. Clinical named entity recognition and relation extraction using natural language processing of medical free text: a systematic review. Int J Med Inform. 2023;177:105122. doi: 10.1016/j.ijmedinf.2023.105122
  14. Jehangir B, Radhakrishnan S, Agarwal R. A survey on named entity recognition—datasets, tools, and methodologies. Nat Lang Process J. 2023;3:100017. doi: 10.1016/j.nlp.2023.100017
  15. Luo X, Deng Z, Yang B, Luo M. Y. Pre-trained language models in medicine: a survey. Artif Intell Med. 2024;154:102904. doi: 10.1016/j.artmed.2024.102904
  16. Wang X, Sanders HM, Liu Y, et al. ChatGPT: promise and challenges for deployment in low-and middle-income countries. Lancet Reg Health West Pac. 2023;41:100905. doi: 10.1016/j.lanwpc.2023.100905
  17. Shaitarova A, Zaghir J, Lavelli A, Krauthammer M, Rinaldi F. Exploring the latest highlights in medical natural language processing across multiple languages: a survey. Yearb Med Inform. 2023;32(01):230-243. doi: 10.1055/s-0043-1768726
  18. Schwartz AS, Hearst MA. A simple algorithm for identifying abbreviation definitions in biomedical text. In: Biocomputing 2003. World Scientific; 2002:451-462. doi: 10.1142/9789812776303_0042
  19. Chapman WW, Bridewell W, Hanbury P, Cooper GF, Buchanan BG. A simple algorithm for identifying negated findings and diseases in discharge summaries. J Biomed Inform. 2001;34(5):301-310. doi: 10.1006/jbin.2001.1029
  20. Leaman R, Khare R, Lu Z. Challenges in clinical natural language processing for automated disorder normalization. J Biomed Inform. 2015;57:28-37. doi: 10.1016/j.jbi.2015.07.010
  21. Lee EB, Heo GE, Choi CM, Song M. MLM-based typographical error correction of unstructured medical texts for named entity recognition. BMC Bioinform. 2022;23(1):486. doi: 10.1186/s12859-022-05035-9
  22. Boudjellal N, Zhang H, Khan A, et al. ABioNER: a BERT-based model for Arabic biomedical named-entity recognition. Complexity. 2021;2021(1):6633213. doi: 10.1155/2021/6633213
  23. Perera N, Dehmer M, Emmert-Streib F. Named entity recognition and relation detection for biomedical information extraction. Front Cell Dev Biol. 2020;8:673. doi: 10.3389/fcell.2020.00673
  24. Berg H, Henriksson A, Dalianis H. The impact of de-identification on downstream named entity recognition in clinical text. In: Proceedings of the 11th International Workshop on Health Text Mining and Information Analysis. Association for Computational Linguistics; 2020:1-11. doi: 10.18653/v1/2020.louhi-1.1
  25. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021:n71. doi: 10.1136/bmj.n71
  26. Rivera-Zavala RM, Martínez P. Analyzing transfer learning impact in biomedical cross-lingual named entity recognition and normalization. BMC Bioinformatics. 2021;22(S1):601. doi: 10.1186/s12859-021-04247-9
  27. Amin S, Goldstein NP, Wixted MK, García-Rudolph A, Martínez-Costa C, Neumann G. Few-Shot Cross-lingual Transfer for Coarse-grained De-identification of Code-Mixed Clinical Texts. arXiv. Preprint posted online 2022. doi: 10.48550/arXiv.2204.04775
  28. Lancheros BS, Corpas Pastor G, Mitkov R. Data augmentation and transfer learning for cross-lingual named entity recognition in the biomedical domain. Lang Resour Eval. 2025;59(2):665-684. doi: 10.1007/s10579-024-09738-8
  29. Vakili T, Henriksson A, Dalianis H. Data-Constrained Synthesis of Training Data for De-Identification. arXiv. Preprint posted online 2025. doi: 10.48550/arXiv.2502.14677
  30. Gallego F, López-García G, Gasco-Sánchez L, Krallinger M, Veredas FJ. ClinLinker: Medical Entity Linking of Clinical Concept Mentions in Spanish. arXiv. Preprint posted online 2024. doi: 10.48550/arXiv.2404.06367
  31. López-García G, Jerez JM, Ribelles N, Alba E, Veredas FJ. Transformers for clinical coding in Spanish. IEEE Access. 2021;9:72387-72397. doi: 10.1109/ACCESS.2021.3080085
  32. Villena F, Bravo-Marquez F, Dunstan J. NLP modeling recommendations for restricted data availability in clinical settings. BMC Med Inform Decis Mak. 2025;25(1):116. doi: 10.1186/s12911-025-02948-2
  33. Sun C, Yang Z, Wang L, Zhang Y, Lin H, Wang J. Deep learning with language models improves named entity recognition for PharmaCoNER. BMC Bioinform. 2021;22(S1). doi: 10.1186/s12859-021-04260-y
  34. Gallego F, Veredas FJ. Recognition and normalization of multilingual symptom entities using in-domain-adapted BERT models and classification layers. Database. 2024;2024:baae087. doi: 10.1093/database/baae087
  35. Akhtyamova L, Martinez P, Verspoor K, Cardiff J. Testing contextualized word embeddings to improve NER in Spanish clinical case narratives. IEEE Access. 2020;8:164717-164726. doi: 10.1109/ACCESS.2020.3018688
  36. Sasikumar N, Mantri KSI. Transfer learning for low-resource clinical named entity recognition. In: Proceedings of the 5th Clinical Natural Language Processing Workshop. Association for Computational Linguistics; 2023:514-518. doi: 10.18653/v1/2023.clinicalnlp-1.53
  37. Fabregat H, Martinez-Romo J, Araujo L. Understanding and improving disability identification in medical documents. IEEE Access. 2020;8:155399-155408. doi: 10.1109/ACCESS.2020.3019178
  38. Mullick A, Mondal I, Ray S, Raghav R, Chaitanya G, Goyal P. Intent identification and entity extraction for healthcare queries in Indic languages. In: Findings of the Association for Computational Linguistics: EACL 2023. Association for Computational Linguistics; 2023:1870-1881. doi: 10.18653/v1/2023.findings-eacl.140
  39. Yada S, Nakamura Y, Wakamiya S, Aramaki E. Real-MedNLP: overview of REAL document-based MEDical natural language processing task. In: Proceedings of the 16th NTCIR Conference on Evaluation of Information Access Technologies. National Institute of Informatics; 2022:285-296.
  40. Phan U, Nguyen N. Simple semantic-based data augmentation for named entity recognition in biomedical texts. In: Proceedings of the 21st Workshop on Biomedical Language Processing. Association for Computational Linguistics; 2022:123-129. doi: 10.18653/v1/2022.bionlp-1.12
  41. Miftahutdinov Z, Alimova I, Tutubalina E. On biomedical named entity recognition: experiments in interlingual transfer for clinical and social media texts. In: Lecture Notes in Computer Science. Springer International Publishing; 2020:281-288. doi: 10.1007/978-3-030-45442-5_35
  42. Pengpun P, Tiankanon K, Chinkamol A, et al. On creating an English-Thai code-switched machine translation in medical domain. In: Findings of the Association for Computational Linguistics: EMNLP 2024. Association for Computational Linguistics; 2024:6055-6073. doi: 10.18653/v1/2024.findings-emnlp.351
  43. Lerner I, Paris N, Tannier X. Terminologies augmented recurrent neural network model for clinical named entity recognition. J Biomed Inform. 2020;102:103356. doi: 10.1016/j.jbi.2019.103356
  44. Idrissi-Yaghir A, Dada A, Schäfer H, et al. Comprehensive Study on German Language Models for Clinical and Biomedical Text Understanding. arXiv. Preprint posted online 2024. doi: 10.48550/arXiv.2404.05694
  45. Fontaine X, Gaschi F, Rastin P, Toussaint Y. Multilingual Clinical NER: Translation or Cross-lingual Transfer? arXiv. Preprint posted online 2023. doi: 10.48550/arXiv.2306.04384
  46. Frei J, Kramer F. Annotated dataset creation through large language models for non-English medical NLP. J Biomed Inform. 2023;145:104478. doi: 10.1016/j.jbi.2023.104478
  47. Schäfer H, Idrissi-Yaghir A, Horn P, Friedrich C. Cross-language transfer of high-quality annotations: combining neural machine translation with cross-linguistic span alignment to apply NER to clinical texts in a low-resource language. In: Proceedings of the 4th Clinical Natural Language Processing Workshop. Association for Computational Linguistics; 2022:53-62. doi: 10.18653/v1/2022.clinicalnlp-1.6
  48. Frei J, Kramer F. GERNERMED: an open German medical NER model. Softw Impacts. 2022;11:100212. doi: 10.1016/j.simpa.2021.100212
  49. Jantscher M, Gunzer F, Kern R, Hassler E, Tschauner S, Reishofer G. Information extraction from German radiological reports for general clinical text and language understanding. Sci Rep. 2023;13(1). doi: 10.1038/s41598-023-29323-3
  50. Abdulnazar A, Roller R, Schulz S, Kreuzthaler M. Large language models for clinical text cleansing enhance medical concept normalization. IEEE Access. 2024;12:147981-147990. doi: 10.1109/ACCESS.2024.3472500
  51. Zaghir J, Bjelogrlic M, Goldman JP, Aananou S, Gaudet-Blavignac C, Lovis C. FRASIMED: a clinical French annotated resource produced through crosslingual BERT-based annotation projection. In: Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). ELRA and ICCL; 2024:7450-7460. doi: 10.63317/4z8mxey3dmxp
  52. Copara J, Knafou J, Naderi N, Moro C, Ruch P, Teodoro D. Contextualized French language models for biomedical named entity recognition. In: Actes de la 6e Conférence Conjointe JEP-TALN-RÉCITAL, Atelier DÉfi Fouille de Textes. ATALA et AFCP; 2020:36-48. Accessed January 31, 2026. https://aclanthology.org/2020.jeptalnrecital-deft.4/
  53. Grabar N, Dalloux C, Claveau V. CAS: corpus of clinical cases in French. J Biomed Semantics. 2020;11(1):7. doi: 10.1186/s13326-020-00225-x
  54. Minh N, Tran VH, Hoang V, Ta HD, Bui TH, Truong SQH. ViHealthBERT: pre-trained language models for Vietnamese in health text mining. In: Proceedings of the 13th Language Resources and Evaluation Conference. European Language Resources Association; 2022:328-337. Accessed January 31, 2026. https://aclanthology.org/2022.lrec-1.35/
  55. Byun S, Hong J, Park S, et al. Korean Bio-Medical Corpus (KBMC) for Medical Named Entity Recognition. arXiv. Published online 2024. doi: 10.48550/arXiv.2403.16158
  56. Wang Y, Lin M, Hu Q, Bao L, Bai S, Li Y. Large and small models for collaborative cross-lingual data augmentation in entity relationship extraction for low-resource languages. J King Saud Univ Comput Inf Sci. 2025;37(4):56. doi: 10.1007/s44443-025-00055-w
  57. Sazzed S. BanglaBioMed: a biomedical named-entity annotated corpus for Bangla (Bengali). In: Proceedings of the 21st Workshop on Biomedical Language Processing. Association for Computational Linguistics; 2022:323-329. doi: 10.18653/v1/2022.bionlp-1.31
  58. Lu F, Qi Q, Qin H. Joint extraction of Uygur medicine knowledge with edge computing. Tsinghua Sci Technol. 2025;30(2):782-795. doi: 10.26599/TST.2024.9010006
  59. Kim YM, Lee TH. Korean clinical entity recognition from diagnosis text using BERT. BMC Med Inform Decis Mak. 2020;20(S7):242. doi: 10.1186/s12911-020-01241-8
  60. Sousa H, Pasquali A, Jorge A, Santos CS, Lopes MA. A biomedical entity extraction pipeline for oncology health records in Portuguese. In: Proceedings of the 38th ACM/ SIGAPP Symposium on Applied Computing. ACM; 2023:950-956. doi: 10.1145/3555776.3578577
  61. Schneider ETR, de Souza JVA, Knafou J, et al. BioBERTpt -a Portuguese neural language model for clinical named entity recognition. In: Proceedings of the 3rd Clinical Natural Language Processing Workshop. Association for Computational Linguistics; 2020:65-72. doi: 10.18653/v1/2020.clinicalnlp-1.7
  62. Šuvalov H, Laur S, Kolde R. Information extraction from medical texts with BERT using human-in-the-loop labeling. In: Studies in Health Technology and Informatics. IOS Press; 2023. doi: 10.3233/SHTI230281
  63. Šuvalov H, Lepson M, Kukk V, et al. Using Synthetic Health Care Data to Leverage Large Language Models for Named Entity Recognition: Development and Validation Study. J Med Internet Res. 2025;27:e66279. doi: 10.2196/66279
  64. Coelho Da Silva T, Fernandes De Macêdo J, Magalhães R. Tracking the evolution of Covid-19 symptoms through clinical conversations. In: Proceedings of the 5th Clinical Natural Language Processing Workshop. Association for Computational Linguistics; 2023:41-47. doi: 10.18653/v1/2023.clinicalnlp-1.6
  65. Vakili T, Hansson M, Henriksson A. SweClinEval: a benchmark for Swedish clinical natural language processing. In: Proceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies (NoDaLiDa/Baltic-HLT 2025). University of Tartu Library; 2025:767-775. Accessed January 31, 2026. https://aclanthology.org/2025. nodalida-1.76/
  66. Catelli R, Gargiulo F, Casola V, De Pietro G, Fujita H, Esposito M. A Novel COVID-19 Data Set and an Effective Deep Learning Approach for the De-Identification of Italian Medical Records. IEEE Access. 2021;9:19097-19110. doi: 10.1109/access.2021.3054479
  67. Bondarenko DA, Ferrod R, Di Caro L. Combining Contrastive Learning and Knowledge Graph Embeddings to develop medical word embeddings for the Italian language. arXiv. Preprint posted online 2022. doi: 10.48550/arXiv.2211.05035
  68. Postiglione M. Towards an Italian healthcare knowledge graph. In: Lecture Notes in Computer Science. Springer International Publishing; 2021:387-394. doi: 10.1007/978-3-030-89657-7_29
  69. Vassileva S, Todorova G, Ivanova K, et al. Automatic Transformation of Clinical Narratives into Structured Format. In: Proceedings of the Student Research Workshop Associated with RANLP 2021. INCOMA Ltd. Shoumen, Bulgaria; 2021:219-227. doi: 10.26615/issn.2603-2821.2021_030
Share
Back to top
Artificial Intelligence in Health, Electronic ISSN: 3029-2387 Print ISSN: 3041-0894, Published by AccScience Publishing