International Journal of Pharmaceutical and Phytopharmacological Research
ISSN (Print): 2250-1029
ISSN (Online): 2249-6084
Publish with eIJPPR Submission
2026   Volume 16   Issue 1

Foundation Models for Drug Discovery: A Systematic Review of Molecular Language, Representation Learning, Target Prediction, Generative Design, and Translational Risk
Download PDF


, ,
  1. Department of AI for Anti-Inflammatory Lectins and Glycosides, Faculty of Pharmacy, Federal University of Minas Gerais, Belo Horizonte, Brazil.
  2. Department of Computational Pharmacology for Autoimmune Diseases, Faculty of Pharmacy, University of Coimbra, Coimbra, Portugal.
Citation
Vancouver
Costa G, Ribeiro L, Alves R. Foundation Models for Drug Discovery: A Systematic Review of Molecular Language, Representation Learning, Target Prediction, Generative Design, and Translational Risk. Int J Pharm Phytopharmacol Res. 2026;16(1):35-46. https://doi.org/10.51847/bWGXsPYdzM
APA
Costa, G., Ribeiro, L., & Alves, R. (2026). Foundation Models for Drug Discovery: A Systematic Review of Molecular Language, Representation Learning, Target Prediction, Generative Design, and Translational Risk. International Journal of Pharmaceutical And Phytopharmacological Research, 16(1), 35-46. https://doi.org/10.51847/bWGXsPYdzM
Download citation:   EndNote   RIS
Article Link:
Downloads: 33
Views: 104
Abstract

Foundation models are increasingly used in computational drug discovery to learn molecular, protein, graph-based, and multimodal representations, but their translational value remains unevenly supported across tasks and validation settings. This systematic review evaluates how foundation models are applied to molecular language, representation learning, target prediction, compound–target interaction prediction, binding-related modelling, generative molecular design, and translational-risk assessment in drug discovery. PubMed, Scopus, Web of Science, IEEE Xplore, ScienceDirect, SpringerLink, Wiley Online Library, ACS Publications, Nature Portfolio journals, Oxford Academic journals, and quality-filtered MDPI and Frontiers journals were searched for peer-reviewed journal articles published from 1 January 2017 to 9 July 2026; 580 records were identified, 301 duplicates were removed, 279 records were screened, 105 full-text articles were assessed, 65 full-text articles were excluded with predefined reasons, and 40 studies were included. Evidence was extracted on model category, input modality, pretraining or adaptation strategy, drug-discovery task, benchmark setting, validation type, limitation, and translational risk. The included evidence shows that molecular language models, protein language models, graph-based self-supervised models, and knowledge-guided representation models may improve molecular representation, target-related prediction, compound–target modelling, and generative design support. However, evidence was dominated by benchmark and internal validation, with limited external, prospective, or experimental confirmation across many model categories. Foundation models are reshaping drug-discovery modelling, but benchmark performance should not be equated with chemical validity, biological plausibility, safety, or clinical utility. Stronger dataset provenance, leakage-resistant evaluation, uncertainty reporting, interpretability, prospective testing, and experimental confirmation are required before translational confidence can be justified.

Related articles:
Most viewed articles:
Naproxen in Pain and Inflammation – A Review
Vol 11 Issue 1, 2021 | Svetoslav Nikolaev Stoev
An Overview on Emulgel
Vol 9 Issue 1, 2019 | Sreevidya V.S
Volume 16
Issue 4
2026

Call for Papers
[email protected]
Issues
Associations
Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.

Copyright © 2026 International Journal of Pharmaceutical and Phytopharmacological Research
Authors retain copyright of their article if they are accepted for publication.