QSAR Datasets: Where to Find Them and How to Clean Them
By BioDockify Computational Research Team · October 11, 2026 · QSAR
Building a QSAR model starts with the dataset: the four public databases that matter (ChEMBL, BindingDB, PubChem, DUD-E), then the cleaning pipeline that decides model quality - activity harmonization, duplicates, cutoffs, and the scaffold split that random splits fake.
Read the complete article with tables, code and references on BioDockify.
Key References
- Gaulton, A. et al. ChEMBL: a large-scale bioactivity database for drug discovery. Nucleic Acids Res. 40, D1100-D1107 (2012). DOI: 10.1093/nar/gkr777
- Liu, T. et al. BindingDB: a web-accessible database of experimentally determined protein-ligand binding affinities. Nucleic Acids Res. 35, D198-D201 (2007).
- Mysinger, M.M., Carchia, M., Irwin, J.J. & Shoichet, B.K. DUD-E. J. Med. Chem. 55, 6582-6594 (2012). DOI: 10.1021/jm300687e
- Tropsha, A. Best practices for QSAR model development, validation, and exploitation. Mol. Inform. 29, 476-488 (2010). DOI: 10.1002/minf.201000062
Scope & Limitations
Public bioactivity data contain assay inconsistencies no cleaning fully removes - model ceilings reflect data quality, not algorithm choice Scaffold-split validation is the honest default, but for very small datasets temporal or leave-cluster-out variants may be needed