Identification of useful genes from multiple microarrays for ulcerative colitis diagnosis based on machine learning methods

Lin Zhang, Rui Mao, Chung Tai Lau, Wai Chak Chung, Jacky C. P. Chan, Feng Liang, Chenchen Zhao, Xuan Zhang*, Zhaoxiang Bian*

*Corresponding author for this work

Research output: Contribution to journalJournal articlepeer-review

11 Citations (Scopus)


Ulcerative colitis (UC) is a chronic relapsing inflammatory bowel disease with an increasing incidence and prevalence worldwide. The diagnosis for UC mainly relies on clinical symptoms and laboratory examinations. As some previous studies have revealed that there is an association between gene expression signature and disease severity, we thereby aim to assess whether genes can help to diagnose UC and predict its correlation with immune regulation. A total of ten eligible microarrays (including 387 UC patients and 139 healthy subjects) were included in this study, specifically with six microarrays (GSE48634, GSE6731, GSE114527, GSE13367, GSE36807, and GSE3629) in the training group and four microarrays (GSE53306, GSE87473, GSE74265, and GSE96665) in the testing group. After the data processing, we found 87 differently expressed genes. Furthermore, a total of six machine learning methods, including support vector machine, least absolute shrinkage and selection operator, random forest, gradient boosting machine, principal component analysis, and neural network were adopted to identify potentially useful genes. The synthetic minority oversampling (SMOTE) was used to adjust the imbalanced sample size for two groups (if any). Consequently, six genes were selected for model establishment. According to the receiver operating characteristic, two genes of OLFM4 and C4BPB were finally identified. The average values of area under curve for these two genes are higher than 0.8, either in the original datasets or SMOTE-adjusted datasets. Besides, these two genes also significantly correlated to six immune cells, namely Macrophages M1, Macrophages M2, Mast cells activated, Mast cells resting, Monocytes, and NK cells activated (P  <  0.05). OLFM4 and C4BPB may be conducive to identifying patients with UC. Further verification studies could be conducted.

Original languageEnglish
Article number9962
Number of pages13
JournalScientific Reports
Issue number1
Publication statusPublished - 15 Jun 2022

Scopus Subject Areas

  • General


Dive into the research topics of 'Identification of useful genes from multiple microarrays for ulcerative colitis diagnosis based on machine learning methods'. Together they form a unique fingerprint.

Cite this