RF-NR: Random Forest Based Approach for Improved Classification of Nuclear Receptors

Hamid D. Ismail, Hiroto Saigo, Dukka B. Kc

Research output: Contribution to journalArticle

3 Citations (Scopus)

Abstract

The Nuclear Receptor (NR) superfamily plays an important role in key biological, developmental, and physiological processes. Developing a method for the classification of NR proteins is an important step towards understanding the structure and functions of the newly discovered NR protein. The recent studies on NR classification are either unable to achieve optimum accuracy or are not designed for all the known NR subfamilies. In this study, we developed RF-NR, which is a Random Forest based approach for improved classification of nuclear receptors. The RF-NR can predict whether a query protein sequence belongs to one of the eight NR subfamilies or it is a non-NR sequence. The RF-NR uses spectrum-like features namely: Amino Acid Composition, Di-peptide Composition, and Tripeptide Composition. Benchmarking on two independent datasets with varying sequence redundancy reduction criteria, the RF-NR achieves better (or comparable) accuracy than other existing methods. The added advantage of our approach is that we can also obtain biological insights about the important features that are required to classify NR subfamilies. RF-NR is freely available at http://bcb.ncat.edu/RF-NR/.

Original languageEnglish
Article number8107505
Pages (from-to)1844-1852
Number of pages9
JournalIEEE/ACM Transactions on Computational Biology and Bioinformatics
Volume15
Issue number6
DOIs
Publication statusPublished - Nov 1 2018

All Science Journal Classification (ASJC) codes

  • Biotechnology
  • Genetics
  • Applied Mathematics

Fingerprint Dive into the research topics of 'RF-NR: Random Forest Based Approach for Improved Classification of Nuclear Receptors'. Together they form a unique fingerprint.

  • Cite this