The need for manual curation of the cotton proteome leads to a solution primed for undergraduate involvement
Jonathan Zirkel, Amanda R. Storm, Grant T. Billings, Amanda M. Hulse‐KempSocietal Impact Statement
Cotton is a global crop cultivated worldwide, yet a significant portion of the cotton genome remains functionally unannotated. Attempts have been made to automate this annotation but have been unable to completely bridge the gap. Human curation is still needed for accurate interpretations and predictions to be made. We present a solution leveraging a course‐based approach to annotating protein function in a format tailored for undergraduate students. This format produces significant insight into individual proteins while introducing students to research in a low‐ to no‐cost footprint. In this way, this work provides a case study to bridge the gap between the next generation of scientists to deliver cutting‐edge science.
Summary
Cotton suffers from a modest fraction of functionally unannotated proteome. This problem is unsolvable by our current automated annotation tools. We set out to develop a solution that brings undergraduates into the research pipeline. We systematically identified all unannotated proteins and their corresponding genes and compiled a set of thousands of proteins suitable for manual downstream analysis. We devised a mutually beneficial solution, which provides manual annotation of unknown function genes through an analysis pipeline that can be used by undergraduate students as an authentic research experience. This approach ultimately provides valuable annotations disseminated in the form of a microPublication that would otherwise remain inaccessible. This analysis pipeline solution is modular and adaptable as technology improves and has successfully produced multiple microPublications, each of which provided novel predictions of form and function for an uncharacterized protein in the cotton genome. These microPublication reports are then shared directly with the cotton community by linking them to the gene in the community database, which ensures the results are broadly findable to provide valuable functional information for downstream research purposes.