← Papers

Unverified paper record

High throughput measurement of Arabidopsis thaliana fitness traits using transfer learning

bioRxiv (Cold Spring Harbor Laboratory) · 1 Jul 2021 · 10.1101/2021.07.01.450758

Abstract

Summary Revealing the contributions of genes to plant phenotype is frequently challenging because the effects of loss of gene function may be subtle or be masked by genetic redundancy. Such effects can potentially be detected by measuring plant fitness, which reflects the cumulative effects of genetic changes over the lifetime of a plant. However, fitness is challenging to measure accurately, particularly in species with high fecundity and relatively small propagule sizes such as Arabidopsis thaliana . An image segmentation-based (ImageJ) and a Faster Region Based Convolutional Neural Network (R-CNN) approach were used for measuring two Arabidopsis fitness traits: seed and fruit counts. Although straightforward to use, ImageJ was error-prone (correlation between true and predicted seed counts, r 2 =0.849) because seeds touching each other were undercounted. In contrast, Faster R-CNN yielded near perfect seed counts (r 2 =0.9996) and highly accurate fruit counts (r 2 =0.980). By examining seed counts, we were able to reveal fitness effects for genes that were previously reported to have no or condition-specific loss-of-function phenotypes. Our study provides models to facilitate the investigation of Arabidopsis fitness traits and demonstrates the importance of examining fitness traits in the study of gene functions.

Plant phenotyping relevance

画像分割とFaster R-CNNを用いて種子数・果実数という植物形質を高スループット測定し、精度比較・検証を行うことが中心であるため。

titleHigh throughput measurement of Arabidopsis thaliana fitness traits using transfer learning
abstractAn image segmentation-based (ImageJ) and a Faster Region Based Convolutional Neural Network (R-CNN) approach were used for measuring two Arabidopsis fitness traits: seed and fruit counts.
abstractIn contrast, Faster R-CNN yielded near perfect seed counts (r 2 =0.9996) and highly accurate fruit counts (r 2 =0.980).

Code and data availability

The paper's Data availability statement explicitly deposits all analysis scripts and the final seed and fruit counting models (trained Faster R-CNN phenotyping models) on the authors' public GitHub repository, which is listed in allowed_urls.

Codepublic

, PD, SH, 713 NLP, EV, EW, JKC, PJK, and MDL performed data collection and analysis. PW, FM, 714 MDL, and SHS wrote the manuscript. All authors read and approved the final manuscript. 715 716 Data availability 717 All the scripts used in this study and the final seed and fruit counting models are available 718 on Github at: 719 https://github.com/ShiuLab/Manuscript_Code/tree/master/2021_Arabidopsis_seed_and_f 720 ruit_count 721 722 References 723 Abadi M, Barham P, Chen JM, Chen ZF, Davis A, Dean J, Devin M, Ghemawat S, 724 Irving G, Isard M et al. 2016. TensorFlow: A system for large-scale machine 725 learning. 12th USENIX Symposium on Operating Systems Design and 726 Implementation. USENIX

Open resource ↗ShiuLab/Manuscript_Code · 2021_Arabidopsis_seed_and_f · pdf-raw-page:33 lines:1-52

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.