← Papers

Unverified paper record

Genomic prediction reveals unexplored variation in grain protein and lysine content across a vast winter wheat genebank collection.

Frontiers in plant science · 11 Jan 2024 · 10.3389/fpls.2023.1270298

Abstract

Globally, wheat ( Triticum aestivum L.) is a major source of proteins in human nutrition despite its unbalanced amino acid composition. The low lysine content in the protein fraction of wheat can lead to protein-energy-malnutrition prominently in developing countries. A promising strategy to overcome this problem is to breed varieties which combine high protein content with high lysine content. Nevertheless, this requires the incorporation of yet undefined donor genotypes into pre-breeding programs. Genebank collections are suspected to harbor the needed genetic diversity. In the 1970s, a large-scale screening of protein traits was conducted for the wheat genebank collection in Gatersleben; however, this data has been poorly mined so far. In the present study, a large historical dataset on protein content and lysine content of 4,971 accessions was curated, strictly corrected for outliers as well as for unreplicated data and consolidated as the corresponding adjusted entry means. Four genomic prediction approaches were compared based on the ability to accurately predict the traits of interest. High-quality phenotypic data of 558 accessions was leveraged by engaging the best performing prediction model, namely EG-BLUP. Finally, this publication incorporates predicted phenotypes of 7,651 accessions of the winter wheat collection. Five accessions were proposed as donor genotypes due to the combination of outstanding high protein content as well as lysine content. Further investigation of the passport data suggested an association of the adjusted lysine content with the elevation of the collecting site. This publicly available information can facilitate future pre-breeding activities.

Plant phenotyping relevance

小麦のタンパク質・リジン含量という植物形質について、歴史的表現型データを整理し、複数のゲノム予測法を比較して大規模コレクションの予測表現型を生成しており、計算的な形質推定とデータセット活用が研究の中心です。

abstracta large historical dataset on protein content and lysine content of 4,971 accessions was curated, strictly corrected for outliers as well as for unreplicated data and consolidated as the corresponding adjusted entry means.
abstractFour genomic prediction approaches were compared based on the ability to accurately predict the traits of interest.
abstractFinally, this publication incorporates predicted phenotypes of 7,651 accessions of the winter wheat collection.

Code and data availability

The authors deposited the paper's curated historical protein/lysine phenotype data (ISA-Tab), the R code for BLUE calculation and genomic prediction with all input files, and key output files (BLUEs and predicted phenotypes) in the public e!DAL repository under DOI 10.5447/ipk/2023/20. This is a paper-specific, public,

Datasetpublic

n with all input files, and the most important output files of the analysis. The output files include BLUEs of protein and lysine content as well as the predictions of protein content, lysine content and adjusted lysine content. The aforementioned information is available via the e!DAL ( Arend et al., 2014 ) online repository ( https://dx.doi.org/10.5447/ipk/2023/20 ). Author contributions MB: Conceptualization, Formal Analysis, Investigation, Methodology, Software, Visualization, Writing – original draft. SW: Data curation, Writing – review & editing. JR: Conceptualization, Methodology, Supervision, Writing – review & editing. AS: Conceptualization, Methodology, Supervision, Validation, W

Open resource ↗e!DAL · 10.5447/ipk/2023/20 · lines:302-323

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.