← Papers

Unverified paper record

Improving Multi-Trait Genomic Prediction Efficiency Through The Incorporation Of Synthetic Traits Selected Based on Co-heritability

bioRxiv (Cold Spring Harbor Laboratory) · 27 Jun 2025 · 10.1101/2025.06.23.661178

Abstract

Abstract Genomic prediction (GP) is an essential tool in the field of plant breeding to accelerate the cultivar development pipeline by predicting the performance of unphenotyped lines. The precision of prediction is constrained by the heritability of the target trait when applying a single-trait genomic prediction model. To overcome this limitation, a multi-trait genomic prediction model leveraging high-heritability secondary traits co-heritable with the target trait can boost predictive ability for the target trait. However, this is practically challenging because it requires additional phenotyping effort and prior knowledge of trait co-heritability. This study aimed to assess the efficiency of multi-trait genomic prediction models powered by secondary traits derived from high-throughput phenotyping data when predicting important leaf functional target traits, i.e., nitrogen (N) content and specific leaf area (SLA) in diverse sorghum accessions. Since these traits can be predicted from hyperspectral reflectance data, there is significant potential for other wavelengths within the existing dataset to meet the criteria needed to improve prediction accuracy using multi-trait approaches. Therefore, experiments were performed on traditional direct measures of leaf N content and SLA, plus partial least squares regression predictions of them (Leaf N-PLSR, SLA-PLSR), i.e., four target traits in total. Three secondary, “synthetic traits” (S1, S2, S3), each a ratio of two wavelengths within the hyperspectral data, were identified based on high co-heritability with a given target trait. Single-trait GBLUP (Genomic Best Linear Unbiased Predictor) was fitted as a baseline model, followed by three multi-trait GBLUP models using synthetic traits and target traits together. Model performance was assessed using k-fold (k=5) cross-validation (CV), which consisted of single-trait, CV1, and CV2 schemes. The synthetic traits’ high genetic correlation and heritability met the requirements for their use as secondary traits. There was a significant increase in accuracy when synthetic traits were used in the multi-trait genomic prediction model compared to a single trait alone for all four target traits. It improved prediction accuracy while using secondary traits derived from hyperspectral high-throughput phenotyping data in the multi-trait genomic prediction model, suggesting that this approach could be broadly applied in a post-hoc fashion to many datasets without any additional phenotyping effort. Our analysis highlights a practical approach to improve multi-trait genomic prediction model performance using synthetic traits with no intrinsic biological meaning selected through co-heritability estimation.

Plant phenotyping relevance

ハイパースペクトル高スループット表現型データから合成形質を抽出し、PLSR予測と遺伝相関に基づいてゲノム予測へ利用する解析手法が研究の中心であるため。

abstractsecondary traits derived from high-throughput phenotyping data
abstractThree secondary, “synthetic traits” (S1, S2, S3), each a ratio of two wavelengths within the hyperspectral data, were identified based on high co-heritability with a given target trait.
abstractIt improved prediction accuracy while using secondary traits derived from hyperspectral high-throughput phenotyping data in the multi-trait genomic prediction model

Code and data availability

The supplied blocks describe sorghum hyperspectral phenotyping and multi-trait genomic prediction methods but contain no data availability statement, no public phenotype/hyperspectral dataset link, and no author code repository or deposit language. The genotypic data is cited from prior publications, not a paper-asset.

No evidence-backed public reproduction asset is currently recorded.

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.