Publicly available datasets were analyzed in this study. This data can be found here: https://ieee-dataport.org/documents/goa-cashew-apple-maturity-grading .
Open resource ↗goa-cashew-apple-maturity-grading · lines:837-851Unverified paper record
Distilled vision transformers with CNN fusion for robust cashew apple maturity prediction.
Frontiers in plant science · 21 Apr 2026 · 10.3389/fpls.2026.1787609
Abstract
Introduction Cashew apple is a nutrient-rich fruit containing abundant minerals, vitamins, and energy. However, its fleshy texture and delicate skin significantly limit its storage life and market value. Accurate maturity grading is therefore essential for improving post-harvest management and transportation efficiency. Methods This study proposes a lightweight vision transformer (ViT) student model trained using multi-granular knowledge distillation (KD) from a stronger data-efficient image transformer (DeiT)-Base teacher. The distillation framework integrates response-based soft-label supervision, attention transfer, and token-level feature regression to enhance representation learning under limited data conditions. Auxiliary lightweight architectures, including MobileNet, ConvNeXt, and EdgeNeXt, were trained independently to provide complementary predictions, and a weighted fusion strategy was employed for ensemble evaluation. Results The proposed ensemble ViT-KD with EdgeNeXt achieved 90% accuracy under the evaluated test split. To ensure statistical reliability and address potential partition bias, a stratified fivefold cross-validation was conducted on the dataset, yielding a mean accuracy of 86.89% ± 2.89% with consistent F1 scores and recall. The relatively low variance across the folds indicates stable internal generalization. Comparative experiments with conventional convolutional neural network (CNN) baselines and lightweight CNN baselines such as MobileViT-S and ShuffleNetV2 were performed, with the proposed ensemble framework achieving improved accuracy while maintaining computational efficiency. Computational analysis indicates that the stand-alone distilled ViT maintains a real-time inference capability of 8.79 ms per image, which supports suitability for edge-oriented agricultural applications. Discussion These results highlight the effectiveness of knowledge-distilled lightweight transformers for data-efficient maturity grading of cashew apples.
Plant phenotyping relevance
カシューナッツ果実の成熟度という植物器官の状態を画像から推定する手法を開発し、交差検証・比較実験・推論速度評価まで行っており、フェノタイピング手法が中心である。
abstractThis study proposes a lightweight vision transformer (ViT) student model trained using multi-granular knowledge distillation (KD) from a stronger data-efficient image transformer (DeiT)-Base teacher.
abstractTo ensure statistical reliability and address potential partition bias, a stratified fivefold cross-validation was conducted on the dataset
abstractThese results highlight the effectiveness of knowledge-distilled lightweight transformers for data-efficient maturity grading of cashew apples.
Code and data availability
The paper's cashew apple maturity grading experiments use a public image dataset from IEEE Dataport (Sawant, 2025), explicitly linked in the data availability statement. No author code or models are shared.
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.