Unverified paper record
Bridging CNNs and vision transformers for efficient tea leaf phytopathogen diagnosis: GL-MobFormer.
Frontiers in plant science · 12 May 2026 · 10.3389/fpls.2026.1814962
Abstract
Background Accurate identification of visible disease symptoms is essential for the sustainable management of tea ( Camellia sinensis ) cultivation. However, balancing high diagnostic accuracy with the computational efficiency required for deployment on agricultural edge devices remains a significant challenge. Methods We propose GL-MobFormer, a lightweight hybrid deep learning framework. This architecture integrates the local feature extraction capabilities of MobileNetV3 with the global contextual modeling of a Transformer Encoder. To improve model robustness in unstructured field environments, we applied the CutMix data augmentation strategy. The framework was evaluated on a dataset comprising 5,278 tea leaf images across seven phytosanitary categories. Results Empirical evaluations demonstrate that GL-MobFormer achieved a classification accuracy of 95.13% and a Matthews Correlation Coefficient (MCC) of 0.9417. Crucially, this performance was maintained with a low computational footprint of merely 0.33 G FLOPs(Floating Point Operations). Importantly, an occlusion-based sensitivity protocol was implemented to provide quantitative grounding for model interpretability. Results revealed that systematically masking only the top 5% of critical activation regions led to an average reduction of 70.61% in classification confidence, empirically confirming that the model's diagnostic logic is faithfully anchored on pathologically relevant lesion features rather than background noise. Conclusion GL-MobFormer achieves an optimal trade-off between diagnostic precision and computational overhead. It provides a practical and highly efficient solution for on-site, real-time phytosanitary monitoring in precision agriculture.
Plant phenotyping relevance
茶葉の病斑画像から病害状態を推定する軽量CNN・Transformer手法を開発し、精度・計算量・解釈性を評価しており、植物フェノタイピング手法が中心です。
abstractWe propose GL-MobFormer, a lightweight hybrid deep learning framework.
abstractan occlusion-based sensitivity protocol was implemented to provide quantitative grounding for model interpretability.
abstractsystematically masking only the top 5% of critical activation regions led to an average reduction of 70.61% in classification confidence
Code and data availability
The paper uses the teaLeafBD image dataset sourced from Mendeley Data, but this is a cited third-party dataset (Alam et al., 2025) with no authors' URL provided in the supplied blocks. No author analysis code, trained model checkpoints, or paper-specific public repository is mentioned with an explicit availability/depo
No evidence-backed public reproduction asset is currently recorded.
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.