← Papers

Unverified paper record

BCAR-Net: A Bidirectional Cross-Attention Network with Auxiliary Reconstruction for Tree Counting in Complex Forest Scenes Using Airborne RGB and LiDAR Data

Plants (Basel, Switzerland) · 6 Jun 2026 · 10.3390/plants15121762

Abstract

Accurate tree counting from remote sensing data is essential for forest inventory, biomass estimation, carbon accounting, and ecological monitoring. However, existing approaches predominantly rely on airborne RGB imagery and often struggle in complex forest scenes where neighboring crowns exhibit highly similar textures and colors and where overlapping crown boundaries become ambiguous. To address this limitation, the LiDAR-derived Canopy Height Model (CHM) is introduced as a complementary modality that provides explicit cues on canopy height variation and vertical structure to support RGB-based analysis. Building on this, we propose BCAR-Net, a broker-guided RGB and depth (RGB-D) multimodal framework that couples bidirectional cross-modal interaction, adaptive tri-branch fusion, and auxiliary reconstruction within a two-stage optimization scheme. Specifically, a bidirectional cross-attention U-Net generates an intermediate broker RGB-D representation from paired RGB images and depth maps through symmetric bidirectional cross-attention between the two modalities and direction-aware gating. The original RGB image, depth map, and broker representation are then jointly encoded by three weight-sharing branches and adaptively aggregated by a spatial fusion gate for density-map regression. To regularize the fused latent feature, a multi-scale cross-attention reconstruction decoder provides auxiliary RGB and depth reconstruction supervision by querying multi-scale BCA-UNet encoder features through 2D cross-attention, and a reconstruction-oriented first stage replaces externally generated fused-image supervision, yielding a task-consistent optimization scheme. Experiments on the NEONTreeEvaluation benchmark show that BCAR-Net consistently outperforms single-modality settings and direct RGB-D concatenation multimodal baseline. Additional experiments on a public UAV RGB-LiDAR dataset provide a small-scale supplementary evaluation under a different acquisition setting, where BCAR-Net achieves modest but consistent improvements over RGB-only and depth-only baselines. These results demonstrate that the proposed framework offers an effective but computationally cautious solution for tree counting in complex forest environments.

Plant phenotyping relevance

RGB画像とLiDAR由来データから樹木数を推定する深層学習手法を開発し、複数ベンチマークで比較評価しており、植物個体の計測手法が研究の中心である。

abstractwe propose BCAR-Net, a broker-guided RGB and depth (RGB-D) multimodal framework
abstractExperiments on the NEONTreeEvaluation benchmark show that BCAR-Net consistently outperforms single-modality settings and direct RGB-D concatenation multimodal baseline.
abstractThese results demonstrate that the proposed framework offers an effective but computationally cautious solution for tree counting in complex forest environments.

Code and data availability

The paper evaluates on two third-party public datasets (NEONTreeEvaluation benchmark and the Dubrovin et al. UAV RGB–LiDAR dataset), but these are cited prior-work resources, not paper-specific deposits. No authors' code, trained models, or supplementary data availability statement appears in the supplied blocks.

No evidence-backed public reproduction asset is currently recorded.

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.