Overview:
This research evaluates individual tree crown (ITC) detection performance of three transformer detectors: the Detection Transformer (DETR), Deformable DETR, and DETR with Improved DeNoising Anchor Boxes (DINO); using either an ImageNet-pretrained ResNet-50 backbone or a Masked Autoencoder (MAE) pretrained on unlabelled canopy heigh model (CHM) data.
The study also examines how data scarcity, forest diversity, spatial resolution, and training strategy affect model performance. Multiple datasets and training protocols are tested, including joint training and cross-dataset fine-tuning. The transformer models are benchmarked against Faster R-CNN and a classical local-maximum and watershed approach across several North American and European forest environments.
Objective(s):
- Evaluate how the architectural developments from standard DETR to Deformable DETR and DINO affect individual tree crown detection.
- Determine whether self-supervised MAE pretraining on unlabelled CHM data provides more effective, forest-specific feature representations than ImageNet-pretrained ResNet-50.
- Assess whether supplementing training with annotated data from diverse forest types and geographic regions improves model stability.
- Compare alternative training protocols, including single-dataset training, joint training on all available datasets, and pretraining on one dataset followed by fine-tuning on another.
Publication(s):
Gourde, J., Hu, B., & Li, Q. (2026). Transformer-Based Individual Tree Crown Detection from Canopy Height Models with Cross-Domain and Self-Supervised Pretraining. Remote Sensing, 18(11), 1674. https://doi.org/10.3390/rs18111674

