HITL-CNN: Human-in-the-Loop Crop Disease Diagnosis Model

An EfficientNet-B4-based convolutional neural network for crop disease classification, combined with Monte Carlo Dropout uncertainty quantification and a confidence-threshold-based Human-in-the-Loop (HITL) referral mechanism. Developed as part of an MRes dissertation at the University of Greater Manchester.

Live demo: hitl-crop-diagnosis.streamlit.app Code: https://github.com/INITREASURE0611/hitl-cnn-app

Model Description

  • Architecture: EfficientNet-B4, ImageNet-pretrained backbone
  • Input: 224×224 RGB leaf images
  • Output: 38-class disease/healthy classification across 14 crop species
  • Uncertainty quantification: Monte Carlo Dropout (T=50 stochastic passes, p=0.1), with post-hoc temperature scaling for calibration
  • Referral mechanism: predictions with entropy above a calibrated threshold (θ=0.45 bits) are flagged for expert review rather than auto-classified

Training Data

Trained on the PlantVillage dataset (Hughes and Salathé, 2015), 54,305 images across 38 classes, licensed CC BY-NC-SA 4.0. Split 70/15/15 (train/val/test), stratified, seed=42.

Training Procedure

Two-stage fine-tuning: Stage 1 trains the classification head only (Adam, lr=1e-3, 2 epochs); Stage 2 fine-tunes the full network (Adam, lr=1e-4, ReduceLROnPlateau scheduler, early stopping patience=5). Training converged at epoch 19 (best checkpoint epoch 14).

Evaluation Results

Evaluated on a held-out test set of 8,146 images.

Metric Value
Accuracy 99.69%
Macro F1 0.9943
Macro Precision 0.9946
Macro Recall 0.9941
AUC-ROC (macro) 1.0000
ECE (post-calibration) 0.0025
HITL Referral Rate 2.6%
Error Capture Rate 88.0%

A controlled, protocol-matched comparison against a self-trained ResNet-50 baseline found ResNet-50 achieved marginally but statistically significantly higher accuracy (99.84% vs. 99.69%, McNemar's p=0.0227). This model's contribution is not claimed to rest on superior raw accuracy, but on the combination of near-ceiling accuracy with the calibrated HITL referral mechanism described above.

Limitations

  • Trained and evaluated exclusively on laboratory-condition images; no field-condition accuracy figure has been obtained
  • The Monte Carlo Dropout uncertainty pass (T=50) requires approximately 10 seconds per prediction on CPU, a real constraint for real-time or low-connectivity
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support