Ramadhani, Queenta Paradissa and Dewanta, Favian and Gunawan, Teddy Surya and Kartiwi, Mira and Nugroho, Bambang Setia (2026) Similar accuracy, different explanations in YOLO-based palm fruit ripeness detection. IIUM Engineering Journal, 27 (3). pp. 340-365. ISSN 1511-788X E-ISSN 2289-7860
|
PDF
- Published Version
Download (2MB) | Preview |
|
|
PDF
- Supplemental Material
Download (208kB) | Preview |
|
|
PDF
- Supplemental Material
Download (264kB) | Preview |
Abstract
Modern palm-fruit detectors can achieve similarly high benchmark accuracy, yet class activation mapping (CAM) methods are often applied without establishing whether their explanatory behavior transfers across detector configurations. This study is a controlled exploratory within-dataset comparison: model scale and target layer were selected on the same 416-image test partition used for final CAM evaluation, so the reported values are not independent estimates of explanation generalization. We evaluated 15 YOLOv8, YOLO11, and YOLO26 variants under a matched training protocol and assessed representative medium models across three random seeds. Backbone representations were then audited on the 416 test images using class-agnostic EigenCAM, backbone-adapted Grad-CAM++ (bGrad-CAM++), and backbone-adapted Score-CAM (bScore-CAM). Explanation behavior was quantified with CAM-box intersection over union (IoU) and Pointing Game for spatial localization and with Mean Absolute Confidence Drop (MACD) and Increase in Confidence as image-level confidence-retention diagnostics. Repeated-run mean mAP50–95 was narrowly separated at 0.903 ± 0.002, 0.899 ± 0.002, and 0.901 ± 0.003 for YOLOv8m, YOLO11m, and YOLO26m, respectively. In contrast, Pointing Game varied by 68.03 percentage points and MACD by 48.62 percentage points across model–explainer configurations. EigenCAM showed the strongest observed profile on YOLOv8m, whereas bScore-CAM achieved the highest localization on YOLO11m; no explainer dominated across all evaluated backbones. The ranking reversal was also present within the difficult underripe, ripe, and overripe subset, and the recorded layer search showed that layer 8 was the strongest compatible backbone target across the layer 6–10 window in all three detectors. Similar detection accuracy therefore did not imply transferable explanation behavior. CAM explanations should be validated jointly with the explanatory target, layer, and detector rather than treated as interchangeable post-hoc visualizations.
Actions (login required)
![]() |
View Item |
