High-Precision Detection of Leather Creases via Dynamic, Cross-Calibrated, and Edge-Enhanced YOLOv8n-Pose
Ran An, Gongchang Ren, Jiangong Sun, Yuan Huan, Jiaxuan Yang, Kaijie Zhang, Yuanbiao WangResidual creases generated during the leather spreading process exhibit highly variable morphologies and irregular feature distributions, causing significant challenges for feature extraction and leading to low localization accuracy. To address these issues, this paper proposes the Dynamic, Cross-Calibrated, and Edge-Enhanced YOLOv8n-Pose (DCE-YOLOv8n-Pose) algorithm. Instead of providing regional approximations, this framework outputs precise spatial coordinates for robotic grasping by integrating three synergistic components in a progressive network flow. First, dynamic snake convolution adaptively perceives the continuous geometric features of elongated creases; subsequently, an efficient multi-scale attention mechanism provides cross-dimensional weight calibration to suppress highly homochromatic background interference and correct spatial misalignments; finally, an edge-enhanced content-aware reassembly of features module preserves high-frequency gradients and prevents feature fracturing during multi-scale fusion. For comprehensive evaluation, a dataset comprising 700 original laboratory images was constructed. To prevent data leakage, the dataset was partitioned into training and validation sets based on individual leather specimens, ensuring that images of the same leather piece do not appear in both sets. Additionally, an independent test set of 500 images collected from an actual processing plant was designed for industrial validation. Experimental results indicate that, at an Intersection over Union (IoU) threshold of 0.5, the DCE-YOLOv8n-Pose model achieves a bounding box mean average precision (mAP@0.5) of 91.8% and a keypoint mAP@0.5 of 85.1%, with a keypoint precision of 87.9%. The computational load is maintained at 9.2 GFLOPs, alongside an inference speed of 114.3 FPS. Furthermore, consistent convergence across four independent training runs substantiates the model’s reliability in reducing missed detection rates and localization deviations. In conclusion, the proposed algorithm demonstrates practical applicability for the visual guidance of automated leather spreading equipment by balancing detection precision and inference speed, thereby offering an effective coordinate reference for subsequent robotic stretching operations.