DOI: 10.3390/s26165213 ISSN: 1424-8220

A Dual-Camera Edge Sensing Framework with Zone-Aware Multi-Object Tracking for Sensorless Smart Vending Cabinets

Abror Shavkatovich Buriboev, Farkhat Rajabov, Shavkat Buriboev, Rustem Allanyazov, Giyosjon Sharipov, Abbos Abduvaytov, Aziza Akhmedova, Ruzimboy Sobirov, Su-Mi Shin, Cheolwon Lee, Heung Seok Jeon

Top-loading smart vending cabinets require precise transaction-level product detection under strict hardware and deployment constraints. In this paper, “sensorless” refers specifically to the absence of auxiliary product-level sensing hardware, such as RFID tags, weight sensors, shelf load cells, or product-slot instrumentation; the system still uses two camera sensors. Conventional snapshot-difference methods compare only a small number of frames at the beginning and end of a transaction and therefore cannot explicitly represent intermediate product motion, such as pickup, return, inspection, occlusion, and shelf resettling. This paper proposes ZAB-Fusion, a dual-camera edge sensing framework with zone-aware multi-object tracking for sensorless smart vending cabinets. The framework combines a YOLO11-seg and RT-DETR detection ensemble with ByteTrack temporal association, projects product tracks into a three-zone vertical cabinet model, interprets compressed zone sequences using a finite-state event classifier, and integrates camera-specific event streams through an evidence-gated cross-camera fusion rule. The proposed method was evaluated on 220 in-service vending transactions containing 227 ground-truth TAKEN events and 87 RETURNED events across seven product classes. Compared with the snapshot-difference baseline, ZAB-Fusion improved recall from 0.665 to 0.925 and F1-score from 0.780 to 0.944, while maintaining a high precision of 0.963. At the transaction level, exact receipt accuracy increased from 0.645 to 0.900. Runtime analysis on an Intel N100 CPU-only edge device showed an average processing latency of 562 ms per transaction under the selected-frame inference protocol. The zone classification, finite-state event interpretation, and evidence-gated cross-camera fusion stages required only 3 ms in total. These results demonstrate that explicit motion semantics and auditable cross-camera evidence gating can improve sensorless retail transaction level recognition in sensorless smart vending cabinets without adding auxiliary product-level sensing hardware.

More from our Archive