DOI: 10.3390/electronics15194442 ISSN: 2079-9292

An Edge-Cloud Infrastructure for Scalable Personalization of AI Models in Community Health Care

Rakesh Suvvari, Awatif Yasmin, Sana Alamgeer, Tarek Mahmud, Anne H. H. Ngu

Modern Internet of Things (IoT) systems increasingly rely on machine learning/deep learning models deployed across edge devices and cloud infrastructure. However, many real-world IoT systems require personalized models for each user or device, making scalable retraining, deployment, and real-time inference challenging. To address this, we propose a practical low-cost edge–cloud framework for concurrent personalization of AI models in IoT systems that can support users in a community-dwelling center. The personalization is achieved with a data-efficient machine learning retraining model (MODELXc) coupled with a layered software architecture that separates edge sensing, communication, cloud processing, and model management. First, retraining of AI model jobs runs as Kubernetes pods with pre-configured CPU and memory limits, and users are grouped into batches to prevent resource oversubscription during concurrent training. Second, edge devices communicate with the cloud using NATS.io request-reply messaging with binary-encoded payloads for low-latency communication. Third, a lifecycle pipeline handles versioned model storage, in-memory caching, and background synchronization. We evaluated this framework using two IoT applications: a wearable fall-detection system and a heart-rate-monitoring system. The second application is used to demonstrate the cross application portability of our infrastructure. Using the MODELXc personalization workflow, we executed 40 user-specific training jobs concurrently on a 40-logical-CPU server within one minute. We also demonstrated experimentally that a 40-CPU server can support personalization of 5000 virtual users in a 6–7 h period using a batch-wise processing strategy. Docker- and Kubernetes-based orchestration provided bounded per-job resource usage and resulted in a predictable system load compared with traditional background execution. For real-time inference, NATS.io with binary encoding and controlled throttling achieved a median latency of approximately 110–120 ms, compared with approximately 220 ms using the default without binary encoding.