Physics of Failure Approach for Reliability Analysis of Liquid Cooled Data Centers
Ibaad Gandikota, Roshith Mittakolu, Vishal Raavi, Patrick McCluskeyAbstract
Data center cooling has evolved from air-cooling to liquid cooling, pushing the boundaries of a conventional cooling system to its limits. As the cooling designs evolve at a rapid pace, the reliability of such a system plays a leading role in avoiding any downtime. Physics-of-failure based reliability and availability modeling were applied in this study to a rack-level cooling system. The failure modes and mechanisms are listed, and based on these failure mechanisms, physics of failure models were applied to estimate the reliability and availability of the secondary cooling loop. The failure models were parameterized, and then these results were extended to different system architectures by adding redundancies to them. The results reveal that high availability can be achieved with reliable components as well as with redundant architectures, as availability increases from 0.9932 to 0.9997 in this study without additional redundancy. Hence, it is concluded that both the system architecture and component lifetimes play a crucial role in maximizing system availability.