User activity growth on modern web resources is rarely linear or predictable. Large-scale marketing campaigns, major international tournaments, or the release of highly anticipated new games can trigger an avalanche of visitors within minutes. If the platform's IT infrastructure is insufficient in flexibility, application servers will quickly run out of RAM, and users will experience critical delays or complete interface inaccessibility. To address this issue, technology companies are implementing horizontal pod autoscaling systems. A detailed analysis of container orchestrators' operating principles is available at https://loot-zino.com . The Kubernetes orchestrator is the primary tool for managing infrastructure under variable loads. The autoscaling logic is based on continuous monitoring of system metrics for each running container—processor (CPU) utilization and RAM consumption. As soon as the load on current microservices exceeds a safe threshold set by engineers (for example, 70%), Kubernetes instantly and automatically deploys new copies (replicas) of containers from a prepared template. |