Large-Scale Patching Automation: Sustaining 160,000 Servers with HP SA
Keeping operating systems updated in global infrastructures is one of the greatest challenges in information security. When the environment encompasses more than 160,000 virtual and physical servers across multiple global data centers, manual execution of updates is no longer a viable option.
In this article, I detail the patching automation architecture developed with HP Server Automation (HP SA) and advanced scripts in Bash/Unix, which reduced the update cycle from weeks to just a few hours.
The Challenge of Scale
In enterprise Data Centers, frequent patching runs into three major pillars: 1. Strict Maintenance Windows: Servers have very short windows (sometimes less than 2 hours per month) to apply updates and reboot. 2. Service Dependencies: Critical applications (such as SAP databases, legacy UNIX systems) need to be stopped and validated in an orchestrated manner. 3. Auditability: Each applied update must be recorded in regulatory compliance systems.HP SA Architecture with Shell Script
The solution was based on encapsulating HP SA APIs (hpsa-client) inside optimized, parallel loops in Bash. The main script runs in three phases:
Phase 1: Prerequisites Validation
check_dependencies() {
local server=$1
echo "[CHECK] Verificando espaço em /var e status do HPSA Agent em $server..."
# Simulated command
hpsa-agent-status --server "$server"
}
Orchestration and Parallelism
To avoid bottlenecks in the Data Center network, we implemented a execution system using semaphore batches, where a maximum of 50 servers were updated simultaneously per network zone.A batch of 50, not the whole estate
The agent answers. What saturates the network is firing 160,000 checks at once. The batch stays at 50 per zone and only continues when a slot frees up.
MAX_PARALLEL=50
while read -r server; do
hpsa-agent-status --server "$server" &
while [ "$(jobs -r | wc -l)" -ge "$MAX_PARALLEL" ]; do
wait -n 2>/dev/null || wait
done
done < servers.txt
wait
echo "[COMPLETE] zone pre-check finished"
servers.txt is one zone, not the global inventory. The core dispatches the work; the satellite runs it close to the server; the agent returns status. That split is what fits inside a two-hour window.
Results
Security Compliance: Rose from 72% to 99.4% across the entire server fleet. Overhead Reduction: The effort of N1/N2 teams in manual operating system update processes was reduced by 90%. * Stability: Zero unavailability failures due to broken dependency packages, thanks to the integrated simulation engine before the final commit.Related Articles
Multi-cloud for Banks: Separating the Core, Data, and the Audit Trail
Multi-cloud architecture for banks: a locked region, a core isolated from the lab, and an audit trail nobody can delete. SCP, CloudTrail, and Object Lock commands.
SAP Business One with SAP HANA on AWS: Architecture and Day-to-Day Operations
How to run SAP Business One on SAP HANA on AWS: instance, data and log volumes, port 30015, and backup. Power BI reads the replica, not production HANA.