How to Keep a Crash-Loop Exploit From Taking Down Both Nodes of an HA Pair
🌐 Networking Intermediate 6 min read

How to Keep a Crash-Loop Exploit From Taking Down Both Nodes of an HA Pair

Failover does not help when the attacker crashes the service on demand. A runbook for stopping the trigger, preserving evidence, and patching an HA pair in the right order.

Published: October 4, 2026 • Updated: October 4, 2026
high availabilityincident responsefailoverrunbooknetwork appliances

High availability protects you from hardware failure, from a bad power supply, from a kernel panic on one node. It does not protect you from an attacker who can crash the service on demand, because the standby node runs the same software with the same bug and inherits the same virtual IP. When Citrix NetScaler appliances started crash-looping on crafted SAML requests in October 2026 (CVE-2026-88779), administrators watched exactly this happen: the primary's authentication daemon died, the watchdog rebooted the box, the secondary took the VIP, received the same traffic, and died too. This guide is the runbook for that situation.

Why failover makes it worse, not better

An HA pair is designed around the assumption that failures are independent. One node's disk fails; the other's does not. An exploitExploit🛡️Code or technique that takes advantage of a vulnerability to cause unintended behavior, such as gaining unauthorized access. is a correlated failure. The attacker sends a request to the service address, not to a node, so whichever node holds the address gets the request. Failover moves the target, and the attacker does not even have to notice.

Two things make it actively worse. First, failover is often triggered by health checks or watchdogs that count consecutive process failures, so a crash loop on the primary forces a failover, and a crash loop on the new primary forces another. The pair can flap back and forth, which looks like instability and tends to prompt a reboot of both nodes at once. Second, every reboot destroys forensic state: in-memory process state, unsaved core files, and often the logs that would have shown what request caused the crash. A pair that has flapped for a few hours may have no usable evidence left.

Step 1: Stop the traffic, not the service

The instinct when an appliance is crash-looping is to reboot it cleanly or fail it over by hand. Neither helps. The first move is to stop the triggering input from reaching either node.

For an authentication-daemon crash this means identifying which virtual server and which authentication path the requests hit. In the NetScaler case Citrix's own guidance was to review Gateway and AAA configurations for SAML authentication actions. If you can confirm the attack path, you have options short of taking the whole gateway down:

  • Bind a responder or request-filtering policy to the affected virtual servers that drops the malformed requests before they reach the vulnerable parser. Citrix made such a policy available through support on 3 October. Verify the feature the policy depends on is actually enabled; the community-maintained NetScaler checker flagged deployments where a responder policyResponder Policy🌐A NetScaler feature that inspects incoming requests against a rule and takes an action such as dropping, resetting, or redirecting them before the request reaches the back-end service. Vendors sometimes supply a responder policy as an interim workaround for a vulnerability until a fixed build is installed. was bound but the Responder feature was off, which makes the mitigation a no-op.
  • Restrict the source addresses that can reach the authentication endpoint, if your user population allows it. A corporate gateway serving known offices and a known VPN concentrator can often be fenced for a day.
  • Temporarily unbind the affected authentication factor and fall back to a different one, accepting the user-experience cost. Why disabling a feature to dodge a zero-dayZero-Day🛡️A security vulnerability that is exploited or publicly disclosed before the software vendor can release a patch, giving developers 'zero days' to fix it. is a business decision covers how to make that call with the business rather than for it.

Do this on both nodes. Configuration on an HA pair should synchronize, but confirm it rather than assume it.

Step 2: Preserve evidence before anything restarts

If the crash loop has been running for a while, some evidence is gone. Save what remains before you take any action that triggers another restart.

  • Copy the crash artifacts off the appliance. On NetScaler, administrators were told to keep the core files from the `nsaaad` crashes and to open a case; the same principle applies to any appliance that writes a core or crash report.
  • Copy the authentication and system logs for the window covering the first crash. Look for the unusual request that preceded each crash: an oversized field, a username with encoded content, a burst of requests from one source.
  • Capture the running configuration from both nodes, including any policies an attacker might have added.
  • Note the uptime and crash count of each node so you know how many restarts have already happened.

How to preserve forensic evidence on a network appliance before you patch it goes deeper on this; the key point here is that in a crash loop the window closes with every reboot.

Step 3: Patch in the right order

When a fixed build exists, the order matters.

  1. **Upgrade the standby node first.** It is not serving traffic, so a failed upgrade costs nothing. Verify it boots, that the configuration synchronized, and that the fixed build number is actually what is running.
  2. **Fail over to the upgraded node deliberately.** Now the patched node holds the VIP and absorbs the attack while you work on the other one.
  3. **Upgrade the former primary.** Do not skip this step because the pair is "stable now." An unpatched standby is a time bomb; the next failover, for any reason, hands the VIP to a vulnerable node.
  4. **Confirm both nodes show the fixed build.** Cyberpress's reporting on the NetScaler incident specifically called out verifying builds on both active and standby nodes, because several organizations had patched only the node that was primary at the time.

If the fix requires a configuration change as well as a build, apply it on both nodes and verify synchronization explicitly.

Step 4: Assume compromise until shown otherwise

A crash is sometimes a failed exploit and sometimes a successful one that was followed by a crash. In the October NetScaler case the vendor rated the bug as denial of service, but researchers observed a downloaded binary running on a patched honeypotHoneypot🛡️A decoy system deployed to be attacked so defenders can observe exploitation attempts safely. Honeypot networks give early warning that a vulnerability has moved from theoretical to actively exploited, often before official catalogs like CISA KEV confirm it. and administrators reported download commands in logs before each crash. Treat a crash-looping edge appliance as a potentially compromised one.

  • Check for new files in web-accessible paths, new scheduled tasks or cron entries, and new administrator accounts on both nodes.
  • Check outbound connections from the appliance. A gateway that initiates HTTP downloads to an unfamiliar address is not behaving as a gateway.
  • If the appliance's configuration directory holds credentials and keys, assume they were read if there is any sign of code execution, and rotate them.

The standby node needs the same inspection as the primary. It held the VIP during part of the attack.

Step 5: Fix the runbook

After the incident, change three things in the HA documentation.

**Add a "correlated failure" branch.** The existing runbook probably says "if the primary fails, verify the secondary took over." Add: "if the secondary fails the same way within minutes, this is an attack or a software bug, not hardware; stop the traffic first, then investigate."

**Make evidence preservation a step, not an afterthought.** Put "copy core files and logs off both nodes" before "reboot" in every procedure.

**Record the patch order.** Standby first, failover, former primary, verify both. Write it down so the person on call at 2 a.m. does not have to reason it out.

How a SAML login flow works and where an unauthenticated request gets parsed explains why this particular class of bug keeps landing on authentication gateways. The HA runbook does not need that detail, but the person deciding whether to unbind SAML for a day probably does.