Why URL Encoding Lets Attackers Slip Past String-Matching WAF Rules
🛡️ Security Intermediate 5 min read

Why URL Encoding Lets Attackers Slip Past String-Matching WAF Rules

A firewall that matches a literal path is defeated by a request spelling the same path differently. How percent-encoding creates a parser mismatch, and why you should normalize before deciding.

Published: September 27, 2026 • Updated: September 27, 2026
wafurl encodingweb securityinput validationevasion

A web application firewallFirewall🌐Security system that monitors and controls network traffic based on predetermined rules. sits in front of your application and inspects requests, blocking the ones that match known-bad patterns. It is a genuinely useful control. It is also, when configured to match literal strings, one of the easiest defenses in the stack to walk past, because the attacker and the firewall almost never read the same request the same way. The renewed ShinyHunters campaign against Oracle PeopleSoft made this concrete: a rule blocking the path `/PSEMHUB/` was defeated by requesting `/%50SEMHUB/`, where `%50` is simply the letter P written in percent-encoded form. This piece explains why that works, and why the fix is a principle rather than another string.

Two Readers, One Request

Every HTTP request passes through several components before it reaches your application code: a load balancer, a reverse proxyReverse Proxy🛡️A server that sits in front of one or more backend services, terminating client connections and forwarding requests to the backend. It is the standard place to add authentication, TLS and access control to a service that lacks its own, without modifying the application., a WAF, the web server, the application server, and finally the framework's routing layer. Each of them parses the request, and each is free to interpret it slightly differently. Security researchers call this a parser differentialParser Differential🛡️A security weakness that arises when two components in a request path interpret the same input differently, such as a firewall matching a raw path while the application server decodes it first. Attackers exploit the gap between the two readings to route a blocked request to a vulnerable endpoint., and it is the root of a whole family of evasion techniques.

The critical detail is *when* each component decodes the URL. A path like `/%50SEMHUB/hub` contains a percent-encoded byte. `%50` is the hexadecimal for the ASCII code of the capital letter P, so after decoding, the path is `/PSEMHUB/hub`. A WAF rule that inspects the raw, still-encoded path sees the literal characters slash-percent-five-zero and does not match a block on `/PSEMHUB/`. The application server, meanwhile, decodes the path as a matter of course and routes the request to exactly the endpoint the WAF meant to protect. Two readers, one request, two different meanings, and the gap between them is where the exploitExploit🛡️Code or technique that takes advantage of a vulnerability to cause unintended behavior, such as gaining unauthorized access. lives.

Encoding Is Ambiguous by Design

Percent-encoding exists so that URLs can carry characters that would otherwise be unsafe or reserved. The specification permits encoding characters that do not strictly need it, which means a single logical path has many valid textual representations. `/PSEMHUB/`, `/%50SEMHUB/`, `/%50%53EMHUB/`, and mixed-case variants where the platform treats the path case-insensitively can all resolve to the same resource. A firewall that blocks one exact spelling has blocked one of dozens of doorways into the same room.

Percent-encoding is only the most common trick. The same principle powers double URL-encoding, where `%2570` decodes once to `%50` and again to `P`; overlong UTF-8 sequences; alternate path separators; and directory-traversal obfuscation. In every case the attacker relies on a downstream component performing a decoding or normalization step that the security control did not anticipate. This is closely related to the path-manipulation problems behind directory traversal, which we cover in how path traversalPath Traversal🛡️A web vulnerability (CWE-22) where user-supplied input in a file path escapes the directory the application intended to serve from, typically via parent-directory references, letting an attacker read or write files elsewhere on the server. bugs let attackers read files outside the web root.

Why Blocklists Lose

The deeper problem is that string matching is a blocklist, and blocklists enumerate badness. To be correct, a blocklist must anticipate every representation of every attack, forever. The attacker only has to find one representation you missed. That asymmetry is why the ShinyHunters operators needed to change exactly one character. The rule was not wrong about the danger; it was wrong to assume the danger would always be spelled the same way.

This is the same reasoning that makes input validation by allowlist stronger than by denylist, and it applies far beyond firewalls. Any control that makes a security decision on a string before that string is fully normalized is vulnerable to the same class of bypass.

The Fix: Decide on the Normalized Form

The durable defense is to make the security decision on the same canonical form the application will act on. In practice that means normalizing the request before you match against it: fully URL-decode the path, resolve case folding if the target platform is case-insensitive, collapse redundant separators and traversal sequences, and only then evaluate your rule against the result. Modern WAFs support this as a transformation or normalization step; the failure in the PeopleSoft incidents was that the deployed rules matched raw input.

Better still, do not rely on matching the malicious pattern at all. Where the goal is to keep an endpoint off the internet, block by allowlisting the routes and clients that are permitted and denying everything else, or remove the exposure at the source by disabling the vulnerable component. Deciding what is allowed to reach a resource is far more robust than trying to enumerate every way an attacker might spell an attack against it. When you must run a firewall rule as a temporary measure, understand its limits, a theme we develop in why a WAF rule is a stopgap, not a substitute for patching.

Turning the Evasion Into a Detection

There is an upside to understanding this technique: the evasion attempt is itself a strong signal. Legitimate clients and browsers have no reason to request `/%50SEMHUB/` when `/PSEMHUB/` works identically. A request that arrives pre-obfuscated is almost always automated and almost always hostile. If your logging captures the raw path as well as the normalized one, a mismatch between the two, or the mere presence of encoded characters in a path segment that is normally plain text, is worth an alert. That detection strategy, applied to server logs, is the subject of how to hunt your web server logs for path-normalization evasion.

The Takeaway

The `/%50SEMHUB/` bypass was not clever in a technical sense; it was clever in exploiting an assumption. The assumption was that a request means what the firewall sees, when in fact a request means what the application ultimately does with it. Close the gap by normalizing before you decide, prefer allowlists to blocklists, and treat obfuscated input as evidence of intent rather than noise. Do that, and a one-byte change stops being a skeleton key.