Networking

Why Most SD-WAN Designs Fail at the Overlay Policy Stage

The overlay is what you are actually buying in an SD-WAN project. It is also where multi-site deployments quietly fall apart. Here are the failure modes we see most often, and the design decisions that separate a network that performs from an expensive set of routers.

E7
Edge7 Networks Team
SD-WAN specialists since 2018
15 September 2026
8 min read
Share

The hardware is racked. The circuits are lit. The underlay is converged and passing traffic. Then someone opens the policy editor, and the project stalls.

This is the pattern we see across multi-site SD-WAN deployments. Organisations invest months in procurement, circuit provisioning, and CPE deployment. The overlay policy design gets compressed into a week. That is where a project that looked on track starts to unravel.

If you are mid-way through an SD-WAN decision, this is the part worth slowing down on. The boxes and the bandwidth are the easy half. The overlay is what you are actually buying, and it is where the difference between a network that performs and one that merely passes traffic is decided.

The short version

In a multi-site SD-WAN project, the hardware and circuits are the easy part. The overlay, the policy layer that decides which application takes which path and what happens when a link degrades, is where deployments succeed or fail. Most designs are rushed at exactly this stage. Below are the failure modes we see most often, and a set of questions to ask before you sign, so you can tell a design built for the demo from one built for production.

The vendor demo problem

Every SD-WAN vendor has a compelling three-site demo. Hub, spoke, spoke. Two transport types. Three application categories. The policies fit on a single screen, and failover works perfectly because there is nothing complex enough to break.

Production looks different. A typical multi-site deployment involves 80 to 200 sites across multiple regions, 3 to 4 transport types per site, more than 40 application categories, and a mix of hub-and-spoke and regional mesh topologies. The policy matrix does not scale linearly. It scales exponentially, because every site type, transport combination, and application category introduces new decision points.

The overlay is where the intelligence lives. Get it wrong, and you have an expensive collection of routers passing traffic with no more sophistication than the MPLS network you replaced.

The disjoint underlay problem

One of the most common failure modes we encounter is the disjoint underlay. This happens when your MPLS and internet transports are not directly connected at the branch. The branch CPE has a connection to the MPLS cloud and a separate connection to the internet, but those two paths do not interconnect locally.

In this scenario, branch-to-branch traffic that needs to cross between transports has to pass through a gateway appliance that bridges both. Without the right overlay reachability rules, the fabric can advertise a path over a transport that the remote branch has no way to reach. The routing says the path exists. The data plane says otherwise. Traffic is blackholed, and the dashboard still shows the tunnel as healthy, because the tunnel between the appliances is fine.

The fix is overlay policy that only advertises a path to the transports that can actually carry it, defined per site. This is not default behaviour on any platform we have worked with. It takes deliberate design, and an understanding of the physical topology detailed enough to know which sites cannot bridge their transports locally. In a 200-site estate, that information is rarely documented accurately.

For the decision

The point is simpler than the mechanism: a healthy dashboard is not proof of a working network. If you are evaluating a design rather than building it, ask how path health is verified end to end, not just whether the tunnel reports as up.

ECMP across gateways: the 50% blackhole

Gateway redundancy is a standard design element. Two or more gateway routers provide path diversity and load distribution for traffic transiting between transport domains. ECMP distributes flows across available gateways, and under normal conditions this works well.

The failure scenario is specific but common. One gateway loses its internet link while keeping its MPLS connectivity. As far as the overlay is concerned the gateway is still up, and it keeps advertising itself as a viable path. ECMP continues to distribute roughly 50% of internet-bound flows towards it. Those flows arrive at the gateway and have nowhere to go.

The gateway is not "down" in any control-plane sense. It is partially reachable. This is exactly the kind of degraded state that default configurations handle poorly. Detecting it takes path-aware health monitoring at the overlay level, combined with policy that stops steering traffic to a gateway over a transport that has failed, even while the gateway itself is still up. We have seen this exact scenario cause intermittent 50% packet loss in production for days before the root cause was identified.

For the decision

This is the failure a buyer never sees in a demo and lives with in production. The question that surfaces it early is a specific one: how does the platform detect a gateway that is half-working, and how quickly does it stop sending traffic to it.

Application-aware routing at scale

Application-aware routing is the headline feature of every SD-WAN platform. Define SLA thresholds per application, and the overlay steers traffic to the best-performing path. In a demo, this is elegant. In production, it is where policy complexity reaches its peak.

Consider a deployment with four transport types: MPLS, broadband internet, dedicated internet access, and 4G/5G backup. You have more than 40 application categories, each with different latency, jitter, and loss tolerances. Some categories have sub-categories with different requirements. Voice and video both fall under "real-time" but have different jitter sensitivity profiles.

Now layer in site-specific exceptions. Your contact centre sites need voice pinned to MPLS regardless of SLA metrics, because the MPLS provider guarantees end-to-end QoS that the internet path cannot match. Your retail sites need PCI traffic on a specific transport for compliance reasons. Your development offices have relaxed requirements but higher bandwidth needs.

The policy set for this environment can run to hundreds of rules. Each rule interacts with every other rule through precedence and inheritance. Testing the full matrix is not something you do in an afternoon.

QoS across mixed transports

MPLS circuits carry QoS markings end to end. Your provider honours DSCP values, and traffic is queued and scheduled according to the service class you have purchased. Internet circuits do not. DSCP markings are typically stripped at the first provider hop.

The SD-WAN overlay has to manage application prioritisation across both transport types at once. On the MPLS path, the overlay needs to mark traffic correctly so the provider QoS policy applies. On the internet path, the overlay itself must perform the queuing and scheduling, because there is no provider QoS to rely on.

This dual-mode QoS requirement means your policy has to define per-application behaviour differently depending on the selected transport. A single QoS policy applied uniformly across all transports either under-serves MPLS, by not using provider QoS, or over-promises on internet, by assuming QoS guarantees that do not exist.

For the decision

In commercial terms, a uniform policy either wastes the MPLS service class you are paying for or promises performance on internet circuits that carry no such guarantee. It is worth confirming the design treats the two transports differently, because your invoice already assumes it does.

The local breakout trap

Direct internet access at the branch is one of the main drivers for SD-WAN adoption. Rather than backhauling all internet traffic to a central data centre, SaaS and web traffic breaks out locally. This reduces latency and offloads the central internet gateway.

The risk is in the classification. Traffic that matches a local breakout rule bypasses VPN encryption by design. It exits the branch router directly onto the local internet circuit, unencrypted and outside the overlay. If your application classification misidentifies traffic, sensitive data leaves the branch without the protection of the SD-WAN tunnel.

This is not a theoretical concern. Application classification engines rely on DNS resolution, SNI inspection, and IP address databases. A miscategorised application, a CDN IP range that shifts, or a DNS response that resolves differently at different branches can all cause traffic to match the wrong policy. EdgeConnect classifies applications on the first packet with First-Packet iQ, which is what makes accurate local breakout viable, but it still depends on the classification data staying current and the breakout rules being scoped tightly. Regular auditing of local breakout traffic is essential, and the policy design needs to err on the side of caution. When in doubt, tunnel it.

Where the design time actually goes

SD-WAN has been Edge7 Networks' core discipline since 2018, and we are one of the established SD-WAN specialists in the Irish and UK market. We know the major SD-WAN policy engines at the level of detail these failure modes demand, because we have spent years designing, migrating, and running them in production.

The pattern is consistent across engagements. Hardware deployment takes a predictable amount of time. Circuit provisioning is largely a waiting game. The policy design phase is where we spend the most time with customers, and it is where most other providers rush through. We have picked up projects mid-deployment where the hardware was installed months earlier but the overlay was never properly designed. The sites were "up". The network was not performing.

Our approach separates the design into discrete phases. We map every application first, identifying each category and its transport requirements, not just the top ten. We define site archetypes, such as hub, regional hub, large branch, small branch, retail, and contact centre, and build a Business Intent Overlay for each, so path selection and failover are defined once per archetype rather than box by box. We then apply site-specific exceptions on top and test the full matrix in a staged rollout. In a recent 200-site EdgeConnect migration, the decisions that took the most time were not about hardware or circuits. They were the path-steering thresholds, and the Path Conditioning and Tunnel Bonding settings, that decide what happens when a link is degraded rather than down.

This is closely tied to underlay design. If your dual underlay architecture has diversity gaps or capacity mismatches, the best overlay policy in the world cannot compensate. We cover that side of the problem in a companion piece on building a resilient WAN for a multi-site organisation.

"The overlay is the product. Everything else is plumbing. Treat the policy design with the engineering rigour it deserves."

Edge7 Networks engineering team

The practical takeaway

Policy design needs to happen before deployment, not during. If your project plan shows "configure policies" as a two-day task after hardware installation, the plan is wrong.

Map your applications, all of them. Define transport preferences per site type, not per site. Build a policy hierarchy that uses inheritance, so an individual site exception does not require rewriting the entire rule set.

Test failover under realistic traffic conditions. A ping test across a backup path tells you nothing about how 40 applications behave when 200 branches fail over at once. And design for the degraded state, not just the outage. A circuit running at 8% packet loss is technically available and practically useless for half your applications. Your policies need to recognise that and respond before users do.

Document every exception and the reason behind it. Six months from now, when someone needs to modify a policy, they need to understand why it was designed that way.

What to ask before you sign

If you are the one signing off the project rather than configuring it, these are the questions that separate a design built for the demo from one built for production. They are worth putting to any vendor, and to your own team.

  • When a link is degraded rather than down, running at 5 to 8% packet loss or high jitter, does the overlay steer away from it before users notice, and what thresholds trigger that?
  • How is path health verified end to end, beyond a tunnel reporting up or down?
  • Which applications are pinned to which transport, and is that decision documented per site type rather than configured box by box?
  • What happens to capacity and application priority when the larger circuit fails, not only when a circuit goes down completely?
  • How is local breakout traffic classified and audited, so sensitive data never leaves a branch outside the tunnel by mistake?

The quality of the answers tells you more about whether a network will perform than the price per site on the quote.

E7
Edge7 Networks Team
Networking & Security Specialists, Ireland & UK

Edge7 Networks is a specialist networking and security provider, founded in 2018. We design and run multi-site SD-WAN, LAN, and WiFi alongside managed security and compliance for IT leaders across Ireland and the UK, as an HPE Aruba Networking Gold Partner. We hold ISO 27001:2022, ISO 9001:2015, and Cyber Essentials certifications.

Scoping an SD-WAN project?

Whether you are choosing a platform or you have inherited a deployment where the sites are up but the network is not performing, we are happy to talk through the overlay design for your environment. No pitch, just a conversation with the engineers who do this work.