Stabilizing a Multi-Uplink ISP Network with Chronic Congestion
A growing ISP experienced recurring outages, customer dissatisfaction, and escalating operational costs despite continuous infrastructure spend.
Risk
- Revenue erosion from churn
- Brand damage from repeated outages
- Capital expenditure driven by guesswork
- Overdependence on vendors for explanations
A full technical and operational review was conducted to determine whether failures were caused by insufficient infrastructure or systemic design weaknesses. The findings showed the issue was not lack of capacity, but lack of visibility and architectural discipline.
Service stability improved without disproportionate capital spend. Outages became predictable and manageable. Management gained clear insight into where money actually needed to be spent.
Risk was reduced by improving how the network was engineered and operated — not by blindly expanding it.
Symptoms
- Peak-hour performance collapse
- Unstable upstream failover
- Customer complaints inconsistent with utilization metrics
Root causes
- Aggregation oversubscription modeled on averages
- Asymmetric routing during BGP failover
- No flow-level telemetry to validate assumptions
- Congestion occurring upstream of monitored interfaces
Actions
- Introduced NetFlow/IPFIX at aggregation and edge
- Reworked BGP local preference and failover behavior
- Modeled capacity using 95th percentile peak-hour traffic
- Aligned ingress and egress paths to reduce asymmetry
Predictable peak-hour behavior. Non-disruptive upstream failover. Targeted capacity upgrades instead of blanket expansion.