Skip to content
ISP Operations

NOC & Network Operations

A NOC is not a room with screens. It is a discipline: knowing what the network is doing, catching faults before customers do, and handling every incident the same way every time.

The NOC is where an ISP either runs its network or reacts to it.

01 — Why It Matters

Why NOC Capability Decides Service Quality

Most small and mid-sized ISPs do not lack monitoring tools — they lack a working NOC practice. Dashboards exist but nobody trusts them. Alerts fire constantly, so everyone ignores them. Outages are discovered by customer calls, escalations travel by personal WhatsApp, and changes go into the network at peak hour with no record and no rollback plan.

The result is predictable: long outages, repeat faults, engineers who firefight instead of engineer, and a management team that cannot answer basic questions about availability or capacity.

Our own career started on the NOC floor — field work, support, and NOC escalations — before moving into core network operations. We build NOC capability from the operator side, not the vendor brochure side: the goal is a team that detects, decides, and resolves, not a bigger wall of screens.

02 — Monitoring

Monitoring That Engineers Actually Trust

An NMS is only effective if it measures the right things and alarms only when a human must act.

What to Measure

  • Full device inventory and discovery — no unmonitored gear
  • Link utilisation, errors, and discards on every core and uplink port
  • Latency, jitter, and packet loss to upstreams and key destinations
  • Service-level checks (DNS, RADIUS, billing, portals), not just ping
  • Power, environment, and backhaul health at remote sites
  • Configuration backup and change tracking on every node

Alarm Design & Alert-Fatigue Engineering

  • Severity levels with defined responses — not one red for everything
  • Dependency-aware suppression: one upstream failure, one alarm
  • Thresholds set from baselines, then reviewed and tuned
  • Flap damping and hold-down timers on noisy sources
  • Escalation when alarms are unacknowledged, not just re-notification
  • A standing rule: every alert leads to an action or gets redesigned

Flow-Level Telemetry

  • NetFlow/IPFIX export from core and edge routers
  • Top-talker and per-customer traffic visibility
  • Traffic mix analysis for peering, caching, and upstream decisions
  • DDoS and anomaly detection from flow baselines
  • Before/after flow evidence for routing and QoS changes

A small, tuned monitoring stack the team trusts beats a sprawling one it ignores.

03 — Discipline

Operational Discipline: Incidents, Tickets, Changes, Runbooks

Tools detect; discipline resolves. These four practices are what separate a NOC from a helpdesk with graphs.

Incident Handling & Escalation

A defined path from detection to resolution, with named tiers and time-bound escalation.

What We Put in Place

  • Incident severity classification tied to customer impact
  • Tiered escalation paths — who is called, when, with what information
  • On-call rotations and handover routines between shifts
  • Major-incident roles: one person leads, one communicates
  • Post-incident reviews focused on cause and prevention, not blame

What It Prevents

  • Outages resolved by process, not by heroics
  • No incident stalls because "the one engineer" is unreachable
  • Repeat faults get engineered out instead of re-fixed

Ticketing Discipline

If it is not in a ticket, it did not happen. The ticket queue is the NOC’s memory and its accountability record.

What We Put in Place

  • Every incident, request, and task tracked — including NOC-detected faults
  • Automatic ticket creation from monitoring alarms
  • Root-cause categorisation so recurring problems become visible
  • SLA clocks and breach escalation on every queue
  • Linkage between tickets, devices, and customers

What It Prevents

  • No customer issue silently dropped
  • Management sees real workload and real problem patterns
  • Fault history survives staff turnover

Change Management

Most self-inflicted outages are unrecorded changes. Change discipline is the cheapest availability upgrade an ISP can buy.

What We Put in Place

  • Change requests with stated purpose, risk, and rollback plan
  • Maintenance windows and customer notification routines
  • Peer review for changes touching core and shared infrastructure
  • Configuration backup before and after every change
  • Post-change verification against monitoring and flow data

What It Prevents

  • Fewer outages caused by the ISP’s own engineers
  • "What changed?" answered in minutes, not days
  • Auditable history for regulators and enterprise customers

Runbooks

Written procedures for known situations, so the 2 a.m. response matches the 2 p.m. response.

What We Put in Place

  • Runbooks for the top recurring faults and routine operations
  • Escalation contact trees kept current and tested
  • Diagnostic checklists per platform and per service
  • Runbooks updated as an output of post-incident reviews

What It Prevents

  • Junior engineers handle more without waking seniors
  • Consistent resolution quality across shifts
  • Tribal knowledge captured before people leave
An alarm nobody acts on is noise. A ticket nobody owns is a complaint waiting to escalate.

05 — Capacity & Team

Capacity Planning & Building the NOC Team

The same telemetry that runs the NOC should drive upgrade decisions — and the team must be able to run it after we leave.

Capacity Planning From Utilisation Data

  • 95th-percentile utilisation tracking on every core link and upstream
  • Growth trending per link, per POP, and per service
  • Defined upgrade thresholds with procurement lead time built in
  • Flow data informing peering, caching, and transit decisions
  • Busy-hour analysis instead of daily averages

NOC Team Training & Capability Building

  • Role and shift structure sized to the operator, not to a template
  • Hands-on training on the monitoring, ticketing, and escalation stack
  • Runbook walk-throughs and incident simulation exercises
  • Escalation practice: what to try, what to record, when to call
  • Follow-up support while the team takes ownership

A NOC that depends on one hero engineer — or on an outside consultant — is a single point of failure.

06 — Experience

Grounded in Real NOC Work

This practice began in the NOC, not in a slide deck. Our founder started as a NOC technician at EdgeNet (2005–07), progressing from field work through support to NOC escalations. At InterSAT (2007–09) he managed and secured the core network and servers, handled escalations from CRM, NOC, and support teams, and coordinated with RF teams across Nairobi, London, and Washington hubs. At KDN (2009–11) he supported a team of IP and system engineers handling NOC and CRM escalations, and later ran ground-monitoring operations for Avanti Communications’ satellite fleet across Kenya and Tanzania.

More recently, through HCS: for DSI in the DRC we deployed network planning, monitoring, and ticketing systems for both internal and customer-facing operations, and trained the technical and business teams that run them. At SkyTrend (2021) we deployed the redesigned systems, monitored and optimised them in production, and trained and supported the technical team.

That trajectory — from answering the escalation to designing the escalation path — is exactly what building NOC capability requires.

07 — Engagement

How We Build NOC Capability

  1. 01

    Phase 0 Review

    Assess monitoring and NMS effectiveness, NOC workflows and escalation paths, ticketing and incident handling, and change management discipline

  2. 02

    Design

    Define the monitoring stack, alarm policy, escalation tiers, and change process that fit your scale

  3. 03

    Implement

    Deploy and integrate NMS, flow telemetry, and ticketing; tune alarms against real baselines

  4. 04

    Train & Hand Over

    Runbooks, shift structure, incident drills, and follow-up support until the team owns it

Every engagement starts with the Phase 0 review — we do not recommend tooling before we have seen how your NOC actually works today.

Field note — 2026

The 2026 NOC conversation is dominated by AIOps — vendors promise alert volumes cut by 90%+ through event correlation, and streaming telemetry over gRPC is displacing 5-minute SNMP polling in large operators. The reality check: correlation engines only work on top of clean inventory, sane alarm design, and disciplined ticketing, which is precisely what most African ISPs are missing — the skills gap is real enough that ITU, APC, and AFRALTI launched a dedicated network-manager training programme for the region in 2025. Our advice is unchanged: tune your alarms, wire flows into your capacity decisions, and write your runbooks first — an AI layer on top of a noisy, undocumented NOC just automates the confusion.

Assess Your NOC Capability

Start with a focused review of your monitoring, escalation paths, ticketing, and change discipline.