The service

Monitoring is easy. Deciding what deserves a phone call is the job.

Any tool can tell you a disk is 80% full. The work is knowing which of your forty services matters at four in the morning, what may be restarted without asking, and who to wake when it cannot be.

Onboarding

Two weeks, and most of it is us listening

We do not install an agent on Monday and call it monitoring. The first fortnight is spent finding out how your systems actually fail, which is rarely how the architecture diagram suggests.

Included in every plan. No setup fee, and if we get to the end of week two and think you do not need us, we say so and you owe nothing.

  1. Day 1–2

    Inventory

    Every service, host, domain, certificate and third-party dependency, written down in one place. Half our clients have never had this list.

  2. Day 3–5

    What breaks, historically

    We read a year of your incidents and ask the awkward question: what actually went wrong last time, and how did you find out?

  3. Day 6–8

    Journeys and thresholds

    We script the three or four things a customer must be able to do, and set thresholds with your engineers rather than from a vendor default.

  4. Day 9–12

    Runbooks

    One page per service: what it does, what usually breaks, what we may restart, roll back or scale on our own, and what we must never touch without you.

  5. Day 13

    We break it on purpose

    In your staging environment, with your team watching. A rehearsal nobody enjoys and everybody remembers.

  6. Day 14

    Live, quietly

    First month we page ourselves and copy you, so you can see what we would have woken you for before we start doing it.

The checks

Four kinds, and only one of them wakes anybody

01 · pages on failure

Synthetic journeys

A scripted robot user that does what your customer does: signs in, searches, adds to basket, pays with a test card, downloads the invoice. Every three minutes, from Frankfurt, Warsaw and Dublin.

Because it runs the whole business process, it catches the failures that individual service checks miss entirely — an expired certificate on step four, a payment provider silently rejecting one card type, a session that dies after a deploy.

This is the only check type allowed to ring a phone at night, and only after it fails twice from two separate locations.

02 · pages on threshold

Infrastructure

Hosts, containers, queues, replication lag, disk, memory, connection pools. Standard stuff, with one difference: your thresholds are set during onboarding by someone who asked what normal looks like on a Monday morning versus a Saturday night.

Most of these produce a ticket, not a phone call. A queue backing up at 10:00 is a conversation. The same queue backing up at 03:00 with a growing customer impact is a page.

03 · pages on shape

Logs and error rates

We do not alert on single lines — that way lies four hundred notifications a day and a team that has stopped looking. We alert on the shape of the curve.

A tenfold jump in 500s inside five minutes matters. One stack trace does not. A new error string that has never appeared before matters more than a familiar one appearing again.

04 · pages 30 days early

The boring expiries

TLS certificates, domain renewals, API keys, OAuth secrets, payment tokens, licence files, and the credit card on the account that pays for all of it.

Deeply unglamorous, and the cause of more outages we are called about than any code deploy. These page thirty days ahead, again at seven days, and to a human being rather than a shared inbox.

Boundaries

What we will not do

Every one of these has been asked for, and refusing them is why the service works. If you need them, you need a different supplier, and we will happily name two.

  1. No

    We do not write your features

    We keep things running. The moment we start shipping product code we stop being the people who can be objective about why it broke.

  2. No

    We do not touch anything outside the runbook

    Even when we are fairly sure we know the fix. Fairly sure at 03:40 is how a thirteen-minute outage becomes a six-hour one.

  3. No

    We do not forward raw alerts to clients

    You get one message written by a person. If we are just relaying what the tool said, you are paying us for a webhook.

  4. No

    We do not take a client whose only problem is cost

    If you are looking to replace an on-call rota purely to save money, and the rota works, keep it. We will tell you that on the first call.

Service levels

What is written into the contract

Severity is agreed with you during onboarding and reviewed every quarter. The definitions below are the defaults; yours may differ, and the contract carries yours, not these.

Sev 1 — customers cannot do the main thing

Checkout down, login down, the product unusable. Acknowledged within 5 minutes, 24/7. Work starts immediately. You hear from a person within 10 minutes and every 30 minutes after that until it is closed.

Sev 2 — degraded, with a workaround

Slow but working, one payment method failing, a non-critical integration down. Acknowledged within 15 minutes, 24/7. You hear from us within the hour.

Sev 3 — something is wrong, nobody is hurting yet

Disk trending full, replication lag growing, a certificate expiring in three weeks. Handled during Kraków working hours, in your ticket queue, with a note.

What happens if we miss it

Miss the acknowledgement time on a Sev 1 and that month is free. No forms, no claim process — it comes off the next invoice automatically, because we are the ones holding the timestamps.

The best time to sort this out is a Tuesday afternoon

Not at 03:00 with the site down and nobody answering. Twenty minutes now, and you will at least know where the gaps are.

Book the call