24/7 monitoring and incident response · Kraków

Your systems break at 03:00. Someone should be awake.

We run the checks, take the alert, and start fixing before your first customer notices. No ticket queue, no "we will look at it in the morning" — a named engineer on a phone, every hour of the year.

On duty right now --:--:--
  • 3 engineers — Kraków operations floor · Live
  • 2 engineers — escalation on-call · Standby
  • 1,284 checks running across 41 client systems
  • Last page acknowledged in 71 seconds

Handover at 07:00, 15:00 and 23:00 CET. Nobody works two nights in a row — that is a policy, not a perk.

The first hour

A certificate expires on your payment gateway. Drag the hour.

Same failure, same Tuesday, two companies. The difference is not talent — it is whether anyone was watching. Move the slider and read both columns.

Without anyone watching

Checkout down
  1. T+00The certificate on the payment gateway expires.Nothing on any screen changes. The renewal reminder went to an inbox that belonged to someone who left in March.
  2. T+04First customer email: "my card keeps getting declined".Support treats it as a one-off and suggests trying another card.
  3. T+11Three more emails and a tweet.Support asks in a general Slack channel. It is 03:11. Nobody is in that channel.
  4. T+19Someone awake in another timezone notices and phones the CTO.
  5. T+26An engineer opens a laptop.No dashboard, no starting point, no idea which of forty services is at fault.
  6. T+38Wrong theory: the payment provider must be down.Twelve minutes spent reading a status page that says everything is fine.
  7. T+52Someone finally opens the checkout in a browser and sees the certificate warning.
  8. T+58Certificate renewed by hand. Checkout works again.No note is written. The same reminder will fail again in ninety days.

With Nightshift on duty

Checkout down
  1. T+00The same certificate expires.
  2. T+01Our synthetic checkout fails at step 4 of 6, then fails again.We buy something on your site every three minutes from three cities. One failure never pages anyone — two in a row does.
  3. T+02Page fires naming the step and the expired chain. On-call acknowledges.Median acknowledgement across all clients last quarter: 94 seconds.
  4. T+05The runbook for this service opens automatically.Written with your team in onboarding, rehearsed twice a year.
  5. T+07Your CTO gets one message: what broke, what we are doing, when to expect the next update.One message. Not forty alerts.
  6. T+11Certificate reissued and deployed straight from the runbook.
  7. T+13Synthetic checkout green from all three cities. Status page updated.Total customer-visible damage: thirteen minutes, at 03:00.
  8. T+40Post-incident note in your inbox.Why the reminder failed, and the check we added so an expiring certificate pages us thirty days early instead.

Lost without anyone watching€0

Lost with Nightshift€0

Straight-line arithmetic: minutes down × your hourly revenue. It ignores refunds, support hours and the customers who do not come back — so treat it as the floor, not the bill.

What we watch

Four kinds of check, and only one of them wakes anybody

Most monitoring tells you a server is at 91% memory. That is not an outage. We watch the things your customers actually touch, and we page on those.

01

Synthetic journeys

A robot user that signs in, adds to basket, pays, downloads the invoice. Every three minutes, from three cities. If the business process breaks, we know before your customer writes in.

02

Infrastructure

Hosts, containers, queues, databases, disk and certificate expiry. Thresholds set with your team during onboarding, not copied from a vendor default.

03

Logs and error rates

We alert on the shape of the curve, not on single lines. A tenfold jump in 500s at 04:00 matters. One stack trace does not.

04

The boring expiries

Certificates, domains, API keys, licences, payment tokens. Unglamorous, and the cause of more outages we get called about than any code deploy.

How a page travels

The escalation ladder

Written down, agreed with you before we start, and followed even at four in the morning when it would be easier to guess.

Why this matters more than the tooling: almost every monitoring product can send an alert. Very few companies have decided, in advance and in writing, who picks it up at 03:00 and what they are allowed to do without permission.

  1. 00:00

    The check fails twice

    One failure is noise — a flaky network, a slow third party. Two consecutive failures from separate locations is a signal.

  2. 00:02

    On-call engineer, by phone

    Not an email, not a Slack message. A phone call that keeps ringing until a human presses a key. Unacknowledged after 3 minutes, it goes to the second engineer.

  3. 00:05

    Runbook, then hands on

    Every service you hand us gets a runbook: what it does, what usually breaks, what we may restart, roll back or scale without asking, and what we must never touch.

  4. 00:07

    You hear from us once

    A single message to the person you nominated, with the cause, the plan and the time of the next update. We do not forward raw alerts to clients.

  5. 00:20

    If we cannot fix it, we wake your team

    With a written handover, not a shrug. You get what we tried, what we ruled out and where we stopped.

  6. next day

    A note, and a new check

    Every incident ends with one page of plain writing and at least one new check, so the same failure pages us earlier next time. No blame, no vendor jargon.

Last four quarters

Numbers we would rather be judged on

94 sMedian time to acknowledge a page
17 minMedian time to recovery, all severities
41Client systems under watch
0Nights without someone on the floor

Demo note: Nightshift is an invented company, built to demonstrate a site where the interaction is a timeline rather than a menu. The figures above are written for the demo, not measured.

Pricing

Three sizes, published, no "contact us for a quote"

Monthly, one month notice, no setup fee. Onboarding takes two weeks and is included — most of it is us reading your systems and writing runbooks with your team.

Watch

€900 / month

For teams who have engineers, but nobody awake at night.

  • Up to 10 services
  • Synthetic journeys every 5 minutes
  • Infrastructure and expiry checks
  • We page your on-call, 24/7
  • We do not touch your systems
  • No runbook development
Talk it through
Most clients

Respond

€2,400 / month

For teams who want the phone to ring somewhere else.

  • Up to 25 services
  • Synthetic journeys every 3 minutes, 3 locations
  • Everything in Watch
  • We take the page and act, within an agreed runbook
  • Written note after every incident
  • Quarterly rehearsal of your worst case
Book 20 minutes

Own it

From €5,800 / month

For regulated companies who need the paperwork as well as the phone call.

  • Unlimited services
  • Everything in Respond
  • Named engineers, introduced to your team
  • NIS2 and DORA incident reporting timelines
  • Evidence pack your auditor will accept
  • 15-minute response written into the contract
Talk it through

Full comparison, and what is deliberately not included →

Field notes

Written on the floor, between pages

No thought leadership. Things we learned the hard way, at an hour when nobody was reading.

All field notes →

Twenty minutes, and you will know if this is for you

No slide deck. We ask what breaks, who gets called, and what happened the last time it broke at night. If you already have that covered, we will say so.

Book the call