NEW CASE
Antana

We create digital solutions that work for businesses


Give us a call +38 (066) 35-14-529

Let's take the first step towards your website — write to us

Close
BB STUDIO 12 min read

Website Uptime Monitoring: How to Configure Checks, Alerts and an SLA

Hosting & tech
Website Uptime Monitoring: How to Configure Checks, Alerts and an SLA

A website can stop accepting leads overnight, return a blank page after an update or fail only for visitors in one region. If customers are the first to report the problem, the business has already lost traffic, advertising spend and trust.

Website availability monitoring sends automated requests, validates the expected response and alerts the responsible people when a condition fails. A five-minute homepage check is useful, but it is only a baseline. A dependable setup distinguishes a network failure from a slow response, a server error, incorrect content and a broken business transaction.

Short answer

For a typical business website, check the homepage and at least one critical landing page every one to five minutes from multiple locations. Success should require the expected HTTP status, acceptable response time, a stable content marker and a valid TLS certificate.

After two or three consecutive failures, notify an owner through at least two independent channels. For an ecommerce site, separately test the catalogue, cart and a safe checkout path with no real charge. Every alert needs an owner, a runbook and a response target.

What uptime means — and why 100% explains very little

Uptime is the share of a measurement period during which a resource meets the defined availability conditions:

uptime = successful time / total measured time × 100%.

The definition of “successful” determines the value of the number. A monitor that accepts any HTTP response can mark a maintenance page returning 200 as healthy. A check from one country will miss some regional DNS or CDN failures. A 30-minute interval may miss short but recurring outages.

Always read an uptime percentage together with its method: monitored URLs, probe regions, interval, success criteria and excluded periods. Reliable website hosting reduces infrastructure risk, but an external monitor is still necessary because it observes the website from the visitor’s side.

Five layers of website monitoring

1. DNS and network path

The domain must resolve to the expected address and the connection must complete within a reasonable time. A failure can happen before the request reaches the web server: an expired domain, removed record, unavailable authoritative nameserver or stale regional configuration can all make the site unreachable.

After a change, verify the relevant A, AAAA, CNAME, MX and TXT DNS records and keep probes active in several regions until propagation is complete.

2. HTTP status and response time

A basic monitor sends a GET or HEAD request to a specific URL. A normal page usually expects 200; a deliberate permanent redirect can expect 301 or 308. Codes such as 500, 502, 503 and 504 indicate server-side failure, but not every broken page returns a 5xx response.

Track latency as well. A page may be technically online while taking 15 seconds to respond, which still wastes paid clicks and drives visitors away. Use the separate guide to TTFB, Core Web Vitals and website speed testing to diagnose performance rather than availability alone.

3. Content validation

Search the response for a stable marker: the company name, a product heading or a technical string rendered only after the required data loads. This catches soft failures where the server returns 200, but the template is empty, the database is unavailable or a registrar parking page replaced the site.

Do not monitor copy that a content editor may change without warning. Add a stable HTML marker or expose a dedicated health endpoint that reveals no secrets.

4. TLS certificate and domain expiry

Monitor the certificate’s remaining validity, hostname coverage and chain. Automatic renewal can fail after a DNS, permission or reverse-proxy change. The SSL certificate and HTTPS troubleshooting guide explains the checks in more detail.

Track domain renewal separately. A valid certificate will not save a site after registration or nameserver settings expire. For a broader inventory, follow the domain, hosting and server checklist.

5. Synthetic transactions

An ecommerce or web application monitor should do more than open the homepage. A synthetic script can search for a test product, add it to the cart, open checkout and verify the expected API response. A service website can open a lead form and send a marked test submission to a technical destination.

Synthetic checks must not create real purchases, charge cards or pollute the CRM. Use a test product, a dedicated address, a clear service marker and automatic cleanup.

What to monitor

Target Success condition Typical interval Priority
Homepage 200, content marker, normal response time 1–5 min high
Critical landing page 200, form or CTA present 1–5 min high
Catalogue or search valid, non-empty result 5 min high for ecommerce
Cart / checkout safe test journey succeeds 5–15 min critical for ecommerce
API / webhook valid response to test request 1–5 min based on business impact
TLS certificate valid with enough time left 6–24 h high
DNS expected answers from several resolvers 15–60 min medium
Domain sufficient time before expiry daily high
Backup job job finished and recovery is verified after every run critical

There is no need to hit every page every minute. Select a few endpoints that represent the real customer journey and run broader crawls less often.

Choosing intervals and probe locations

One-minute monitoring detects incidents faster but produces more requests and is more sensitive to short network interruptions. A five-minute interval is often sufficient for a small site with no major paid campaign. A checkout, public API or high-spend landing page usually justifies one-minute checks.

A single probe cannot distinguish a website outage from a problem at the monitoring location. A practical sequence is:

  1. retry immediately after the first failure;
  2. confirm from a second region;
  3. alert after two or three consecutive failures;
  4. use a lower threshold for a critical endpoint;
  5. send a recovery notification when service returns.

This reduces noise without hiding a sustained incident.

Reducing false alerts

The worst monitoring system is the one everybody learns to ignore. False positives come from brief network interruptions, blocked probe IPs, unrealistic latency limits, maintenance windows or fragile browser scripts.

For every check, define:

  • expected status and content marker;
  • acceptable response time;
  • consecutive failures before an alert;
  • confirmation from another region;
  • planned maintenance windows;
  • severity and business impact;
  • recovery condition;
  • escalation timeout.

Avoid unconditional allowlisting that lets a probe bypass the protections faced by real customers. Keep the route realistic, but mark synthetic requests clearly in the User-Agent and logs.

Alerts and escalation

Email is suitable for reports but weak as the only critical channel. A checkout outage should also reach a messenger, push or phone path. At the same time, one slow request should not wake the whole team.

Severity Example Initial response Channels
P1 critical site or payment unavailable 5–10 min messenger + phone
P2 high form, catalogue or API broken 30 min messenger + email
P3 medium slow site, TLS expiry approaching business hours email / task tracker
P4 informational one brief failure review in report event log

Include the URL, incident start, locations, status, latency, latest deployment and runbook link in the alert. Ongoing website support and maintenance becomes valuable when there is both a signal and a person with access and responsibility to recover the service.

SLI, SLO and SLA

An SLI is the measured indicator, such as the ratio of successful checkout requests. An SLO is the team’s internal target, such as 99.9% monthly availability. An SLA is an agreement with a customer or provider that defines service levels, measurement rules, exclusions and consequences.

An SLA should not be reduced to a percentage. It should define the monitored component, time zone, calculation period, maintenance windows, third-party dependencies, incident start, response time and remedy.

The downtime hidden behind the nines

Availability Maximum in 30 days Maximum in 365 days
99% 7 h 12 min 3 d 15 h 36 min
99.5% 3 h 36 min 1 d 19 h 48 min
99.9% 43 min 12 sec 8 h 45 min 36 sec
99.95% 21 min 36 sec 4 h 22 min 48 sec
99.99% 4 min 19 sec 52 min 34 sec

These are mathematical values before contractual exclusions. If planned maintenance, third-party failures or attacks are excluded, the user experience can be worse than the reported SLA.

What to do when an alert fires

Monitoring without a response plan only reports failure faster. Keep a short runbook:

  1. Confirm from an independent network and another region.
  2. Identify the layer: DNS, CDN, TLS, server, CMS, database, third-party API or feature.
  3. Review the latest deployment, update, import and configuration change.
  4. Roll back the latest change or redirect traffic when safe.
  5. Preserve symptoms and logs before restarting components.
  6. Tell the business what is affected and when the next update will arrive.
  7. After recovery, test forms, payment, email, CRM and analytics.
  8. Record the cause, duration, impact and preventive action.

If compromise is suspected, follow the website malware-checking checklist. If recovery needs a rollback, use a tested website backup and restore process, not the first archive you find.

Set it up in one working day

Step 1. Define critical journeys

List what creates revenue: a landing page, lead form, phone number, catalogue, cart, payment, account area or API. Estimate the impact of 15 minutes, one hour and one day of downtime.

Step 2. Create baseline checks

Add the homepage, one commercial page and one critical endpoint. Validate status, content, TLS and latency from two or three geographic locations.

Step 3. Configure thresholds

For a normal page, start with a five-minute interval and three failures. For checkout, start with one minute and two failures. Base performance limits on your stable baseline rather than an arbitrary ideal.

Step 4. Route notifications

Assign a primary and backup owner. Test email, messenger and phone channels with a controlled alert. Keep credentials in a protected vault, not in the alert message.

Step 5. Run a drill

Change the expected marker on a test check or use a safe nonexistent URL. Measure detection, acknowledgement and diagnosis time. Restore the correct condition immediately after the drill.

Common mistakes

  • monitoring only the homepage;
  • accepting any HTTP status as success;
  • never validating page content;
  • checking from one region only;
  • alerting one person by email;
  • sending no recovery notice;
  • ignoring planned maintenance;
  • running the monitor on the same server it checks;
  • never testing notification channels;
  • reporting uptime without measurement rules;
  • closing incidents without preventive action.

Owner’s checklist

  • the homepage and a conversion-critical page are checked;
  • success requires status, latency and content;
  • at least two independent regions are used;
  • TLS and domain expiry are monitored early;
  • ecommerce has a safe cart transaction check;
  • retries do not hide sustained failure;
  • critical alerts use two channels;
  • primary and backup owners are assigned;
  • the runbook is current;
  • backup recovery has been tested;
  • uptime, latency and incident causes are reviewed monthly;
  • every incident creates a preventive action.

Conclusion

Useful monitoring does not ask whether a server returned something. It asks whether a customer can complete the intended action now. That requires multi-location probes, status and content validation, sensible thresholds, independent alerts, named owners and a rehearsed response plan.

Start with three checks and a controlled alert. Add synthetic transactions, an SLO and monthly incident review as the site grows. If you need a monitoring design, cleaner alerts or an SLA that matches business risk, contact BB STUDIO for a practical setup.

Часті питання

One to five minutes is suitable for most business websites. Use a one-minute interval for checkout, APIs or high-spend campaign pages. Confirm failure with a retry and another location so a brief network issue does not trigger a false incident.

No. A server can return 200 OK with an empty template, an error message or a parking page. Also validate a content marker, latency, TLS and the critical user journey.

The target depends on the cost of downtime. 99.9% is a practical internal goal for many sites, while an active ecommerce operation may require more. The percentage must be accompanied by clear measurement rules and response targets.

The event may have been brief, regional or limited to one URL. A probe IP may also be blocked, or latency may exceed the threshold. Compare timestamps, locations, status codes and server or CDN logs before changing the monitor.

Assign primary and backup owners with the necessary access and a short runbook. For critical incidents, define acknowledgement time, escalation channels and the person responsible for business updates.
Rate this article
It helps us write better content
Be the first to rate 5.0 of 5 0 votes
Поділитися статтею:

Схожі статті

Let’s create something amazing together

Become a clientBecome a client
Telegram Viber Call us