We create digital solutions that work for businesses
A website can stop accepting leads overnight, return a blank page after an update or fail only for visitors in one region. If customers are the first to report the problem, the business has already lost traffic, advertising spend and trust.
Website availability monitoring sends automated requests, validates the expected response and alerts the responsible people when a condition fails. A five-minute homepage check is useful, but it is only a baseline. A dependable setup distinguishes a network failure from a slow response, a server error, incorrect content and a broken business transaction.
For a typical business website, check the homepage and at least one critical landing page every one to five minutes from multiple locations. Success should require the expected HTTP status, acceptable response time, a stable content marker and a valid TLS certificate.
After two or three consecutive failures, notify an owner through at least two independent channels. For an ecommerce site, separately test the catalogue, cart and a safe checkout path with no real charge. Every alert needs an owner, a runbook and a response target.
Uptime is the share of a measurement period during which a resource meets the defined availability conditions:
uptime = successful time / total measured time × 100%.
The definition of “successful” determines the value of the number. A monitor that accepts any HTTP response can mark a maintenance page returning 200 as healthy. A check from one country will miss some regional DNS or CDN failures. A 30-minute interval may miss short but recurring outages.
Always read an uptime percentage together with its method: monitored URLs, probe regions, interval, success criteria and excluded periods. Reliable website hosting reduces infrastructure risk, but an external monitor is still necessary because it observes the website from the visitor’s side.
The domain must resolve to the expected address and the connection must complete within a reasonable time. A failure can happen before the request reaches the web server: an expired domain, removed record, unavailable authoritative nameserver or stale regional configuration can all make the site unreachable.
After a change, verify the relevant A, AAAA, CNAME, MX and TXT DNS records and keep probes active in several regions until propagation is complete.
A basic monitor sends a GET or HEAD request to a specific URL. A normal page usually expects 200; a deliberate permanent redirect can expect 301 or 308. Codes such as 500, 502, 503 and 504 indicate server-side failure, but not every broken page returns a 5xx response.
Track latency as well. A page may be technically online while taking 15 seconds to respond, which still wastes paid clicks and drives visitors away. Use the separate guide to TTFB, Core Web Vitals and website speed testing to diagnose performance rather than availability alone.
Search the response for a stable marker: the company name, a product heading or a technical string rendered only after the required data loads. This catches soft failures where the server returns 200, but the template is empty, the database is unavailable or a registrar parking page replaced the site.
Do not monitor copy that a content editor may change without warning. Add a stable HTML marker or expose a dedicated health endpoint that reveals no secrets.
Monitor the certificate’s remaining validity, hostname coverage and chain. Automatic renewal can fail after a DNS, permission or reverse-proxy change. The SSL certificate and HTTPS troubleshooting guide explains the checks in more detail.
Track domain renewal separately. A valid certificate will not save a site after registration or nameserver settings expire. For a broader inventory, follow the domain, hosting and server checklist.
An ecommerce or web application monitor should do more than open the homepage. A synthetic script can search for a test product, add it to the cart, open checkout and verify the expected API response. A service website can open a lead form and send a marked test submission to a technical destination.
Synthetic checks must not create real purchases, charge cards or pollute the CRM. Use a test product, a dedicated address, a clear service marker and automatic cleanup.
| Target | Success condition | Typical interval | Priority |
|---|---|---|---|
| Homepage | 200, content marker, normal response time |
1–5 min | high |
| Critical landing page | 200, form or CTA present |
1–5 min | high |
| Catalogue or search | valid, non-empty result | 5 min | high for ecommerce |
| Cart / checkout | safe test journey succeeds | 5–15 min | critical for ecommerce |
| API / webhook | valid response to test request | 1–5 min | based on business impact |
| TLS | certificate valid with enough time left | 6–24 h | high |
| DNS | expected answers from several resolvers | 15–60 min | medium |
| Domain | sufficient time before expiry | daily | high |
| Backup job | job finished and recovery is verified | after every run | critical |
There is no need to hit every page every minute. Select a few endpoints that represent the real customer journey and run broader crawls less often.
One-minute monitoring detects incidents faster but produces more requests and is more sensitive to short network interruptions. A five-minute interval is often sufficient for a small site with no major paid campaign. A checkout, public API or high-spend landing page usually justifies one-minute checks.
A single probe cannot distinguish a website outage from a problem at the monitoring location. A practical sequence is:
This reduces noise without hiding a sustained incident.
The worst monitoring system is the one everybody learns to ignore. False positives come from brief network interruptions, blocked probe IPs, unrealistic latency limits, maintenance windows or fragile browser scripts.
For every check, define:
Avoid unconditional allowlisting that lets a probe bypass the protections faced by real customers. Keep the route realistic, but mark synthetic requests clearly in the User-Agent and logs.
Email is suitable for reports but weak as the only critical channel. A checkout outage should also reach a messenger, push or phone path. At the same time, one slow request should not wake the whole team.
| Severity | Example | Initial response | Channels |
|---|---|---|---|
| P1 critical | site or payment unavailable | 5–10 min | messenger + phone |
| P2 high | form, catalogue or API broken | 30 min | messenger + email |
| P3 medium | slow site, TLS expiry approaching | business hours | email / task tracker |
| P4 informational | one brief failure | review in report | event log |
Include the URL, incident start, locations, status, latency, latest deployment and runbook link in the alert. Ongoing website support and maintenance becomes valuable when there is both a signal and a person with access and responsibility to recover the service.
An SLI is the measured indicator, such as the ratio of successful checkout requests. An SLO is the team’s internal target, such as 99.9% monthly availability. An SLA is an agreement with a customer or provider that defines service levels, measurement rules, exclusions and consequences.
An SLA should not be reduced to a percentage. It should define the monitored component, time zone, calculation period, maintenance windows, third-party dependencies, incident start, response time and remedy.
| Availability | Maximum in 30 days | Maximum in 365 days |
|---|---|---|
| 99% | 7 h 12 min | 3 d 15 h 36 min |
| 99.5% | 3 h 36 min | 1 d 19 h 48 min |
| 99.9% | 43 min 12 sec | 8 h 45 min 36 sec |
| 99.95% | 21 min 36 sec | 4 h 22 min 48 sec |
| 99.99% | 4 min 19 sec | 52 min 34 sec |
These are mathematical values before contractual exclusions. If planned maintenance, third-party failures or attacks are excluded, the user experience can be worse than the reported SLA.
Monitoring without a response plan only reports failure faster. Keep a short runbook:
If compromise is suspected, follow the website malware-checking checklist. If recovery needs a rollback, use a tested website backup and restore process, not the first archive you find.
List what creates revenue: a landing page, lead form, phone number, catalogue, cart, payment, account area or API. Estimate the impact of 15 minutes, one hour and one day of downtime.
Add the homepage, one commercial page and one critical endpoint. Validate status, content, TLS and latency from two or three geographic locations.
For a normal page, start with a five-minute interval and three failures. For checkout, start with one minute and two failures. Base performance limits on your stable baseline rather than an arbitrary ideal.
Assign a primary and backup owner. Test email, messenger and phone channels with a controlled alert. Keep credentials in a protected vault, not in the alert message.
Change the expected marker on a test check or use a safe nonexistent URL. Measure detection, acknowledgement and diagnosis time. Restore the correct condition immediately after the drill.
Useful monitoring does not ask whether a server returned something. It asks whether a customer can complete the intended action now. That requires multi-location probes, status and content validation, sensible thresholds, independent alerts, named owners and a rehearsed response plan.
Start with three checks and a controlled alert. Add synthetic transactions, an SLO and monthly incident review as the site grows. If you need a monitoring design, cleaner alerts or an SLA that matches business risk, contact BB STUDIO for a practical setup.
Let’s create something amazing together