Monitoring Benefit
Service Level Availability turns ENow's monitoring results into an availability record. Every availability test ENow runs against your tenant feeds it: when a test fails, ENow begins counting how long that workload stays down, and reports the resulting uptime percentage per workload.
It answers a question the dashboard cannot. A dashboard tells you what is broken now. This tells you how often each workload has been broken, and for how long — which is what you need when someone asks whether a service has been reliable.
There are two surfaces, and they answer different questions
| SLA Status page | SLA Report | |
|---|---|---|
| Question | Is anything wrong right now? | How reliable has this been? |
| Window | Fixed at the last 24 hours | 7, 30, 90, 180 or 365 days, or YTD |
| Columns | WorkLoad, Up Time, Status | Workload, Number of Checks, Uptime, Downtime, Microsoft SLA, PoP Trend |
| Drill-down | None | Expand a workload for every individual failure |
| Use it for | Alerting and the daily check | Reporting, trend analysis, and finding the cause |
Both cover the same five workloads: Exchange Online, Microsoft Entra, Microsoft Teams, OneDrive and SharePoint.
What it measures, and what it does not
This is ENow's own measurement, calculated from the synthetic availability tests ENow performs against your tenant from your own infrastructure. Three consequences worth being clear about:
- It reflects your experience of the service, including your network path, your service account and your configuration — not a global average.
- The Microsoft SLA column is Microsoft's published commitment, shown for comparison. It is not what ENow measured. Your figure sitting below it does not by itself mean Microsoft breached anything — the two are measured differently, and yours includes conditions specific to your tenant and your network.
- It is not evidence for a Microsoft service credit. Claiming a credit requires evidence gathered to Microsoft's own definitions across the categories they publish. This is a sound operational record; it is not a substitute for that process.
Node type is the most useful thing on the report
The SLA figure blends results from every place ENow tests from. The Node Type filter separates them:
| Node Type | Where the test ran |
|---|---|
| EMS | The ENow web server itself |
| ECloud | The ENow-hosted external probe |
| RemoteProbe | Your own probes, with the Site column naming the location |
This matters because a single unhealthy location drags the blended figure down. Before treating a dip as a Microsoft problem, filter by Node Type. If the failures are confined to one probe or one site, the service was fine and that location was not — which is a network conversation, not a Microsoft one.
Reading the report
Number of Checks varies enormously between workloads, because each workload runs a different number of tests from a different number of nodes. A workload with several test types across three node types accumulates checks far faster than one with a single test. Compare a workload against its own history rather than against another workload's check count.
PoP Trend is direction, not health. It compares the selected period against the one before it. A workload at very high availability can show a red trend because it slipped fractionally, while a workload recovering from a bad period shows green while still being the worst on the page. Read it with the uptime figure, never instead of it.
Expanding a workload lists the individual failures — test name, node type, site, server, how long the test took, when it ran in UTC, and the error returned. That detail grid is where a dip stops being a number and becomes a diagnosis. Export to Excel takes the summary and the detail together.
How uptime accrues
A workload counts as down from the moment a test fails until a test succeeds again. Two consequences cause most of the confusion:
- A brief outage rounds away over a long period. Minutes of failure inside a 90-day window barely move the percentage. The figure is not an outage detector — the workload's own check pages are.
- The figure lags a recovery. Restoring a service does not restore the percentage. Accrued downtime stays in the window until it rolls out.
Where the thresholds are set
Thresholds are configured under Settings > Alert Settings > Alert Thresholds > Universal Policy > Microsoft 365 > Configuration.
Each threshold carries a Warning and a Critical value with its unit, plus an Include In Observation switch that governs whether Observation Mode watches it. These are set once under Configuration — not on each workload's own page.
If you enable Observation Mode, remember it only produces recommendations. Nothing is changed automatically unless Tuning Mode is enabled as well — see Observation and Tuning Mode.
When 0% does not mean an outage
A workload showing 0% uptime for an extended period is more often a collection problem than a real one. A data collection failure, a service connectivity problem, or a delay in the underlying results will all present as a sustained zero, because no successful test result has been recorded.
Before treating a sustained 0% as downtime, open that workload's own check pages. If those checks are passing, the workload is healthy and the SLA figure is not being fed — work the collection, not the service.
Common warning or error results, and potential solutions
| Result | Potential solution |
|---|---|
| One workload below target | Expand it on the report and read the failures. The test name and node type usually identify the cause in seconds. |
| Failures all carry the same Site or Server | One location, not the service. Work that probe's connectivity — see Remote Probe - Connectivity Status. |
| Failures span every node type at the same timestamps | Genuinely tenant-wide. Compare against Microsoft's service health for that window. |
| Uptime below the Microsoft SLA column | Not automatically a breach. Establish first whether the failures were tenant-wide or local to one node. |
| Sustained 0% on one workload | Almost always collection rather than service. Confirm that workload's checks are running and reporting. |
| Every workload at 0% | Not five simultaneous outages. Look at the service account, tenant connectivity, or the collection process shared by all five. |
| Users reported an outage but uptime looks high | Expected over a long period. Shorten the window to 7 days, or use the 24-hour status page. |
For Microsoft's own service health reporting and how to consume it, see Microsoft 365 Service Health and Continuity.
Comments
0 comments
Please sign in to leave a comment.