> For the complete documentation index, see [llms.txt](https://docs.powermonitor.com.br/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.powermonitor.com.br/en/technical-documentation/documentacao-tecnica/monitoramento-continuo.md).

# Continuous Monitoring

How Power Monitor continuously checks the health of the environment, opens and resolves incidents, and sends alerts.

Continuous monitoring is separate from [inventory scans](/en/technical-documentation/documentacao-tecnica/fluxo-de-scans.md). While scans collect **what exists** in the environment, monitoring checks the **health** of resources and detects **failures** in short cycles.

## Check cycle

```
Health check job (every 30 seconds)
    │  selects the due monitoring tasks of each organization
    ▼
Monitoring task (by type, with its own cadence)
    │  queries the Microsoft APIs with the organization's Service Principal
    ▼
Resource health state ──▶ incident opened / kept / resolved
    │
    ▼
Notification policy ──▶ email, Teams, Slack, Telegram
```

| Task                       | What it checks                                                          | API                                                               | Cadence    |
| -------------------------- | ----------------------------------------------------------------------- | ----------------------------------------------------------------- | ---------- |
| **Gateway status**         | Each gateway on which the Service Principal is an administrator         | `GET /v1.0/myorg/gateways/{id}`                                   | 5 minutes  |
| **Semantic model refresh** | Latest refreshes of models on capacity in monitored workspaces          | `GET /v1.0/myorg/admin/capacities/refreshables`                   | 15 minutes |
| **Fabric item runs**       | Pipelines, notebooks, Copy jobs, and Dataflows Gen2                     | Fabric job instance APIs                                          | 2 hours    |
| **Fabric item schedules**  | Collection of schedules (no alert)                                      | Fabric scheduling APIs                                            | 6 hours    |
| **Capacity consumption**   | Interactive, background, and accumulated usage, by capacity and by item | DAX query (`executeQueries`) to the Fabric Capacity Metrics model | 5 minutes  |

Other scheduled jobs complement monitoring:

| Job                                        | Cadence                                 |
| ------------------------------------------ | --------------------------------------- |
| Detection of consumption anomalies by item | 5 minutes (baseline recalculated daily) |
| Execution time deviation                   | Every hour                              |
| Data Freshness and Fabric Mirroring        | 10 minutes                              |
| Capacity time-window alerts                | 5 minutes                               |
| Pause, resume, and SKU change schedules    | Every minute                            |
| Capacity auto-scale                        | 5 minutes                               |
| Capacity cost alert                        | Daily                                   |
| Hourly checklist / Daily checklist         | Every hour / daily                      |
| Reconciliation of monitoring tasks         | 30 minutes                              |

Monitoring tasks are created and maintained automatically when the administrator turns on each monitoring in **Settings › Monitoring**.

## Health state and incidents

Each monitored resource (gateway, semantic model, Fabric item, capacity) has a health state:

| State         | Meaning                                                                                     |
| ------------- | ------------------------------------------------------------------------------------------- |
| **Unknown**   | Not yet evaluated or undetermined; does not generate a notification                         |
| **Healthy**   | Working normally                                                                            |
| **Degraded**  | Situation requiring attention (for example, high background consumption or run in progress) |
| **Unhealthy** | Failure: gateway offline, failed refresh, failed or canceled run, capacity above the limit  |

| Transition        | Effect                                                                |
| ----------------- | --------------------------------------------------------------------- |
| Healthy → failure | **Opens an incident** and sends the opening notification              |
| Failure → failure | The incident remains open; recurrences are counted                    |
| Failure → healthy | **Closes the incident** and sends the **Alert Resolved** notification |

## Notification policy

* **Gateways, semantic models, and Fabric items:** email **when the incident is opened and when it is resolved**. Recurrences go into the failure counter and appear in the **Hourly Checklist**, avoiding excessive messages.
* **Capacity usage limit (80%):** repeated notices while the situation persists.
* **Background usage:** at most one notice per day in each threshold (50–69%, 70–79%, 80–98%, and 100% or more).
* **Consumption data unavailable:** one alert per episode, closed automatically when the data returns.
* **Recipients:** defined per alert type in **Settings › Notifications**, respecting each user's workspace scope. Unmonitored workspaces do not generate refresh or Fabric item alerts.

## Logging

* Each alert is recorded and available in **Monitoring › Alerts**.
* Each email sent is recorded in **Audit › Email Audit**.
* Emails use templates rendered on the server and are sent through SendGrid or the SMTP account configured by the organization. Teams, Slack, and Telegram use the Power Tuning bot service.

## Related pages

* [Real-time alerts](/en/principais-funcionalidades/alertas-em-tempo-real.md)
* [Scan Flow](/en/technical-documentation/documentacao-tecnica/fluxo-de-scans.md)
* [Data Model](/en/technical-documentation/documentacao-tecnica/modelo-de-dados.md)


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation by asking a question.

Perform an HTTP GET request on the following URL with the `ask` and `goal` query parameters:

```
GET https://docs.powermonitor.com.br/en/technical-documentation/documentacao-tecnica/monitoramento-continuo.md?ask=<question>&goal=<user_goal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is what the user is ultimately trying to achieve, the reason they need the answer. Sharing it helps GitBook give you a better, more relevant answer. A goal is most helpful when it describes the outcome the user wants rather than restating the question. For example, with `ask=how do I create an API token`, a goal like `automate deployments from our CI pipeline` lets GitBook tailor the answer to that use case.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
