Azure Monitor for AZ-104: Metrics, Logs and Alerts
An Azure virtual machine is running but users are getting timeouts. CPU utilization is normal, and the dashboard is green. That does not prove the application is healthy; it means the team has not yet identified the relevant signal. A storage permission change, a DNS problem or a failed service dependency may never appear as high CPU.
The AZ-104 exam covers Azure Monitor metrics, logs, diagnostic settings, Log Analytics queries, alert rules, action groups, alert processing rules, VM/storage/network Insights and Network Watcher. The practical question is which evidence will explain the failure, not how many dashboards the administrator can create.
Platform metrics summarize numerical behavior such as CPU percentage, transaction rate and response latency over time. Resource logs describe service-specific events, often useful for explaining a denied request or malformed operation. The Azure Activity Log records subscription-level management operations, including deployments and policy-related changes. Those streams have different purposes, formats and availability.
A running VM with an unhealthy guest process may need information from inside the operating system; an application with a failing API dependency may require request tracing. Ask whether you need to know what changed, what failed, whether a condition is increasing, or which network hop broke. Each question should lead to an appropriate source rather than a reflex to add more metrics.
A useful monitoring design begins with one real service objective and a few indicators that distinguish its likely failure modes. More collection is not necessarily better if nobody can interpret or respond to it.
For a concrete exercise, choose an Azure storage account that a test application reads through an identity-based connection. Enable only the diagnostic category needed to investigate access operations, send it to a known workspace and run one permitted request followed by one denied request. Compare what appears in the workspace, whether it contains useful correlation information, and how long ingestion takes. A request blocked before it reaches the storage service may leave different evidence from one authenticated request denied by the service itself. That contrast prevents investigators from treating every missing event as proof that nothing occurred.
Document the data pipeline as an operational contract: which resource emits the event, which category was enabled, which workspace or destination receives it, who may inspect it, and what retention period applies. A later change to collection can then be reviewed against the questions the operations team actually needs to answer. Telemetry should be selected for diagnostic value, not accumulated merely because the portal offers an additional checkbox.
Azure resource types expose different diagnostic categories and destinations. An administrator can enable supported categories and route the information to Log Analytics, storage or another supported destination. The specific log categories and tables must be checked on the resource; copying another team’s settings blindly can omit the only information that would explain an incident.
Log Analytics workspaces provide query capabilities and shared investigation context, but also create access-control, ingestion and retention costs. Before sending a high-volume event stream to a workspace, determine who needs it, how often it is queried and how long the evidence must be kept. Development telemetry and regulated production audit data may have very different requirements.
A query returning no results does not prove there were no incidents. Verify the diagnostic setting is enabled, the right category and workspace were selected, the query uses the correct time range and the data has had time to arrive. Missing data can be a collection failure.
After a corrective deployment, repeat the same query over the new time window and compare the expected operation statuses. A decrease in failures is useful, but confirm the application’s actual outcome as well. An API may stop generating authorization errors while still returning stale data because of a separate caching or dependency problem. When the KQL result is noisy, group by resource identifier or narrow it to a single deployment correlation rather than selecting a wider time range hoping the cause becomes visible. For a recurring incident, save a documented query with its assumptions and expected data table so another operator can reproduce the analysis. Explain how missing logs would affect the interpretation; absence of a recorded failure is not proof that every request succeeded.
Kusto Query Language (KQL) lets administrators filter, group, project and summarize records in Log Analytics. An AZ-104 candidate should be able to inspect a bounded time range and narrow results to a particular resource or operation. Table availability depends on the actual collected data; example query tables are not present by default in every workspace.
If subscription Activity Log events have been routed to a workspace, this illustrative query groups recent management operations by name and status:
AzureActivity
| where TimeGenerated > ago(24h)
| summarize Operations = count()
by OperationNameValue, ActivityStatusValue
| order by Operations desc
A row showing several failed operations is a starting point, not a root cause. Filter by the affected resource and review an individual deployment operation. Was it refused by a missing role, a deny policy, a lock, quota or invalid service configuration? Compare the result with the reported incident time and the team’s change history. A good query replaces a guess with testable evidence.
Azure Monitor Insights offers curated views for supported resources such as virtual machines, storage accounts and networks. A VM Insights dashboard can accelerate investigation, but its detail depends on the necessary agent and data collection configuration. An Azure platform metric may show CPU pressure while guest telemetry is needed to explain a full disk or stopped service.
Suppose a web server slows after a deployment while CPU remains normal. The useful next observations might be disk pressure, failing downstream connections or a guest process repeatedly restarting. A generic availability status will not settle those possibilities. Use Insights to identify a lead and confirm it against the underlying event and a reproducible user transaction.
If an expected chart is empty, check its collection path before assuming that Azure Monitor itself is broken. The most dangerous observability gap is one that silently looks like healthy zero activity.
A metric alert evaluates a numerical condition using supported metrics and thresholds. A log search alert evaluates a query over a configured window and schedule. The choice affects freshness, cost and the detail available. An uncomplicated sustained CPU warning may be best expressed as a metric rule; a rule that must distinguish a specific API error or event field may require logs.
Define an alert around something a responder can act on. A few seconds of elevated latency during a scheduled batch may be normal; sustained customer-facing failure is not. Ask how long the symptom must persist, which resources are in scope, and what evidence an on-call administrator should inspect. A rule firing constantly during healthy operation teaches people to disregard alerts.
The Azure Monitor alerts guide covers the configuration process in depth. Here the important skill is connecting the chosen condition to a response and recovery test.
An action group defines notifications or automation for an alert, such as email, webhook or a supported response workflow. The alert rule determines when the condition fires. That distinction matters: a technically correct production warning delivered only to an unattended mailbox is not an effective operational control.
Alert processing rules can add or suppress action groups for matching alerts, often during a maintenance window. They do not necessarily prevent the underlying alert condition from existing. Document suppression windows and verify that important unrelated alerts still reach the appropriate on-call person. Never use a broad suppression just because troubleshooting an incident is inconvenient.
Test one alert end to end in a controlled environment. Cause a known condition, verify the alert instance, confirm the notification reaches a responsible person and check that it can be acknowledged and resolved. A saved action-group setting is not proof of successful delivery.
A service may be healthy but unreachable because a route points to the wrong next hop, an NSG denies traffic or DNS resolves a private endpoint to the wrong address. Azure Network Watcher offers relevant tools such as effective security rules, next hop, IP flow verify, NSG diagnostics, topology and Connection Monitor for supported scenarios. A CPU chart cannot replace those network observations.
If an application fails to connect to Azure Storage over a private endpoint, first check the address its client resolves, then the intended route and effective security controls. A valid Microsoft Entra data role will not fix an incorrect network path; an allowed network path does not grant data permission. Keeping those checks separate avoids repeatedly changing unrelated settings.
Connection Monitor can help observe ongoing connectivity between selected endpoints and recognize a recurring problem. After a fix, test the application’s actual transaction as well as the network probe. Network reachability and functional recovery are different outcomes.
Deploy a small disposable service with a storage dependency, enable one relevant platform metric and one diagnostic data stream, and configure an alert with a named recipient. Introduce a harmless performance problem and note which indicators respond. Restore the service, then deliberately block access to its storage dependency through an isolated network setting. The second fault should require a different explanation than the first.
Record a chronology: user symptom, first useful telemetry, recent resource change, hypothesis, verification step and smallest safe correction. Evaluate which signals were missing or misleading. Running through that sequence builds diagnostic judgment in a way that repeatedly opening the same dashboard cannot.
After remediation, repeat the original user action and confirm the indicators have returned to an expected range. Clearing an alert without restoring the transaction is not success.
Telemetry creates ongoing cost and maintenance work. Decide who can query each workspace, whether the chosen retention period is justified, and which alert owners will still be present in six months. When a resource moves subscriptions or is retired, review its diagnostic settings and alert scopes so that missing resources do not generate unexplained noise.
Monitoring should be considered during deployment planning, not after the first incident. A new workload should have a known set of critical symptoms, suitable signals and a response path before release. For AZ-104, understanding how evidence leads to action is more valuable than memorizing the exact arrangement of portal menus.