AZ-104 Network Troubleshooting: DNS, Routes and NSGs
A virtual machine can reach an Azure storage account by one hostname but not another. A web server responds on its own private IP while the load balancer reports its backend as unhealthy. A new VNet peering appears connected, yet traffic between two application tiers still times out. These are not interchangeable network failures.
The AZ-104 exam expects administrators to configure and troubleshoot Azure virtual networks, routes, NSGs, service and private endpoints, DNS, load balancing and Network Watcher. Rather than memorizing every network product, build a repeatable diagnostic sequence that starts with the requested application flow and isolates one layer at a time.
Before changing an NSG or route table, write down the source and destination. Which VM, service or client initiated the connection? Which hostname or IP address was used? What transport protocol and destination port are required? Does the client see a timeout, connection refused, TLS error, name-resolution error or an application-level authorization response? Those symptoms can point to different components.
A timeout might indicate a dropped packet, an unreachable route or an unresponsive service. A connection refused can mean the host is reachable but the requested port is not accepting connections. An HTTP 403 from Azure Storage may mean the request reached the service but failed a data permission or firewall check. A successful ping does not prove that the application port and protocol are permitted.
Record a baseline from a working client if one exists. Compare its resolved address, source subnet, effective route and identity with the failing client. That contrast is often more informative than repeatedly reopening the same portal rule panel.
Azure private endpoints map a supported PaaS subresource to a private IP address in a virtual network. Applications usually continue using the service’s normal hostname. DNS is what makes that name resolve to the correct private address for clients whose traffic should use the endpoint. A private endpoint can show an approved connection state while one client still resolves the public address.
Check the complete resolution path from the actual client. Identify whether its resolver uses an Azure private DNS zone, a custom DNS server or a hybrid conditional forwarder. Ensure the recommended private zone is linked to the networks that need it, with the intended record and IP. A peered VNet does not automatically provide every private DNS link. On-premises clients may need a deliberate forwarding path through supported Azure DNS infrastructure.
Do not point an application at a raw private-endpoint IP as a permanent workaround for an incorrect service hostname. It may bypass normal TLS hostname behavior and is hard to maintain when endpoints change. Fix name resolution so the application uses the expected service name and a valid private route.
The AZ-104 virtual networking guide explains the components; this runbook focuses on proving which component owns an actual failure.
Once you know the destination IP, inspect the source network interface’s effective routes. Azure applies system routes alongside any configured user-defined routes and applicable gateway or peering paths. A user-defined route might send traffic toward an inspection appliance, a virtual network gateway or another supported next hop. If that appliance is unavailable or the return path is asymmetric, the request can fail even when the NSG allows it.
A common error in hub-and-spoke designs is assuming VNet peering is automatically transitive. Connecting spoke A to a hub and spoke B to the same hub does not by itself create unrestricted spoke-to-spoke routing through the hub. The intended path depends on routing, gateway transit or an appropriate forwarding appliance and the network configuration. Design and verify both outbound and return traffic paths before assuming that peering status proves reachability.
Use Azure Network Watcher next hop, effective routes or connection troubleshooting where supported, together with source-side route inspection. A route table that looks correct at one subnet may not be associated with the failing VM’s subnet. Always inspect the effective state for the actual source.
Network security groups contain priority-ordered allow and deny rules. A lower number represents higher priority among custom rules in the relevant direction. A subnet-associated NSG and a network-interface NSG may both influence a VM’s traffic; a packet must be permitted through applicable checks. Application security groups help express rules around related interfaces but are not standalone firewalls or substitutes for network topology.
For a VM listening on TCP 443, verify the inbound rules from the actual source address as well as any relevant outbound rules along the return path. The VM’s guest firewall and application process can still refuse the connection even when Azure allows the packet. If a connection fails, inspect the effective rule decision or use a supported Network Watcher diagnostics tool rather than creating an allow-any-any exception.
A particularly misleading situation occurs when the NSG allows the intended service port but a DNS lookup returns an unexpected IP. Changing security-group rules then treats the symptom while leaving the name-resolution error in place. Always trace the destination chosen by the client before evaluating the filtering decision.
IP flow verify or NSG diagnostics can help determine whether the relevant NSG rules allow or deny a specified network flow in supported environments. Next hop helps identify where a packet is expected to be routed for a destination. Effective security rules show the applicable rule set for a network interface. Connection Monitor can observe configured endpoint reachability over time, and topology helps visualize selected network resources and relationships.
These tools are strongest when used in a sequence. Begin with the client’s DNS answer, inspect the effective route and filtering decision, and then test the actual service port. A tool proving a packet is allowed by an NSG does not prove that an application accepts the connection. Likewise, a successful route diagnosis does not establish that a downstream storage service authorizes the requested blob read.
Keep each diagnostic artifact with the incident: source IP, destination IP, port, timestamp, observed next hop and responsible rule. That record makes it much easier to distinguish a real corrective action from a lucky retry.
Use a simple controlled example with two instances of an internal service. Configure the probe to test the expected port or HTTP endpoint and observe which backend remains healthy. Change the test application so it stops answering that probe while continuing to respond on a different local port. The load balancer may remove that instance from the eligible pool even though the guest OS is still running. Next restore the probe and break the actual business endpoint while leaving a shallow health endpoint healthy; that reveals the opposite risk. The result helps you understand why health checks need to reflect genuine service readiness. Probe success is valuable infrastructure evidence, but administrators should still verify a real user transaction before declaring an outage resolved.
An Azure Load Balancer distributes supported traffic among configured backend targets according to rules and health. A backend can be running but considered unhealthy because the configured probe port or path is wrong, the guest firewall blocks it or the application does not answer with an acceptable response. Conversely, a healthy probe may demonstrate only a narrow endpoint’s response, not that every business transaction works.
Take two VMs behind an internal load balancer. If only one backend receives traffic, first inspect backend pool membership, probe configuration and observed health. Then test the application locally on each VM and from an appropriate client path. Review NSG and guest firewall rules that affect the probe and data traffic. Merely increasing VM size or restarting the load balancer is not a disciplined diagnosis.
For application traffic that exits Azure, outbound connectivity and source NAT behavior may create another class of failure. Determine whether the workload has the intended egress design rather than assuming the load balancer’s inbound frontend also grants arbitrary outbound reachability. The Azure Load Balancer article covers its components in detail; the exam-level skill is verifying the complete traffic path.
Azure Bastion provides administrative connectivity to supported Azure virtual machines without requiring a public RDP or SSH endpoint on every VM. That is useful for reducing exposed management surfaces, but it does not make the VM’s application reachable to end users. An administrator connecting successfully through Bastion cannot conclude that the application’s load-balancer or private-service path is healthy.
When a server cannot be administered, distinguish the Bastion deployment and its permissions from the target VM’s guest account, network connectivity and supported access configuration. When an application is unreachable, test its service port independently of Bastion. A service can be fully operational while its management access is misconfigured, or the reverse.
Use least-privilege resource roles and an approved administrative identity for the lab. Avoid opening broad public RDP or SSH rules as a permanent response to a temporary troubleshooting inconvenience.
Create a disposable two-subnet environment with one test VM and one reachable service. Record that the client resolves the correct name, has a valid route and is allowed by the effective NSG. Then introduce a private DNS record or zone-link error and capture the new resolution result. Correct it, and separately add a scoped NSG rule that denies the service’s required port. Finally, restore the rule and introduce a test user-defined route to an unavailable next hop in a controlled network.
For each failure, do not announce the cause before testing. Write the client symptom, the first tool you would use, the evidence that should appear and the minimum safe correction. Compare the resulting observations. A network specialist who can show why a specific route or rule was wrong has more transferable skill than someone who repeatedly toggles settings until traffic happens to pass.
Clean up the test configuration through a documented rollback sequence. Confirm the source can again perform the intended application action, not merely resolve a hostname.
A useful network incident report contains the failing flow, expected destination, DNS answer, effective route, filtering decision, application endpoint result and any relevant change preceding the incident. That structure also helps identify whether the issue belongs to the network team, application team, identity administrators or service owner. One team should not fix another team’s configuration by permanently weakening its own controls.
For AZ-104, approach every networking scenario as an ordered investigation: identify the flow, resolve the name, examine the next hop, evaluate filtering, verify the listener and test the user transaction. The tools are important because they provide evidence for those steps. The sequence is what prevents technically plausible but irrelevant configuration changes.