A warm DNS cache can hide a broken lookup path until applications need fresh answers. Effective DNS troubleshooting identifies the failing layer before you restart services or replace resolvers.
Start with evidence from one affected client, separating name-resolution failures from broken routes, filtering, and application errors. Follow its query path when an application requests a fresh answer for the affected FQDN.
Key Takeaways
- Compare the client’s normal lookup with a direct query to its configured resolver before changing settings.
- Test recursive resolution, authoritative answers, and network reachability separately; each exposes a different failure domain.
- Preserve logs and cache evidence before flushing, restarting, or changing forwarding. Validate recovery with uncached lookups as well as application requests.
Start DNS troubleshooting by isolating the failing layer
Record the affected FQDN, query type, client location, timestamp, and application symptom. A “DNS server is not responding” message alone doesn’t identify whether the client, recursive resolver, authoritative DNS, or network path failed.
Compare an affected endpoint with a healthy endpoint on the same network, using the configured DNS resolver.
The failure boundary determines where to investigate first.
| Layer | Evidence to collect | First check |
|---|---|---|
| Client | One endpoint fails while peers work | Adapter settings, cache, local resolver policy |
| Recursive resolver | Clients sharing a resolver fail | Service health, forwarders, validation, load |
| Authoritative DNS | One zone fails across independent resolvers | Records, delegation, authoritative availability |
| Network | Queries time out along a particular path | Routing, firewall rules, transport behavior |
Treat these as starting points, because multiple faults can coexist. A VPN policy can select the correct resolver while a firewall blocks access to it; some network errors, however, are unrelated to DNS.
Also separate lookup time from connection time. If DNS returns the expected address promptly, investigate TCP establishment, TLS negotiation, or authentication next.
Don’t use ping as a DNS verdict. ICMP may be blocked even when applications work. Conversely, a successful ping doesn’t prove the application port is reachable or its DNS answer is current.
Check the client before changing DNS settings
Capture local configuration before applying a fix. Browser, operating-system, and application resolvers may follow different paths.
Inspect resolver selection and local policy
On Windows, ipconfig /all shows network adapter addresses, DNS servers, and suffix configuration. Look for unexpected VPN adapters, obsolete resolver assignments, or DHCP settings that differ from healthy peers. Confirm the expected DNS server for a domain controller, and compare DHCP assignments with any static configuration.
Use Resolve-DnsName -Name $fqdn -Server $resolver -DnsOnly to query an approved server directly. Here, $fqdn is the affected FQDN and $resolver is the DNS resolver address under investigation.
Compare that result with Resolve-DnsName -Name $fqdn -DnsOnly. Different results point toward local cache, DNS resolver selection, or policy.
On Linux, inspect /etc/resolv.conf for nameserver entries. If systemd-resolved manages resolution, use resolvectl status to inspect per-link servers and routing domains. The file may be generated or point only to a local stub, so check the active settings in resolv.conf.
Inspect cache and service state
On Windows, ipconfig /displaydns exposes cached entries, while sc.exe query Dnscache checks DNS Client service state. This checks local DNS client services, not recursive DNS service health.
Follow Microsoft’s DNS client troubleshooting guidance when checking adapter and DHCP configuration. Greyed-out service controls alone don’t establish a service failure. Don’t modify service permissions to force a restart.
After preserving evidence, flush DNS with ipconfig /flushdns to clear the Windows resolver cache. It doesn’t clear browser, JVM, or application caches.
Avoid fleet-wide flushing or network resets during triage. They destroy evidence and can increase resolver load. Likewise, reserve DHCP renewal for confirmed configuration problems.
Validate VPN routes and internal names
Split tunneling VPN failures often affect internal names while public websites continue working. Test the affected internal name with the VPN connected, then inspect the selected resolver and route to its address.
On Windows, Get-DnsClientNrptPolicy -Effective exposes effective namespace resolution policy. Use Get-NetRoute or route print to inspect routing. Confirm that internal DNS traffic follows the intended tunnel.
Next, compare the short hostname with its FQDN. If only the short name fails, investigate suffix search configuration rather than changing upstream providers. Use an absolute name with a trailing dot when testing search-suffix behavior.
For a slow SMB mapping, time resolution before investigating file-transfer throughput. An address-based TCP test can isolate connectivity, but it doesn’t verify Kerberos authentication or share permissions.
Migration work can also change private zones, routes, and forwarding targets. Include these dependencies in cloud migration readiness planning rather than treating DNS as a final deployment check.
Test the recursive resolver directly
Run queries against the DNS resolver from both the affected client and the resolver’s network. Use $resolver for the approved nameserver address and $fqdn for the affected fully qualified domain name (FQDN).
Interpret responses before changing providers
Run dig @"$resolver" "$fqdn" A, then repeat the query for each record type the application needs. Check the status, answer contents, TTL, and query timing.
NXDOMAIN means the responding system reports that the name doesn’t exist. Confirm the queried name and DNS view before assuming a missing record.
SERVFAIL can indicate validation failure, unreachable authoritative servers, or upstream processing problems. A timeout provides less information: it can reflect packet loss, filtering, or an unresponsive service.
A NOERROR response without the requested record isn’t equivalent to NXDOMAIN. Check whether the name exists with another record type.
Check forwarding, caching, and validation
Compare results across enterprise resolvers. For public names, and only where egress is permitted, compare with public services such as Cloudflare or Quad9. Never send private names outside the organization or use this comparison as a reason to replace the enterprise resolver.
Review DNS query forwarding behavior, including failures involving upstream providers, alongside resolver logs for resource exhaustion and DNSSEC errors. For BIND 9, check forwarders, recursion permissions, and validation configuration against ISC’s BIND documentation. Configuration-specific behavior can also vary in pfSense.
A cached answer can conceal an upstream failure. Test an approved uncached name under a zone you control, alongside the affected name. Avoid purging the entire production cache.
Verify authoritative DNS and delegation
When independent recursive systems return the same failure for a public zone, investigate DNS resolver results and authoritative DNS.
Query the zone’s NS records with Resolve-DnsName -Name "$zone" -Type NS -Server "$resolver". Then query each authoritative nameserver directly using dig @"$authoritative" "$fqdn" A +norecurse. Populate $zone and $authoritative with the actual zone and server address.
Compare answers, authoritative flags, and SOA serials across servers. Different answers can reveal incomplete replication or a partially deployed zone change.
For public names, dig +trace "$fqdn" A follows the delegation chain. It requires direct DNS access and doesn’t reproduce private split-horizon resolution, so interpret it within those limits.
Check parent delegation, glue where required, and DNSSEC DS records. A stale parent DS record after a signing change can break DNSSEC validation, even when an authoritative server returns records. That response alone doesn’t prove validating resolvers will succeed.
During recovery, compare TTLs with the change timeline. Lowering a TTL now won’t shorten the lifetime of records already cached under an earlier TTL.
Inspect the network path with packet evidence
Compare Resolve-DnsName -Name "$fqdn" -Type A -Server "$resolver" with Resolve-DnsName -Name "$fqdn" -Type A -Server "$resolver" -TcpOnly. Ordinary DNS uses UDP 53 and TCP port 53, so clients must be able to use TCP when needed.
If UDP fails while TCP succeeds, investigate filtering, fragmentation, or path MTU behavior. A transport difference is evidence, not a complete diagnosis.
Use packet capture at the client and resolver when possible. In Wireshark, follow the request and response for the intended FQDN, then check retransmissions, truncation, and TCP fallback.
A client-side outgoing query without a response can’t identify the dropping device. Correlated captures help locate the boundary.
Encrypted forwarding needs separate inspection. Cloudflare’s DNS over TLS documentation describes TLS-protected queries; the protocol normally uses TCP port 853. Inspect handshake errors and resolver logs because captures won’t expose encrypted query contents.
For SSH timeouts, establish whether resolution finishes before the connection stalls. A stalled TCP connection belongs to routing, filtering, or service investigation.
Audit pfSense forwarding without removing safeguards
In pfSense, distinguish Unbound’s recursive operation from DNS Query Forwarding. Check the configured DNS resolver’s mode, because forwarding changes the upstream dependencies and diagnostic path.
Verify the encrypted forwarding configuration
Review Services > DNS Resolver and System > General Setup against the intended design. Confirm upstream servers, outgoing interfaces, certificate-verification hostnames, and firewall access in pfSense.
An upstream IP address alone doesn’t provide the hostname needed for TLS certificate verification. Also check system time, because certificate validation depends on it. During DNS over TLS troubleshooting, verify that the connection handshake can validate the certificate.
Use Netgate’s pfSense DNS over TLS configuration for the supported setup. Preserve the existing pfSense configuration before editing it, and validate changes on a controlled path before broad deployment.
Treat DNSSEC as a mode-specific decision
Netgate states that DNSSEC is not generally compatible with forwarding mode, with or without DNS over TLS. This is a pfSense configuration caveat.
DNS over TLS protects transport. DNSSEC validates signed DNS data. Decide where validation occurs within the supported forwarding design.
Don’t disable DNSSEC across production to clear unexplained SERVFAIL responses. Investigate signing, delegation, time, and upstream behavior first. Likewise, don’t disable certificate verification, replace corporate resolvers, or bypass security filtering during an incident without an approved rollback plan.
Monitor fresh lookups and validate recovery
Monitor DNS from branch networks, VPN clients, cloud workloads, and recovery regions. A probe beside the DNS resolver can miss failures affecting remote clients.
Track lookup latency, timeout rates, and response codes separately for internal and public zones. Test required record types, including AAAA and SRV where applications use them.
Include approved uncached lookups. Repeatedly querying one warm record measures cache availability more than the upstream resolution path.
A resolver can return cached answers while its upstream path is broken. Recovery checks need both fresh lookups and application transactions.
Connect DNS ownership to disaster recovery dependency planning. Healthy recovery servers remain unusable when private zones or forwarding paths are missing.
Before closing the incident, verify authoritative consistency, fresh resolution, and application access from affected locations. Record the failing boundary and evidence that confirmed the fix.
Frequently Asked Questions
How do I tell whether DNS is causing the failure?
Query the affected FQDN through the client’s configured resolver and compare the result with a direct query to that resolver. If DNS returns the expected address promptly, investigate routing, filtering, or the application connection instead.
What does a SERVFAIL response mean?
SERVFAIL means the resolver couldn’t complete the query; possible causes include DNSSEC validation, unreachable authoritative servers, or upstream processing problems. Check resolver logs and compare results across approved resolvers before changing settings.
Should I flush the DNS cache during troubleshooting?
First preserve cache and log evidence, since flushing can remove clues and won’t fix an upstream failure. Flush the local cache only when evidence points to a stale client entry, and remember that browsers and applications may maintain separate caches.
Why should I test DNS over both UDP and TCP?
DNS commonly uses UDP, but clients must also be able to use TCP when needed. If one transport works and the other fails, investigate filtering, fragmentation, or path MTU behavior along that path.
Keep the next lookup working
Reliable DNS troubleshooting follows the query path and separates each failure domain. Direct queries, configuration checks, and paired packet captures provide stronger evidence than broad resets.
Preserve incident evidence before changing caches or security controls. Then verify fresh lookups through the repaired path, so a warm cache doesn’t hide the next application outage.

