← DNS: understand and diagnose resolution
08 / 8 · 45 MIN

Cache, transport and evidence-based rollback

Calculate known-entry windows and confirm how the client handles truncation while limiting lab conclusions.

Build a timeline per consumer

Use one row per observed entry, recording value, receipt time and received TTL. R1 gets an A at 10:00 with TTL 600; R2 gets the same A at 10:04. At 10:05, 300 and 540 seconds remain respectively. The difference does not imply an error. This model excludes early refresh, serve-stale and local policies. Write those assumptions alongside the calculation. A useful deadline for one specific entry does not represent every existing cache and application.

Prepare the change before the window

If the old TTL was 600 and all authorities change to 60 at 11:00, an entry received just before the change can still last until roughly 11:10 in the simple model. Lowering TTL at cutover is too late for that population. For new names, include possible negative answers cached before creation. Define observations at relevant consumers, overlap capacity and a communication plan. Do not use the portal write time as evidence that every client has updated.

Rollback also leaves a remaining population

R2 receives the new destination at 14:05 with TTL 60. The team rolls authority back at 14:05:20 after a failure. Forty seconds remain in that entry without early refresh. Plan mitigation for that consumer and base incident status on observed recovery. An application can also retain an already established TCP connection: a new DNS response does not change that session’s destination. Investigation then moves to the reconnection policy and the application lifecycle rather than repeatedly editing a correct record.

Observe truncation and actual retry

The runner uses only 127.0.0.1 on an ephemeral port and sends fictional responses to the installed dig 9.10.6. For large.fund.test, the UDP response has TC; normal retry uses TCP and receives A. With +ignore only the truncated response remains. With +tcp the query starts directly over TCP. Fixture events show the transports actually used. This distinguishes incomplete output from working transport and confirms that an A query over TCP is not thereby an AXFR transfer.

Experiment limits and production criteria

The eight observations confirm client interpretation and transport choices against controlled messages. There is no recursive resolver, real cache expiration, external delegation traversal or DNSSEC validation. Timeline calculations are exercises, not cache measurements. The consulted BIND documentation identifies 9.20.29; that was not the executed version. In production, closure requires authorized tests from relevant origins, including responses needing TCP. Serve-stale can alter post-expiry expectations when refresh fails; confirm the policy before attributing the result to corruption.

# Execute the loopback fixture from the project root:
python3 content/labs/dns-evidence/run.py
# evidence.json records fallback: [udp, tcp]
# ignoredTruncation: [udp]; forcedTCP: [tcp]
# TTL examples are calculations, not measured resolver behavior.
IN PRACTICE

Fictional case: a rule allows short UDP queries but blocks TCP after TC. Escalation supplies the packet sequence and effective rule; closure includes a complete answer from the affected origin.

Common pitfalls

Do not declare global propagation from one TTL, confuse +ignore with a complete answer, or use loopback success as proof of production connectivity.

Related topics: Resolver, authority, and client context · TTL, negative caching, and controlled change · DNS transport and DNSSEC validation

Take this idea with you

Calculate per entry, observe per origin and report only behavior actually demonstrated by commands, events and context.

Create account

Reference: RFC 1035 · DNS RFC 1034/1035 with RFC 2181, 2308, 3596, 4033, 7766 and 8767; dig BIND 9.20