When a program asks for api.northstar.example, the C library consults /etc/nsswitch.conf to decide the order, usually files then dns. Files means /etc/hosts. DNS means the nameserver lines in /etc/resolv.conf, queried in order. The search line appends domains to short names. Containers usually get their own copy of these files, baked into the image or injected by the runtime, which is how a container can disagree with its host.
The response codes carry meaning:
- NOERROR with an A record. The name resolved.
- NXDOMAIN. An authoritative "this name does not exist". The stub resolver stops rather than asking the next server. A public resolver that cannot see your internal zone answers this way for internal names (split-horizon DNS).
- SERVFAIL. The server could not produce an answer, often because its own upstream is broken. Resolvers typically try the next configured server.
- Timeout. Nobody answered. Every lookup pays the full timeout before failing over, which turns into mysterious latency.
Because "the IP works" and "the name works" are independent stages, test them independently:
- Query each resolver directly with
dig @SERVER NAMEand compare answers. - Ask the way the application asks (
getent hosts NAME), which honourshostsand search domains. - Then test HTTP by name.
Pinning a name in /etc/hosts or hardcoding an IP makes today's symptom disappear and guarantees a surprise when the service moves. Fix the resolver configuration where it is generated, whether that is the image, the runtime, or DHCP, so a rebuild does not bring the fault back.