Skip to content

DNS resolver

Source files: 25 · checked against Etemenanki 596916d · katana v3.0.1
  • Etemenanki/protocols/src/dns/mod.rs
  • Etemenanki/protocols/src/dns/message.rs
  • Etemenanki/protocols/src/error.rs
  • Etemenanki/protocols/src/helpers/parse.rs
  • Etemenanki/protocols/src/helpers/address_family.rs
  • Etemenanki/protocols/src/transports/tls/config.rs
  • Etemenanki/protocols/src/transports/connect.rs
  • Etemenanki/protocols/src/hysteria/connector.rs
  • Etemenanki/protocols/src/wireguard/connector.rs
  • Etemenanki/protocols/src/wireguard/device.rs
  • Etemenanki/app/src/config.rs
  • Etemenanki/app/src/instance.rs
  • Etemenanki/app/src/main.rs
  • Etemenanki/app/src/outbound/mod.rs
  • Etemenanki/app/src/outbound/freedom.rs
  • Etemenanki/app/src/balancer.rs
  • Etemenanki/protocols/tests/unit/dns/mod.rs
  • Etemenanki/protocols/tests/unit/dns/message.rs
  • Etemenanki/protocols/tests/pipeline/dns_secure.rs
  • Etemenanki/app/tests/integration/e2e_dns.rs
  • Etemenanki/app/tests/support/mod.rs
  • Etemenanki/protocols/tests/support/mod.rs
  • katana/src/config.rs
  • katana/src/runtime.rs
  • katana/src/outbound/mod.rs

etemenanki_protocols::dns turns a domain name into the list of addresses an outbound dials. It puts a bounded, TTL-aware cache in front of one of four backends: the host resolver (getaddrinfo), plain DNS over UDP, DNS over TLS (RFC 7858) or DNS over HTTPS (RFC 8484). For the three protocol backends it carries its own small wire codec, which encodes one A or AAAA question and reads the address records out of the response.

This page is for contributors who change the resolver, its codec, or the way the app and katana build and share it. The operator view of the same feature, the [dns] table, is in the user guide under DNS.

The module does four things:

  • Build a resolver from configuration strings. Resolver::from_spec is the one place where "udp", "tls" and "https" get their meaning, and where a missing field is refused. The app and katana both call it, so the two programs cannot disagree about what a [dns] table means.
  • Resolve a name to addresses. Resolver::resolve returns every address, de-duplicated, in backend order. It returns address literals unchanged and serves repeat lookups from the cache.
  • Cache answers within bounds. moka bounds the cache to CACHE_CAPACITY names. It enforces the bound during its housekeeping, so the entry count can briefly run past it. Each entry expires after the answer’s TTL, clamped to [MIN_TTL, MAX_TTL]; a system answer, which has no TTL, is held for SYSTEM_TTL.
  • Speak just enough DNS. message.rs encodes a recursive query and decodes A and AAAA records from untrusted bytes without panicking, looping or reading out of bounds.

Address-family policy is outside this module. The resolver returns everything it learned. protocols/src/helpers/address_family.rs → resolve_candidates then applies the outbound’s AddressFamilyStrategy and the caller’s socket capabilities. The resolver also makes no routing decisions and never resolves names for the router. It serves outbounds and balancer health probes only.

protocols/src/dns/mod.rs
pub const SYSTEM_TTL: Duration = Duration::from_secs(60);
pub const MIN_TTL: Duration = Duration::from_secs(5);
pub const MAX_TTL: Duration = Duration::from_secs(3600);
pub const QUERY_TIMEOUT: Duration = Duration::from_secs(5);
const CACHE_CAPACITY: u64 = 8192;
const UDP_BUF: usize = 512;

The Limits section below lists them together with the codec constants.

#[derive(Debug, Default, Clone, Copy)]
pub struct ResolverSpec<'a> {
pub backend: Option<&'a str>,
pub server: Option<&'a str>,
pub server_name: Option<&'a str>,
pub url: Option<&'a str>,
pub ca_pem: Option<&'a [u8]>,
}
impl Resolver {
pub fn from_spec(spec: ResolverSpec<'_>) -> io::Result<Self>;
}

ResolverSpec holds the strings a configuration file carries. It is a struct rather than a parameter list, so adding a backend field does not break every consumer’s call site. ca_pem is PEM bytes, not a path: the caller reads the file. Both app/src/instance.rs → build and katana’s src/runtime.rs → build_outbounds call std::fs::read on ca_file before calling from_spec. They read the file whatever the backend, so a missing ca_file fails the build with the bare OS error, for example No such file or directory (os error 2).

from_spec matches backend exactly. The match is case-sensitive, and None means "system".

backend Requires Builds Ignored fields
None or "system" nothing Resolver::system() server, server_name, url, ca_pem
"udp" server Backend::Udp(addr) server_name, url, ca_pem
"tls" server, server_name Backend::Tls { .. } url
"https" server, url Backend::https(server, url, ca_pem) server_name
anything else error

server is parsed as a SocketAddr ("192.0.2.53:53", "[2001:db8::53]:853"), so it needs an explicit port and is never a host name. The resolver would otherwise have to resolve its own server, and it has nothing to resolve it with. Every refusal is io::ErrorKind::InvalidInput. None of them falls back to the host resolver: an operator who asked for a particular resolver must not get a different one without being told. The error texts are listed in the Failure paths section below.

#[derive(Debug, Clone)]
pub enum Backend {
System,
Udp(SocketAddr),
Tls {
server: SocketAddr,
server_name: CompactString,
ca_pem: Option<Vec<u8>>,
},
Https {
server: SocketAddr,
host: CompactString,
path: CompactString,
ca_pem: Option<Vec<u8>>,
},
}
impl Backend {
pub fn https(server: SocketAddr, url: &str, ca_pem: Option<Vec<u8>>) -> io::Result<Self>;
}

Backend::https splits the URL itself instead of pulling in a URL parser:

  1. The URL must start with https://, or it is refused with dns: "<url>" is not an https:// url.
  2. The authority runs up to the first /. The rest is the request path. With no /, the path is /dns-query.
  3. An empty authority, or one containing @, is refused with dns: "<url>" has no usable host. Userinfo would otherwise end up inside the verified name.
  4. Anything after the first : in the authority is dropped. server is the address that gets dialed, so a URL port has no meaning here and must not leak into the certificate name.

host is used twice: as the TLS server name to verify and as the HTTP Host header. Because of the plain split at :, a bracketed IPv6 literal cannot serve as the URL authority: https://[2001:db8::53]/dns-query builds, but host becomes [2001, which no certificate matches. Use a DNS name in the URL and put the address in server.

#[derive(Clone)]
pub struct Resolver {
inner: Arc<Inner>,
}
struct Inner {
backend: Backend,
min_ttl: Duration,
max_ttl: Duration,
tls: Option<ClientConfig>,
cache: moka::future::Cache<CompactString, Cached>,
}
#[derive(Clone)]
struct Cached {
addresses: Arc<Vec<IpAddr>>,
expires_at: Instant,
}
impl Resolver {
pub fn system() -> Self;
pub fn new(backend: Backend) -> io::Result<Self>;
pub fn from_spec(spec: ResolverSpec<'_>) -> io::Result<Self>;
pub fn with_ttl_bounds(self, min: Duration, max: Duration) -> Self;
pub async fn resolve(&self, host: &str) -> io::Result<Vec<IpAddr>>;
}

A Resolver is one Arc. Cloning it shares the backend, the TLS client configuration and the cache, and every consumer relies on that. Default is Resolver::system(). The Debug impl prints only the backend.

  • Resolver::new builds the ClientConfig once, for Tls and Https only, and stores it in Inner.tls. It fails when ClientConfig::with_verify_mode fails, which the TLS section below describes.
  • Resolver::build (private) creates the cache with max_capacity(CACHE_CAPACITY) and time_to_live(MAX_TTL), and sets min_ttl/max_ttl to MIN_TTL/MAX_TTL.
  • with_ttl_bounds builds a new Inner with the same backend and TLS configuration and a clone of the same moka handle. Anything already cached carries over. Only the clamp changes: moka’s time_to_live backstop stays at the MAX_TTL the cache was built with. At the pinned revisions neither the app nor katana calls it. The tests use it to force expiry.

The lookup machinery is private:

fn build(backend: Backend, tls: Option<ClientConfig>) -> Self;
async fn lookup(&self, host: &str) -> io::Result<(Vec<IpAddr>, Option<Duration>)>;
async fn query_both(&self, host: &str, server: SocketAddr) -> io::Result<(Vec<IpAddr>, Option<Duration>)>;
async fn query(&self, host: &str, server: SocketAddr, qtype: u16) -> io::Result<Answer>;
async fn exchange_https(&self, server: SocketAddr, host: &str, path: &str, query: &[u8]) -> io::Result<Vec<u8>>;
async fn exchange_tls(&self, server: SocketAddr, query: &[u8]) -> io::Result<Vec<u8>>;
// free functions in the same module
async fn exchange_https_over(
mut stream: impl tokio::io::AsyncRead + tokio::io::AsyncWrite + Unpin,
host: &str,
path: &str,
query: &[u8],
) -> io::Result<Vec<u8>>;
async fn exchange_udp(server: SocketAddr, query: &[u8]) -> io::Result<Vec<u8>>;

A cache only helps if the callers that resolve overlapping names share it. Both programs therefore build exactly one resolver per outbound set and hand clones of it to everything that resolves.

flowchart LR
  cfg["[dns] table"] --> spec["ResolverSpec"]
  spec --> fs["Resolver::from_spec"]
  fs --> r["Resolver (one Arc)"]
  r --> free["FreedomConnector"]
  r --> tc["TransportConnector"]
  r --> hy["Hy2Connector::with_resolver"]
  r --> wg["WgConnector::with_resolver"]
  r --> bal["Balancer::spawn_probe"]
Program Where the resolver is built Lifetime Who receives a clone
etemenanki-app app/src/instance.rs → build, before any outbound One per generation, stored in Built.resolver Every outbound that dials, through build_outbound(ob, &resolver) (blackhole takes none): freedom (TCP dial and UDP ResolvingUdp), the TransportConnector of every stream-based proxy outbound, Hysteria 2 (the server name) and WireGuard (tunnelled destinations). Also every balancer’s health probe.
katana src/runtime.rs → build_outbounds, before the reserved direct/block handlers One per outbound pool The built-in direct/freedom handlers and every configured [[outbound]]. All nodes share the pool, so they share the resolver.

Consequences for contributors:

  • A reload starts with an empty cache. The app builds a new generation, and with it a new resolver, on every successful reload. The old generation keeps its resolver until it is dropped. When build fails, the old generation and its warm cache stay in place. The reload path is described in Generations and reload.
  • katana rebuilds the resolver only with the pool. katana’s apply_reload calls build_outbounds only when new.outbounds != cfg.outbounds. An edit to [dns] alone is therefore neither applied nor validated on reload. It takes effect the next time the [[outbound]] list changes, or at restart, and an invalid [dns] table surfaces only then, as reload: bad outbounds, keeping current config: <message>.
  • A default-constructed consumer has a private cache. FreedomConnector::default(), TransportConnector::tcp(), and WgConnector and Hy2Connector until with_resolver is called, start with Resolver::default(), which is a new system resolver with its own cache. Anything built outside build/build_outbounds must receive the shared resolver explicitly, through FreedomConnector::new, TransportConnector::new or with_resolver, or it resolves on its own.
  • A WireGuard peer endpoint is not resolved here. WgConnector uses the shared resolver for destinations reached through the tunnel. The peer’s own endpoint, when it is a name, is looked up by protocols/src/wireguard/device.rs → resolve_endpoint with tokio::net::lookup_host, which is the host resolver without this cache, and only its first address is used.
  • The resolver dials directly. DNS traffic goes out through tokio::net::UdpSocket::bind and TcpStream::connect, not through the etemenanki-environment dialers, so the dialers’ SocketOptions and connect timeout do not apply to it. QUERY_TIMEOUT is its only deadline. See Dialers.
flowchart TB
  start["resolve(host)"] --> lit{"host parses as IpAddr?"}
  lit -- yes --> retlit["Ok(vec![ip]), no cache, no backend"]
  lit -- no --> key["key = host.to_ascii_lowercase()"]
  key --> hit{"cached and expires_at > now?"}
  hit -- yes --> rethit["Ok(cached addresses)"]
  hit -- no --> be{"backend"}
  be -- System --> gai["lookup_host((host, 0)), ttl None"]
  be -- "Udp, Tls, Https" --> both["query_both: A and AAAA"]
  gai --> empty{"no addresses?"}
  both --> empty
  empty -- yes --> nf["Err NotFound: did not resolve"]
  empty -- no --> ttl["ttl.unwrap_or(SYSTEM_TTL), clamp(min_ttl, max_ttl)"]
  ttl --> ins["cache.insert(key, Cached)"]
  ins --> ok["Ok(addresses)"]

Details the diagram leaves out:

  • The key is lowercased, the query is not. key is host.to_ascii_lowercase(), but lookup(host) sends the name as the caller spelled it. Example.COM and example.com share one entry.
  • Expiry is per entry. moka’s own time_to_live is a single value for the whole cache and serves only as a backstop at MAX_TTL. The real deadline is Cached.expires_at, which resolve compares with tokio::time::Instant::now(). An expired entry is not removed on read. The next successful lookup overwrites it, or moka evicts it.
  • The deadline cannot overflow. Instant::now().checked_add(ttl) falls back to Instant::now(). The entry is then already stale, which fails safe.
  • The System backend dedupes and drops ports. It calls tokio::net::lookup_host((host, 0)). That runs getaddrinfo on tokio’s blocking pool. The backend drops the port, keeps the first occurrence of every IP, and reports no TTL, so SYSTEM_TTL applies. It is the only backend that honours /etc/hosts, nsswitch.conf and search domains, which is why it stays the default.
  • Failures are not cached. An error, including NotFound for an empty answer, is returned without touching the cache, and the next call asks the backend again.
  • Concurrent misses are not merged. resolve uses cache.get followed by cache.insert, not a get-or-insert-with. Two callers that miss on the same name at the same moment each run a lookup. The later insert wins.
  • A hit returns a copy. The cached Arc<Vec<IpAddr>> is cloned into a fresh Vec for the caller.

Every protocol backend asks two independent questions with tokio::join!, each under its own QUERY_TIMEOUT, each with a fresh random 16-bit ID (rand::random):

flowchart TB
  qb["query_both(host, server)"] --> join["tokio::join!"]
  join --> ea["encode_query(id_a, host, TYPE_A)"]
  join --> e6["encode_query(id_b, host, TYPE_AAAA)"]
  ea --> xa["timeout(QUERY_TIMEOUT, exchange), own socket or connection"]
  e6 --> x6["timeout(QUERY_TIMEOUT, exchange), own socket or connection"]
  xa --> da["decode_answer(id_a, response)"]
  x6 --> d6["decode_answer(id_b, response)"]
  da -- "Ok(Answer) or Err" --> merge["walk v4 then v6: dedupe addresses, min TTL, keep last_err"]
  d6 -- "Ok(Answer) or Err" --> merge
  merge --> none{"no address and an error?"}
  none -- yes --> err["Err(last_err)"]
  none -- no --> ok["Ok((addresses, ttl))"]

query_both walks the results in the order [v4, v6]:

  • A successful Answer contributes its addresses, skipping duplicates, and its ttl. The combined TTL is the minimum over both answers.
  • A failed query only records its error in last_err.
  • The call fails only when no address came back and at least one query failed. It then returns last_err, which is the AAAA error when both failed. A host with only an A record resolves even if the AAAA query times out or returns an error.
  • When both queries succeed with zero records, the result is Ok((vec![], ttl)), and resolve turns it into dns: <host> did not resolve (NotFound).

Two consequences follow from the ordering and the merge:

  • With a protocol backend, IPv4 addresses always come before IPv6 addresses. AddressFamilyStrategy::Auto keeps the resolver’s order, so such an outbound tries IPv4 first.
  • A lookup where one family failed is cached like any other. It stays single-family until its TTL runs out.

protocols/src/dns/message.rs is a stub resolver’s codec: one question out, and the address records of the answer section in. It uses crate::helpers::parse::{take, take_array} and checked arithmetic for every offset, so a hostile response ends in an error, never a panic or an out-of-bounds read. The shared parsing helpers and ProtocolError are described in Protocol foundations.

pub const TYPE_A: u16 = 1;
pub const TYPE_AAAA: u16 = 28;
const CLASS_IN: u16 = 1;
const HEADER_LEN: usize = 12;
const FLAG_QR: u16 = 0x8000;
const FLAG_RD: u16 = 0x0100;
const RCODE_MASK: u16 = 0x000F;
const MAX_POINTER_HOPS: usize = 16;
const MAX_NAME_LEN: usize = 255;
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct Answer {
pub addresses: Vec<IpAddr>,
pub ttl: Option<Duration>,
}
pub fn encode_query(id: u16, name: &str, qtype: u16) -> io::Result<Vec<u8>>;
pub fn decode_answer(id: u16, data: &[u8]) -> io::Result<Answer>;
fn skip_name(data: &[u8], mut at: usize) -> Result<usize, ProtocolError>;

encode_query writes a standard recursive query. All integers are big-endian.

Field Size Value written
ID 2 id, random per query
Flags 2 FLAG_RD (0x0100): a standard query with recursion desired
QDCOUNT 2 1
ANCOUNT, NSCOUNT, ARCOUNT 6 0
QNAME variable one length byte plus the label bytes for each non-empty label, then a 0 root label
QTYPE 2 TYPE_A (1) or TYPE_AAAA (28)
QCLASS 2 CLASS_IN (1)

Encoding rules:

  • Name length. The textual name may be at most MAX_NAME_LEN (255) bytes, or the call fails with dns: name too long (<n> bytes).
  • Label length. Each label may be at most 63 bytes, or the call fails with dns: label longer than 63.
  • Empty labels. The encoder drops empty labels, so example.com. and example.com encode identically.
  • No escaping or IDNA. The encoder copies label bytes as given.
  • No EDNS. The query has no OPT record, so a UDP response is limited to the classic 512 bytes.

decode_answer reads the header, walks past the questions, and then reads ANCOUNT resource records. It never reads the authority and additional sections.

Field Size Check
ID 2 must equal the query’s id, or dns: response id <got> does not match query <id> (InvalidData)
Flags 2 QR (0x8000) must be set, or dns: not a response (InvalidData). The RCODE (low 4 bits) must be 0, or dns: server returned rcode <n> (NotFound).
QDCOUNT 2 number of questions to skip
ANCOUNT 2 number of records to read
NSCOUNT, ARCOUNT 4 not read
Question × QDCOUNT variable skip_name, then 4 bytes of QTYPE and QCLASS
RR NAME variable skip_name
RR TYPE 2 TYPE_A or TYPE_AAAA are used; anything else is skipped
RR CLASS 2 not checked
RR TTL 4 seconds; the minimum over all A and AAAA records becomes Answer.ttl
RDLENGTH 2 the RDATA range must lie inside the buffer, or truncated input: dns rdata
RDATA RDLENGTH 4 bytes for A, 16 for AAAA; any other length is skipped

The record loop is lenient on purpose:

  • Unusable records are skipped, not fatal. A CNAME, an unknown type, or an A record with the wrong RDLENGTH is stepped over with at = rdata_at + rdlen, and the next record is read. A typical answer has a CNAME before the addresses, and it must still resolve.
  • Every A or AAAA record counts toward the TTL, including one skipped for a bad length.
  • The result can be empty. Answer.addresses may be empty with ttl either set or None. The empty case is decided one level up, in query_both and resolve.

A structural failure in the header, the question section or a record boundary fails the whole response. Each failure is a ProtocolError converted to io::Error: Truncated becomes UnexpectedEof, and Overflow becomes InvalidData.

The decoder never needs a name’s text, only where the name ends, so skip_name walks a name in place and never follows a compression pointer:

First byte of a label Meaning skip_name does
0x00 root label returns the offset after it
0b11xx_xxxx compression pointer (2 bytes) returns the offset after the 2-byte pointer, without dereferencing it
0b01xx_xxxx or 0b10xx_xxxx reserved label types fails with truncated input: dns label type
0b00xx_xxxx (1 to 63) literal label advances by 1 plus the label length (checked) and loops

The loop runs at most MAX_POINTER_HOPS (16) times, one iteration per label. Because a pointer ends the walk and is never followed, a pointer cycle cannot loop. The bound caps the labels read in place: a name written with more than 15 literal labels before its terminator fails with truncated input: dns name pointer chain, and the whole response fails with it. That includes the question a server echoes back, so a query for a name with 16 or more labels does not resolve through a protocol backend.

query wraps one exchange in tokio::time::timeout(QUERY_TIMEOUT, ..). The timeout covers everything from socket creation to the last byte read, including the TCP connect and the TLS handshake. On expiry the error is dns: query timed out (TimedOut).

exchange_udp:

  1. Bind a new UdpSocket to 0.0.0.0:0 or [::]:0, the same family as the server, so an IPv6 resolver is reachable.
  2. connect(server), so the kernel delivers only datagrams from the server’s address and port.
  3. send the query once.
  4. recv one datagram into a buffer of UDP_BUF (512) bytes and truncate it to the received length.

There is no retransmission. A lost datagram surfaces as dns: query timed out after QUERY_TIMEOUT. The TC bit is not examined and there is no retry over TCP. A truncated response is decoded as it arrived: the records it still contains are used, and a response that ends inside a record fails with a truncated input error. A datagram longer than 512 bytes is cut to 512 by the receive buffer.

The DoT and DoH backends pool nothing. Each query opens its own TCP connection and completes its own TLS handshake. Because A and AAAA run concurrently, one cache miss costs two connections. For DoT:

sequenceDiagram
  participant A as query(TYPE_A)
  participant B as query(TYPE_AAAA)
  participant S as DoT server
  A->>S: TcpStream::connect, TLS handshake (SNI server_name)
  B->>S: TcpStream::connect, TLS handshake (SNI server_name)
  A->>S: u16 length, A query
  B->>S: u16 length, AAAA query
  S-->>A: u16 length, response
  S-->>B: u16 length, response
  Note over A,B: each stream is dropped after its one response

DoH follows the same shape, with the POST request in place of the length prefix and read_to_end in place of the length-delimited read. What they do share is the SslConnector inside Inner.tls, built once in Resolver::new. The rationale in the source is that a warm cache rarely consults the resolver, and a held-open connection would need revalidation anyway. If you add pooling, keep the whole exchange inside QUERY_TIMEOUT, and keep one response per request ID.

Resolver::new calls protocols/src/transports/tls/config.rs → ClientConfig::with_verify_mode:

pub fn with_verify_mode(
sni: impl AsRef<str>,
verify_mode: VerifyMode,
alpn: Alpn,
) -> io::Result<Self>;
Backend sni verify_mode alpn
Tls server_name VerifyMode::CustomCa(pem) when ca_pem is set, otherwise VerifyMode::System Alpn::None
Https host from the URL same Alpn::Http1

Both modes verify the chain and the host name, and both require TLS 1.2 or newer. The resolver never uses VerifyMode::Insecure. ClientConfig::wrap hands the name to OpenSSL’s into_ssl, so a name that parses as an IP address (a server_name such as 192.0.2.53, or an IPv4 URL host) is sent without SNI and verified against the certificate’s IP address entries.

VerifyMode::CustomCa trusts the configured CA in addition to the system roots. It calls set_default_verify_paths() first and then adds every certificate parsed from the PEM to the connector’s own store; the host’s trust store is not modified. A private resolver with its own CA therefore works, and a public resolver still verifies against the system roots. A PEM with no certificate in it fails the build with no certificate in CA PEM bundle.

Invariant Enforced by Pinned by
A protocol backend asks the configured server Backend::Udp(addr) passed to exchange_udp the_udp_backend_resolves_against_the_configured_server
DoT uses the 2-byte length framing and DoH an RFC 8484 POST, both over a verified handshake exchange_tls, exchange_https_over dot_resolves_over_tls, doh_resolves_over_https (the DoH test server asserts the request line, Content-Type and Host)
An address literal never reaches the backend or the cache the host.parse::<IpAddr>() early return in resolve an_address_literal_never_reaches_the_backend
A repeat lookup inside the TTL is answered from the cache Cached.expires_at compared with Instant::now() a_second_lookup_is_served_from_cache
An expired entry is looked up again, not served stale the same comparison an_expired_entry_is_looked_up_again (uses with_ttl_bounds)
Clones share one cache Resolver { inner: Arc<Inner> } a_clone_shares_the_same_cache
The cache stays bounded however many distinct names clients send max_capacity(CACHE_CAPACITY) the_cache_stays_bounded_under_distinct_names, which inserts CACHE_CAPACITY / 4 + 500 names, fewer than the capacity, so it would also pass without the bound; it does not exercise eviction
A lookup that yields no address is an error, not Ok(vec![]) the last_err return in query_both and the addresses.is_empty() check in resolve a_name_that_does_not_resolve_is_an_error_not_an_empty_answer, which points the resolver at an unreachable server, so it pins the both-queries-failed case
The default backend is the host resolver impl Default for Resolver → Resolver::system() the_default_backend_is_the_host_resolver
A backend missing a required field is refused, never replaced by system the needed closure and the ok_or_else checks in from_spec a_backend_missing_its_fields_is_refused; end to end, a_backend_without_its_server_is_rejected
A DoH URL yields only a host and a path; a URL port and userinfo never reach the verified name Backend::https a_doh_url_is_split_into_authority_and_path
A response must carry the query’s ID, set QR, and have RCODE 0 decode_answer header checks rejects_a_mismatched_id_a_question_and_an_error_rcode
Truncated or malformed input returns an error and never panics take, take_array and checked_add on every offset truncated_and_malformed_input_never_panics
A compression-pointer cycle terminates skip_name never follows pointers, and its loop is bounded by MAX_POINTER_HOPS a_compression_pointer_loop_terminates
One unrecognised record does not discard the addresses beside it the _ => {} arm in the record loop an_unrecognised_record_is_skipped_not_fatal
The shortest address TTL governs min over A and AAAA TTLs decodes_addresses_and_the_smallest_ttl
Queries are well-formed and bounded encode_query checks a_query_round_trips_through_its_own_encoder, a_trailing_dot_encodes_the_same_as_without, rejects_an_over_long_name_and_label
An HTTP error from a DoH server is an error, not an answer the status-line check in exchange_https_over doh_surfaces_an_http_error
A resolver pinned to the wrong CA fails the handshake chain verification under VerifyMode::CustomCa: the server’s certificate chains to neither the pinned CA nor a system root a_resolver_pinned_to_the_wrong_ca_is_refused
The configured backend is what the app actually uses one Resolver::from_spec in build, passed to every outbound the_udp_backend_is_used_for_resolution together with its negative control the_default_backend_does_not_know_the_test_name
One failed family does not discard the other last_err is returned only when no address came back the_udp_backend_is_used_for_resolution: its fake server answers A with an address and AAAA with NXDOMAIN (RCODE 3)

Three behaviours have no dedicated test:

  • Two successful answers with no address records. Every fake server answers A with an address, and the negative tests make both queries fail, so the path where both queries succeed empty and resolve returns dns: <host> did not resolve is exercised only in production.
  • The DoH Connection: close header. The DoH test server asserts the request line, Content-Type and Host, but not Connection.
  • Case-insensitive cache keys. No test resolves the same name in two spellings.

resolve returns io::Result, and callers such as resolve_candidates pass the error up unchanged. What happens to the connection then belongs to the outbound. All the messages below start with dns: , except the ProtocolError and TLS ones.

Stage Message io::ErrorKind
build dns: the udp backend needs a server address (also tls, https) InvalidInput
build dns: invalid server address: <parse error> InvalidInput
build dns: the tls backend needs a server name to verify against InvalidInput
build dns: the https backend needs the resolver's url InvalidInput
build dns: unknown backend "<name>" (expected "system", "udp", "tls" or "https") InvalidInput
build dns: "<url>" is not an https:// url, dns: "<url>" has no usable host InvalidInput
build no certificate in CA PEM bundle InvalidInput
build an OpenSSL error from ClientConfig::with_verify_mode, for example a certificate block that does not parse Other (ossl_err)
resolve dns: <host> did not resolve NotFound
resolve (System) the host resolver’s error, passed through unchanged from tokio::net::lookup_host as returned by the standard library
query dns: name too long (<n> bytes), dns: label longer than 63 InvalidInput
query dns: query timed out TimedOut
query socket, connect or TLS errors as returned by tokio and OpenSSL
decode dns: response id <got> does not match query <id>, dns: not a response InvalidData
decode dns: server returned rcode <n> (NXDOMAIN is 3, SERVFAIL is 2) NotFound
decode truncated input: dns name, truncated input: dns label type, truncated input: dns name pointer chain, truncated input: dns rdata, truncated input: fixed-size field UnexpectedEof
decode integer overflow: dns name, integer overflow: dns rr, and similar InvalidData
DoH dns: no http header, dns: doh server answered "<status line>" InvalidData

The app logs build errors as configuration invalid: <message> under --test and as failed to start: <message> at start, then exits with a failure status. On reload it logs reload: build failed, keeping current config: <message>. katana’s build_outbounds returns them as io::Error from the pool build. At start katana logs failed to build outbounds: <message> and exits with a failure status; its --test prints configuration error: <message> to standard error. On a reload that rebuilds the pool it logs reload: bad outbounds, keeping current config: <message>, and the current pool stays in place.

Cancellation. The resolver spawns no task. resolve is a plain future: dropping it drops the join! of both queries, and with them the sockets, TLS streams and timers. A lookup cancelled before it finishes inserts nothing into the cache. moka’s future::Cache does its housekeeping inside get and insert calls, so no background worker outlives a generation. A resolver dies when the last clone held by the generation’s outbounds and probes is dropped.

Timeout scope. QUERY_TIMEOUT applies to each protocol query, not to resolve as a whole. The A and AAAA queries run concurrently, so a protocol lookup takes at most about QUERY_TIMEOUT. The System backend is not wrapped in QUERY_TIMEOUT. It takes as long as the host resolver’s own timeout and retry settings allow.

Name Value Where Effect
CACHE_CAPACITY 8192 entries dns/mod.rs moka max_capacity; the key is a client-chosen name, so the cache must be bounded
MIN_TTL 5 s dns/mod.rs floor for every TTL, so a zero-TTL answer is still cached briefly
MAX_TTL 3600 s dns/mod.rs ceiling for every TTL, and moka’s per-cache time_to_live backstop
SYSTEM_TTL 60 s dns/mod.rs TTL assumed for getaddrinfo answers, which report none
QUERY_TIMEOUT 5 s dns/mod.rs per A or AAAA query, from socket creation to the last byte
UDP_BUF 512 bytes dns/mod.rs UDP receive buffer; no EDNS
HEADER_LEN 12 bytes dns/message.rs fixed DNS header
MAX_NAME_LEN 255 bytes dns/message.rs longest name encode_query accepts (textual length)
label length 63 bytes dns/message.rs → encode_query longest single label
MAX_POINTER_HOPS 16 dns/message.rs iterations of skip_name; at most 15 literal labels before a terminator
DoT length prefix u16 dns/mod.rs → exchange_tls largest DoT message is 65535 bytes
TLS versions 1.2 and newer transports/tls/config.rs set by with_verify_mode
File What it covers
protocols/tests/unit/dns/mod.rs Resolver against an in-process UDP server that counts queries: literals, cache hits, shared clones, the capacity bound, expiry, an unreachable server, from_spec refusals, Backend::https URL splitting, and the default backend resolving localhost
protocols/tests/unit/dns/message.rs the codec: round trip, trailing dot, A and AAAA decoding, smallest TTL, CNAME skipping, ID, QR and RCODE rejection, every truncation point, a self-referencing pointer, and over-long names and labels
protocols/tests/pipeline/dns_secure.rs DoT and DoH against a real OpenSSL server with a self-signed certificate passed as ca_pem: dot_resolves_over_tls, doh_resolves_over_https, doh_surfaces_an_http_error, a_resolver_pinned_to_the_wrong_ca_is_refused
app/tests/integration/e2e_dns.rs the built etemenanki-app binary behind a SOCKS inbound: a name only the fake UDP server knows resolves with backend = "udp" (the server answers its AAAA query with NXDOMAIN, so this also covers one failed family) and fails with the default backend; backend = "udp" without server fails --test

The unit tests are compiled into the library through #[path] modules, so they run with the crate’s lib tests:

Terminal window
cargo test -p etemenanki-protocols --lib dns::
cargo test -p etemenanki-protocols --test pipeline dns_secure
cargo test -p etemenanki-app --test integration e2e_dns

an_expired_entry_is_looked_up_again uses real time rather than a paused clock. A paused runtime auto-advances while waiting on the real socket, which would trip QUERY_TIMEOUT before the TTL. Keep that in mind before converting DNS tests to tokio::time::pause.

When you change the codec, add a malformed-input case next to truncated_and_malformed_input_never_panics. When you change a backend, extend dns_secure.rs: a live server is the only place where framing, the handshake and header handling are all exercised together.