Skip to content

Generations and hot reload

Source files: 18 · checked against Etemenanki 596916d
  • Etemenanki/app/src/instance.rs
  • Etemenanki/app/src/main.rs
  • Etemenanki/app/src/serve.rs
  • Etemenanki/app/src/config.rs
  • Etemenanki/app/src/inbound/mod.rs
  • Etemenanki/app/src/inbound/tun.rs
  • Etemenanki/app/src/balancer.rs
  • Etemenanki/protocols/src/hysteria/server/inbound.rs
  • Etemenanki/protocols/src/tun/inbound.rs
  • Etemenanki/protocols/src/tun/device.rs
  • Etemenanki/app/tests/integration/e2e_hysteria_inbound.rs
  • Etemenanki/app/tests/integration/e2e_unix.rs
  • Etemenanki/app/tests/integration/e2e_tun.rs
  • Etemenanki/app/tests/support/mod.rs
  • Etemenanki/app/tests/unit/config.rs
  • Etemenanki/app/tests/unit/inbound.rs
  • Etemenanki/app/tests/unit/serve.rs
  • Etemenanki/protocols/tests/pipeline/tun.rs

A running etemenanki-app holds exactly one generation: the bound listeners, UDP sockets and TUN devices of one configuration, together with every task that serves them, all hanging off a single CancellationToken. Hot reload never patches a generation in place. It builds a complete replacement from the file, cancels the old generation, waits until the old generation has given back its ports, socket paths and interfaces, and then binds the new one.

This page follows that cycle through app/src/instance.rs and app/src/main.rs, down to the locks, tokens, timeouts and log lines involved. Read it before you change how inbounds bind, add an inbound kind that owns a handle, or touch the watcher. What build does with each config section is covered in the build pipeline page; how an accepted socket becomes a relay is covered in the serving page.

Component Owns Does not do
app/src/main.rs → main CLI parsing, tracing setup, --test, the watcher, waiting for signals Any config interpretation beyond [log].level
app/src/main.rs → spawn_watcher Watching the config’s directory and turning events into Instance::reload calls Deduplication (left to reload)
app/src/instance.rs → Instance The current Config, the bytes it came from, and the live Generation Serving connections
app/src/instance.rs → build Validating a Config into a Built: resolver, outbounds, balancers, router, inbounds Binding anything
app/src/instance.rs → spawn_generation Binding each inbound and spawning its accept task under a fresh token Tearing down the previous generation
app/src/serve.rs, app/src/inbound/tun.rs The accept tasks, and releasing their handle before they return Deciding when a generation ends
app/src/instance.rs
pub struct Instance {
path: PathBuf,
state: tokio::sync::Mutex<State>,
}
struct State {
config: Config,
last_bytes: Vec<u8>,
generation: Generation,
}
struct Generation {
token: CancellationToken,
accept_handles: Vec<JoinHandle<()>>,
}
  • path is the -c path exactly as given (default config.toml). It never changes for the life of the process; every reload re-reads the same path.
  • state is a tokio::sync::Mutex because reload and shutdown hold it across .await points: while the old generation drains and while the new one binds. Holding it that long is deliberate. It serialises reloads against each other and against shutdown, so two generations never try to bind at once and a signal that arrives mid-reload waits for the reload to finish.
  • last_bytes is the byte fingerprint of the last file content reload acted on, which is not always the content the running generation came from (see The byte fingerprint).
  • Generation is deliberately small: one token that every task of the generation observes, and the JoinHandle of each accept task (one per successfully bound inbound). Connection tasks, sub-stream tasks and balancer probes are not in accept_handles. They are reached only through the token.
app/src/instance.rs
pub struct Built {
router: Arc<Router>,
inbounds: Vec<BuiltInbound>,
balancers: Vec<(Arc<Balancer>, Duration, Duration)>,
resolver: Resolver,
}
struct BuiltInbound {
tag: CompactString,
bind: BindSpec,
kind: InboundKind,
}
pub fn build(cfg: &Config) -> io::Result<Built>

build is the whole of validation. It is shared by Instance::start, Instance::reload and the --test dry run (main.rs → test_config), so a config that --test accepts is one that start and reload accept too, up to the point of binding. It reads every file the config references (certificates, keys, the DNS CA file, geodata) and constructs a new Resolver, new outbounds, new balancers and a new Router. Nothing in a Built is shared with the previous generation, so a reload also starts with an empty DNS cache and with every balancer member marked healthy (Member::new initialises healthy to true).

A Built owns the balancers’ probe parameters rather than starting the probes itself, so that the probes can be started against the token of the generation that will own them.

build_inbound (in app/src/inbound/mod.rs) lowers each inbound into an InboundKind plus a BindSpec that says what the runtime must bind for it. Deciding this at build time is what lets --test refuse an inbound that a start would refuse, for example a Unix listen with a port, or a TUN inbound with a listen.

app/src/inbound/mod.rs
pub enum BindSpec {
Tcp { host: String, port: u16 },
Udp { host: String, port: u16 },
Unix(PathBuf),
Tun(DeviceSpec),
}
pub enum InboundKind {
Stream(Arc<StreamInbound>),
Hysteria2(Hy2Inbound<()>),
Tun(TunInbound<()>),
}
app/src/instance.rs
enum Listener {
Stream(StreamListener),
Udp(std::net::UdpSocket),
Tun(OwnedFd),
}
async fn bind_inbound(ib: &BuiltInbound) -> io::Result<Listener>

BindSpec’s Display implementation is what appears in the log lines: 127.0.0.1:1080, udp 0.0.0.0:443, unix:/run/etemenanki/in.sock, tun auto or tun ete0.

app/src/serve.rs
pub fn spawn_scoped<F>(token: CancellationToken, fut: F) -> JoinHandle<()>
where
F: Future + Send + 'static,
F::Output: Send,

spawn_scoped wraps fut in a tokio::select! against token.cancelled(). When the token fires, the select completes and fut is dropped, together with everything it owns: the client socket, the runtime’s buffers, the outbound stream, the permits. This is the single mechanism that ends every per-connection task of a stream inbound. The stream accept loop spawns each accepted socket with it, and serve_socket spawns each byte stream a transport yields (for example every gRPC stream of one connection) with it too.

The Hysteria 2 and TUN inbounds do not use spawn_scoped for their per-connection work. They keep it in a JoinSet owned by their run future, so dropping that future aborts every connection. Their run future returns when the token fires, or earlier when the endpoint closes or the device goes away.

app/src/serve.rs
pub async fn run_stream_inbound(
inbound: Arc<StreamInbound>,
tag: CompactString,
listener: StreamListener,
router: Arc<Router>,
token: CancellationToken,
)
pub async fn run_hysteria_inbound(
inbound: Hy2Inbound<()>,
tag: CompactString,
socket: std::net::UdpSocket,
router: Arc<Router>,
token: CancellationToken,
)
app/src/inbound/tun.rs
pub async fn run_tun_inbound(
inbound: TunInbound<()>,
tag: CompactString,
fd: OwnedFd,
router: Arc<Router>,
token: CancellationToken,
)

Each accept task has the same contract, and reload depends on it: the task returns only after the handle it was given has been released. For a TCP listener that means the listener has been dropped. For a Unix listener the socket file has also been unlinked. For Hysteria 2 the UDP port can be bound again, and for TUN the device descriptor has closed. The release step for each kind:

app/src/serve.rs
impl StreamListener {
fn release(self)
}
protocols/src/hysteria/server/inbound.rs
impl<T> Hy2Inbound<T> {
pub async fn shutdown(&self)
}
protocols/src/tun/inbound.rs
impl<T> TunInbound<T> {
pub async fn shutdown(&self)
}

main initialises tracing first. init_tracing loads the config once, only to read [log].level (defaulting to info, also when the file does not load); RUST_LOG, when it is set to a valid filter, takes precedence. The filter is installed once and never replaced, which is why [log].level is not reloaded.

With --test, main calls test_config, which is config::load followed by instance::build, and exits. It prints Configuration OK. on success, and otherwise logs configuration invalid: … and exits with status 1. It binds nothing, so errors that only binding can reveal (a port in use, a non-socket file at a Unix path, missing TUN privileges) do not show up.

Otherwise main calls Instance::start:

app/src/instance.rs
impl Instance {
pub async fn start(path: PathBuf) -> io::Result<Self>
pub async fn reload(&self)
pub async fn shutdown(&self)
}

start reads the file, runs config::parse_bytes and build, then calls spawn_generation(built, true). Any error is returned, main logs failed to start: …, and the process exits with status 1. On success the read bytes become the first last_bytes. main then wraps the Instance in an Arc, starts the watcher, and waits for a signal.

app/src/instance.rs
async fn spawn_generation(built: Built, strict: bool) -> io::Result<Generation>

spawn_generation creates a fresh CancellationToken, starts each balancer’s probes under it, then binds the inbounds in config order, one at a time, and spawns an accept task for each one that binds.

flowchart TB
  start["new CancellationToken"] --> probes["Balancer::spawn_probe per balancer"]
  probes --> next{"next inbound?"}
  next -- "no" --> done["Ok(Generation)"]
  next -- "yes" --> bind["bind_inbound"]
  bind -- "Ok" --> pair["match InboundKind with Listener"]
  pair --> spawn["tokio::spawn the accept task, push its JoinHandle"]
  spawn --> next
  bind -- "Err" --> strict{"strict?"}
  strict -- "yes (startup)" --> abort["token.cancel(), return Err"]
  strict -- "no (reload)" --> log["log the bind failure"]
  log --> next
app/src/balancer.rs
pub fn spawn_probe(
self: &Arc<Self>,
token: CancellationToken,
interval: Duration,
timeout: Duration,
resolver: Resolver,
)

spawn_probe starts one task per member. Each task probes at once, records the result, and then waits in a select! between token.cancelled() and a sleep(interval), so a probe loop ends at the first wait after its generation is cancelled. The interval and timeout default to DEFAULT_PROBE_INTERVAL (30 s) and DEFAULT_PROBE_TIMEOUT (5 s) when the balancer does not set probe_interval or probe_timeout. Probes are not in accept_handles, and nothing waits for them.

BindSpec Bind call Listener Accept task
Tcp std::net::TcpListener::bind, then set_nonblocking(true) and tokio::net::TcpListener::from_std Stream(StreamListener::Tcp) run_stream_inbound
Unix stale-socket check, then std::os::unix::net::UnixListener::bind, set_nonblocking(true) and tokio::net::UnixListener::from_std Stream(StreamListener::Unix) (keeps the path) run_stream_inbound
Udp std::net::UdpSocket::bind Udp run_hysteria_inbound, which turns the socket into a quinn endpoint inside Hy2Inbound::run
Tun etemenanki_protocols::tun::open(spec), which creates, addresses and routes the interface, then logs inbound <tag> owns tun device <name> Tun(OwnedFd) run_tun_inbound

On success spawn_generation logs inbound <tag> listening on <bind>. build_inbound pairs every kind with its bind, so a mismatched (InboundKind, Listener) pair cannot occur. The match still has a fallback arm that logs inbound <tag>: listener does not match its kind and skips the inbound rather than panicking.

A Unix listen path can already exist when the inbound binds: the process crashed, a failed start left it behind, or the previous generation’s socket is still there. bind_inbound inspects the path with std::fs::symlink_metadata (so a symlink is judged as itself, not as its target):

What is at the path Result
Nothing (NotFound) Bind.
A socket (FileTypeExt::is_socket) Remove it, then bind. The socket is treated as stale and taken over.
Anything else Fail with ErrorKind::AlreadyExists and <path> exists and is not a socket.
Metadata error other than NotFound Fail with that error.

A regular file at the path therefore stops a start (timestamp removed):

ERROR etemenanki_app: failed to start: inbound u bind unix:/run/etemenanki/in.sock failed: /run/etemenanki/in.sock exists and is not a socket

The check does not ask whether another process is still listening on the socket. A socket file at the configured path is always considered this inbound’s to take.

Startup (strict = true) Reload (strict = false)
First bind failure Cancels the token, returns Err with inbound <tag> bind <bind> failed: <error> Logs inbound <tag> bind <bind> failed: <error> at ERROR
Remaining inbounds Not attempted Still bound and started
Result failed to start: …, exit status 1 Ok(Generation) with fewer accept handles

Startup is all-or-nothing, so a process that is running is running its whole configuration. Reload is lenient because by the time the new generation binds, the old one is already gone: failing the whole generation would take every inbound down to punish one. With strict = false, spawn_generation cannot return Err, so the Err arm after it in reload (which logs reload: failed to start new generation: …) is defensive.

When a strict start fails partway, the inbounds bound before the failing one have already had their accept tasks spawned. The token cancels them, and the process exits right after. A Unix socket file they created can be left on disk, and the next start takes it over.

app/src/main.rs
fn spawn_watcher(
path: PathBuf,
instance: Arc<Instance>,
) -> notify::Result<notify::RecommendedWatcher>

The watcher uses notify 8 (notify::recommended_watcher, which is inotify on Linux) and watches the config file’s parent directory, non-recursively. It watches the directory rather than the file because many editors save by writing a new file and renaming it over the old one. A watch on the old inode would see the rename once and then go quiet. When the path has no parent component (-c config.toml), the directory is ..

The notify callback runs on notify’s own thread. It sends () into an unbounded tokio::sync::mpsc channel for every event that is Ok, and drops error events. It does not look at which file the event names or what kind of event it is: any event in the directory is a reason to try. That includes a change to a certificate or geodata file that lives next to the config, and, because the watch mask that notify 8.2 installs on Linux includes IN_OPEN, a file in the directory merely being opened.

A task spawned on the runtime consumes the channel:

  1. Wait for one event.
  2. Drain everything already queued with try_recv.
  3. Sleep 200 ms (Duration::from_millis(200), not a named constant).
  4. Drain again, collapsing an editor’s burst of create, write, rename and attribute events into one attempt.
  5. instance.reload().await, then go back to 1.

Because the task awaits reload, attempts never overlap. Events that arrive during a long reload (a Hysteria 2 drain can take seconds) queue up and cause one more attempt afterwards, which the byte check usually turns into a no-op.

main keeps the returned RecommendedWatcher alive in _watcher for the life of the process; dropping it would stop the watch. If creating the watcher or the watch fails, main logs config hot-reload disabled: <error> and runs on without hot reload. That failure is not fatal.

sequenceDiagram
  participant W as watcher task
  participant R as Instance::reload
  participant F as config file
  participant G0 as old generation
  participant G1 as new generation
  W->>R: reload().await
  R->>R: lock state (tokio Mutex)
  R->>F: std::fs::read(path)
  F-->>R: bytes
  alt read error or bytes == last_bytes
    R-->>W: return, generation untouched
  else changed bytes
    R->>R: config::parse_bytes, then build
    alt parse or build error
      R->>R: last_bytes = bytes
      R-->>W: return, old generation keeps serving
    else built
      R->>R: log "config reload: " + config::diff
      R->>G0: token.cancel()
      G0->>G0: accept tasks release listener, socket file, UDP port, TUN fd
      R->>G0: await every accept JoinHandle
      R->>G1: spawn_generation(built, false)
      G1-->>R: Generation (failed binds logged, skipped)
      R->>R: store generation, config and last_bytes
    end
  end

In detail:

  1. Lock. reload takes state.lock().await and holds it to the end.
  2. Read. std::fs::read(&self.path). On error it logs reload: cannot read <path>: <error> and returns without touching last_bytes. A config that is briefly missing during an editor’s save therefore costs nothing: when the file reappears with the same content, the next attempt sees identical bytes.
  3. Byte dedup. If bytes == state.last_bytes, it returns silently. This absorbs every watcher event that did not change the config, for example events for other files in the directory.
  4. Parse. config::parse_bytes: UTF-8 check, then toml::from_str into Config. The config structs use deny_unknown_fields, so a misspelt key fails here rather than falling back to a default. The per-protocol settings tables are kept as opaque TOML values at this stage, so a misspelt key inside settings fails in step 5 instead. On error it logs reload: parse failed, keeping current config: <error>, sets last_bytes = bytes, and returns.
  5. Build. build(&new_cfg). On error it logs reload: build failed, keeping current config: <error>, sets last_bytes = bytes, and returns. Build reads referenced files, so an unreadable certificate fails here.
  6. Log the diff. config reload: <diff> at INFO (see The diff line).
  7. Cancel the old generation. state.generation.token.cancel(). From this point on the running configuration is gone, whatever happens next.
  8. Await the old accept tasks. for h in accept_handles.drain(..) { let _ = h.await; }. All accept tasks saw the cancellation at the same moment and release concurrently, so the loop waits for the slowest one, not the sum of all of them. A panicked task’s JoinError is ignored.
  9. Spawn the new generation. spawn_generation(built, false).
  10. Commit. Store the new Generation, config = new_cfg and last_bytes = bytes.

Steps 4 and 5 are the transactional part: nothing is torn down until the new configuration has parsed and built completely. Steps 7 to 9 are not transactional. Binding happens after the old generation has released its handles, because the new generation usually needs the very same ports, paths and interface names.

Case Log last_bytes Running generation
File unreadable ERROR reload: cannot read … unchanged old, untouched
Bytes identical to last_bytes none unchanged old, untouched
Parse error ERROR reload: parse failed, keeping current config: … set to the new bytes old, untouched
Build error ERROR reload: build failed, keeping current config: … set to the new bytes old, untouched
Success INFO config reload: …, then one listening on or bind … failed line per inbound set to the new bytes new, possibly with some inbounds down

A failed parse, captured from a run of the debug binary (timestamps and colours removed):

INFO etemenanki_app::instance: inbound a listening on 127.0.0.1:18731
ERROR etemenanki_app::instance: reload: parse failed, keeping current config: TOML parse error at line 12, column 12
|
12 | garbage = [
| ^
unclosed array, expected `]`

last_bytes records the last content reload made a decision about, and it is updated on a parse or build failure as well as on success. That choice has consequences that are easy to trip over:

  • The same broken content is not retried. Touching the file, or saving it again unchanged, is a no-op after a failure. A build failure caused by something outside the file (a certificate that is not there yet) is therefore not retried when the certificate appears; the config bytes must change.
  • A failed bind is not retried. After a successful reload in which one inbound failed to bind, last_bytes holds the new content. Freeing the port does not bring the inbound back until the config bytes change again, for example by editing a comment.
  • Reverting a broken edit reloads. If the file is restored byte-for-byte to the running configuration after a failed parse, the bytes differ from last_bytes (the broken content), so a full generation swap follows. The diff line reads no changes, and connections are still dropped.
  • Files referenced by the config are not watched as content. Replacing a certificate, key, CA file or geodata file triggers a watcher event (when it lives in the same directory) but no reload, because the config bytes are the same. The new file is read on the next reload that the config file itself causes.
app/src/config.rs
pub fn parse_bytes(bytes: &[u8]) -> io::Result<Config>
pub struct ConfigDiff {
pub inbounds_added: Vec<CompactString>,
pub inbounds_removed: Vec<CompactString>,
pub inbounds_changed: Vec<CompactString>,
pub outbounds_added: Vec<CompactString>,
pub outbounds_removed: Vec<CompactString>,
pub outbounds_changed: Vec<CompactString>,
pub route_changed: bool,
pub log_changed: bool,
}
pub fn diff(old: &Config, new: &Config) -> ConfigDiff

diff matches inbounds and outbounds by tag and compares the parsed structs with PartialEq. Its Display form joins groups with ; : inbounds +[b] -[a] ~[c], outbounds ~[proxy], route changed, log changed, or no changes when nothing differs. Renaming a tag shows up as one addition and one removal.

The diff is for the log only. It does not decide what is rebuilt: every successful reload replaces everything. It also does not look at [dns] or [[balancer]], so a reload that changes only those logs config reload: no changes and still swaps the generation. Formatting-only edits (comments, whitespace) log no changes as well.

Cancelling the token reaches each kind of task differently:

Task How it ends What its accept task waits for before returning
Stream accept loop (run_stream_inbound) Its select! takes the token.cancelled() branch and breaks. A loop sleeping ACCEPT_ERROR_BACKOFF after an accept error wakes on the token too (backoff_or_cancelled). StreamListener::release: drops the listener and, for a Unix listener, remove_files the path (a failure is logged at DEBUG).
Stream connections and transport sub-streams Each spawn_scoped wrapper drops its future. Nothing. They are not awaited, and they finish on their own as the runtime polls them.
Hysteria 2 (run_hysteria_inbound) Hy2Inbound::run breaks out of its accept loop and returns, dropping its JoinSet and with it every connection task. Hy2Inbound::shutdown: endpoint.close(CLOSE_CODE, b""), wait_idle bounded by DRAIN_TIMEOUT, drop the endpoint, then try UdpSocket::bind(local) every RELEASE_POLL until it succeeds or RELEASE_TIMEOUT passes.
TUN (run_tun_inbound) TunInbound::run returns, dropping the IP stack and every flow. TunInbound::shutdown: poll every RELEASE_POLL until the device’s Weak handle no longer upgrades and the live TCP count is zero, bounded by RELEASE_TIMEOUT.
Balancer probes Their select! takes the cancelled branch at the next wait. Not awaited.

The Hysteria 2 wait exists because dropping a quinn endpoint does not release its UDP socket while connections are live. The socket belongs to the endpoint driver task, which drops it on a later poll. Without the explicit drain and the bind check, the new generation’s UdpSocket::bind would fail with “address in use”, and with lenient binding the inbound would stay down, visible only as one bind-failure line in the log. The shutdown checks the port itself, because a free port is the condition the next generation needs. If the port is still busy after RELEASE_TIMEOUT, it logs hysteria2: <addr> did not come free within 3s at WARN and returns anyway.

The TUN wait exists for the same reason at the interface level. Dropping the IpStack aborts the task that owns the device, but the abort lands only when the runtime schedules it, so the descriptor briefly outlives the drop. Recreating an interface with the same name before that fails. On timeout it logs tun: device fd or <n> tcp flows still open after 3s at WARN.

A TCP listener needs no such wait. Dropping it closes the listening socket at once, and the standard library sets SO_REUSEADDR on Unix, so connections of the old generation lingering in TIME_WAIT do not block the rebind.

Hy2Inbound::shutdown closes the old endpoint with QUIC application error code 0x100 (CLOSE_CODE). The upstream hysteria client reconnects to the new generation on its own, which a_reload_rebinds_the_udp_port relies on.

stateDiagram-v2
  [*] --> Built: parse_bytes and build succeed
  Built --> Binding: spawn_generation
  Binding --> Serving: all inbounds attempted
  Binding --> [*]: strict bind failure at startup
  Serving --> Cancelled: reload commits or shutdown
  Cancelled --> Released: every accept task returned
  Released --> [*]

A Built that is never spawned (because the process is only running --test) is dropped without having bound anything. Between Cancelled and the next generation’s Binding, no inbound is listening.

  • Every successful reload drops every connection, on every inbound, including inbounds whose configuration did not change and edits that only touch a comment. There is no connection hand-over between generations.
  • There is a gap. Between cancelling the old generation and binding the new one, nothing listens. For TCP and Unix inbounds the gap is short. With a Hysteria 2 inbound holding live connections or a TUN inbound it can last up to the release timeouts above, and it applies to every inbound, because binding starts only after the slowest accept task has returned.
  • A bind failure on reload leaves that inbound down. The other inbounds start, the old listener is already gone, and the inbound stays down until a later reload with different config bytes binds it.
  • A broken config keeps the old generation running, with its connections intact, and logs why at ERROR.
  • [log].level is not reloaded. The diff still reports log changed.
  • The -c path is fixed for the life of the process.

The user guide’s hot-reload page describes the same behaviour from the operator’s side.

app/src/main.rs
async fn wait_for_shutdown()

On Unix, wait_for_shutdown registers SIGTERM with tokio::signal::unix::signal and waits in a select! for either SIGTERM or tokio::signal::ctrl_c() (SIGINT). If SIGTERM cannot be registered, it waits for SIGINT alone. On other platforms it waits for Ctrl-C only.

When a signal arrives, main logs shutting down and calls Instance::shutdown, which takes the state lock (so an in-progress reload finishes first), cancels the current token and awaits every accept handle, exactly like steps 7 and 8 of a reload. main then returns ExitCode::SUCCESS. Dropping the runtime at the end of main drops every task that is still alive.

Going through the accept tasks rather than just exiting matters for the handles that outlive the process otherwise. The Unix socket file is unlinked by StreamListener::release. The TUN device and the routes the inbound installed disappear once the descriptor has closed, which TunInbound::shutdown waits for.

Invariant Enforced by Pinned by
A config that fails to parse or build never replaces the running generation. reload returns before token.cancel() on any parse or build error. Not covered by an automated test. a_mistyped_key_is_rejected_rather_than_ignored in app/tests/unit/config.rs pins that typos fail at parse time.
--test, start and reload accept the same configs up to binding. All three call the same config::parse_bytes and build. BindSpec is chosen inside build_inbound. a_unix_listen_takes_no_port, a_socket_listen_needs_a_port, hysteria2_refuses_a_unix_listen, tun_owns_its_interface_and_takes_no_listener in app/tests/unit/inbound.rs.
Every task of a generation ends with it. spawn_scoped for stream connections and sub-streams, the JoinSets owned by Hy2Inbound::run and TunInbound::run, the token passed to Balancer::spawn_probe. a_reload_rebinds_the_udp_port in app/tests/integration/e2e_hysteria_inbound.rs reloads while the old generation holds a live client connection.
The old generation has released its ports, socket paths and devices before the new one binds. reload awaits every accept handle, and each accept task returns only after release or shutdown. Hy2Inbound::shutdown waits on the port itself. a_reload_rebinds_the_udp_port (UDP port reused after reload). socks_over_a_unix_socket_relays_and_cleans_up in app/tests/integration/e2e_unix.rs and a_routed_connect_is_answered_while_the_app_runs in app/tests/integration/e2e_tun.rs pin the release on shutdown.
The old Unix listener’s cleanup cannot unlink the new generation’s socket. Ordering: the old accept task’s remove_file has completed before spawn_generation binds the path again. Follows from the previous row. No dedicated test.
Only a socket at a Unix listen path is taken over. symlink_metadata and is_socket in bind_inbound; anything else is AlreadyExists. Not covered by an automated test.
Reloads never overlap each other or a shutdown. The watcher task awaits each reload, and reload and shutdown both hold the tokio::sync::Mutex<State> throughout. Structural.
A running process has bound its whole configuration at startup. spawn_generation(built, true) aborts on the first bind failure. Not covered by an automated test.
Per-inbound caps start fresh with each generation. The accept semaphores (MAX_HANDSHAKES_PER_INBOUND, MAX_LIVE_CONNECTIONS_PER_INBOUND), the Hysteria 2 connection semaphore and the TUN flow semaphore are created inside the accept task. Structural.
Failure Where Log Effect
Watcher cannot be created spawn_watcher config hot-reload disabled: … Proxy runs without hot reload.
Watch error event notify callback none Event ignored.
Config unreadable on reload reload step 2 reload: cannot read … Old generation kept, last_bytes unchanged.
Parse or build error on reload reload steps 4 and 5 reload: parse failed, … / reload: build failed, … Old generation kept, bytes remembered.
Bind error on reload spawn_generation, lenient inbound <tag> bind <bind> failed: … That inbound down, others started.
Bind error at startup spawn_generation, strict failed to start: inbound <tag> bind <bind> failed: … Exit status 1.
Hysteria 2 endpoint cannot be created from the bound socket Hy2Inbound::run hysteria2 inbound failed: … That inbound’s accept task ends; the inbound is down until the next generation.
TUN stack cannot start on the device TunInbound::run tun inbound failed: … Same as above.
Hysteria 2 port or TUN device not released in time Hy2Inbound::shutdown, TunInbound::shutdown WARN lines quoted above Teardown continues; the new bind may fail and is then logged as a bind failure.
Accept task panics reload, shutdown none from reload JoinError ignored; teardown continues.
Constant Value Defined in Role here
watcher debounce 200 ms app/src/main.rs → spawn_watcher (literal) Settle time between the first event and reload, and so also the period of the watcher’s self-triggered re-reads.
ACCEPT_ERROR_BACKOFF 100 ms app/src/serve.rs Sleep after an accept error; interrupted by cancellation.
MAX_HANDSHAKES_PER_INBOUND 2048 app/src/serve.rs Per accept task, so per generation.
MAX_LIVE_CONNECTIONS_PER_INBOUND 65 536 app/src/serve.rs Per accept task, so per generation.
DRAIN_TIMEOUT 3 s protocols/src/hysteria/server/inbound.rs Bound on Endpoint::wait_idle during teardown.
RELEASE_TIMEOUT 3 s protocols/src/hysteria/server/inbound.rs Bound on waiting for the UDP port to come free.
RELEASE_POLL 20 ms protocols/src/hysteria/server/inbound.rs Interval between bind attempts.
CLOSE_CODE 0x100 protocols/src/hysteria/server/inbound.rs QUIC application error code sent to clients on teardown.
RELEASE_TIMEOUT 3 s protocols/src/tun/inbound.rs Bound on waiting for the device and TCP flows to close.
RELEASE_POLL 20 ms protocols/src/tun/inbound.rs Interval between checks.
DEFAULT_PROBE_INTERVAL 30 s app/src/balancer.rs Probe period when probe_interval is unset.
DEFAULT_PROBE_TIMEOUT 5 s app/src/balancer.rs Probe timeout when probe_timeout is unset.

The worst-case teardown of one generation is therefore about 6 s for a Hysteria 2 inbound (drain plus port release) and about 3 s for a TUN inbound, and it runs concurrently across inbounds.

Test File What it pins
a_reload_rebinds_the_udp_port app/tests/integration/e2e_hysteria_inbound.rs A reload that changes the Hysteria 2 settings on the same UDP port, with a real client connected to the old generation, ends with the inbound serving again on that port. Skips when go is unavailable or the upstream hysteria build fails.
socks_over_a_unix_socket_relays_and_cleans_up app/tests/integration/e2e_unix.rs A Unix listener relays, and after SIGTERM the socket file is gone.
http_connect_over_a_unix_socket_relays app/tests/integration/e2e_unix.rs An HTTP CONNECT over a Unix listener.
a_routed_connect_is_answered_while_the_app_runs app/tests/integration/e2e_tun.rs The TUN interface and its route exist while the app runs and are gone after SIGTERM. Needs CAP_NET_ADMIN, and skips without it.
udp_flows_share_one_association_per_source, a_tcp_connection_becomes_a_stream protocols/tests/pipeline/tun.rs Both end with Fake::stop, which mirrors the app’s teardown: cancel the token, await the run task, then TunInbound::shutdown.
accept_error_backoff_classification app/tests/unit/serve.rs ConnectionAborted and Interrupted retry at once; other accept errors back off.
a_mistyped_key_is_rejected_rather_than_ignored app/tests/unit/config.rs parse_bytes fails on unknown keys, and a misspelt key inside settings fails when the settings are deserialized. That is what makes a typo a failed reload instead of a silent change.
a_unix_listen_takes_no_port, a_stream_protocol_builds_a_unix_stack, tun_owns_its_interface_and_takes_no_listener, tun_parses_addresses_and_routes app/tests/unit/inbound.rs What build_inbound refuses at build time (a Unix listen with a port, a TUN inbound with listen or port) and the BindSpec it produces for Unix and TUN inbounds.

The integration tests run the real binary through spawn_app in app/tests/support/mod.rs, with a relative app.toml in the test directory, so the watcher takes the “no parent component” path and watches .. Proc::terminate sends SIGTERM and waits up to 10 s, which exercises wait_for_shutdown and Instance::shutdown; dropping a Proc sends SIGKILL instead.

The byte dedup, the remembered bytes after a parse or build failure, lenient binding on reload, stale-socket takeover and the watcher’s self-triggered re-reads have no automated test. They were checked against the debug binary for this page. When you change reload or bind_inbound, adding a test for the path you touch is the expected first step.