Generations and hot reload
Source files: 18 · checked against Etemenanki 596916d
Etemenanki/app/src/instance.rsEtemenanki/app/src/main.rsEtemenanki/app/src/serve.rsEtemenanki/app/src/config.rsEtemenanki/app/src/inbound/mod.rsEtemenanki/app/src/inbound/tun.rsEtemenanki/app/src/balancer.rsEtemenanki/protocols/src/hysteria/server/inbound.rsEtemenanki/protocols/src/tun/inbound.rsEtemenanki/protocols/src/tun/device.rsEtemenanki/app/tests/integration/e2e_hysteria_inbound.rsEtemenanki/app/tests/integration/e2e_unix.rsEtemenanki/app/tests/integration/e2e_tun.rsEtemenanki/app/tests/support/mod.rsEtemenanki/app/tests/unit/config.rsEtemenanki/app/tests/unit/inbound.rsEtemenanki/app/tests/unit/serve.rsEtemenanki/protocols/tests/pipeline/tun.rs
A running etemenanki-app holds exactly one generation: the bound listeners, UDP sockets and TUN devices of one configuration, together with every task that serves them, all hanging off a single CancellationToken. Hot reload never patches a generation in place. It builds a complete replacement from the file, cancels the old generation, waits until the old generation has given back its ports, socket paths and interfaces, and then binds the new one.
This page follows that cycle through app/src/instance.rs and app/src/main.rs, down to the locks, tokens, timeouts and log lines involved. Read it before you change how inbounds bind, add an inbound kind that owns a handle, or touch the watcher. What build does with each config section is covered in the build pipeline page; how an accepted socket becomes a relay is covered in the serving page.
Responsibilities
Section titled “Responsibilities”| Component | Owns | Does not do |
|---|---|---|
app/src/main.rs → main |
CLI parsing, tracing setup, --test, the watcher, waiting for signals |
Any config interpretation beyond [log].level |
app/src/main.rs → spawn_watcher |
Watching the config’s directory and turning events into Instance::reload calls |
Deduplication (left to reload) |
app/src/instance.rs → Instance |
The current Config, the bytes it came from, and the live Generation |
Serving connections |
app/src/instance.rs → build |
Validating a Config into a Built: resolver, outbounds, balancers, router, inbounds |
Binding anything |
app/src/instance.rs → spawn_generation |
Binding each inbound and spawning its accept task under a fresh token | Tearing down the previous generation |
app/src/serve.rs, app/src/inbound/tun.rs |
The accept tasks, and releasing their handle before they return | Deciding when a generation ends |
Key types
Section titled “Key types”Instance, State and Generation
Section titled “Instance, State and Generation”pub struct Instance { path: PathBuf, state: tokio::sync::Mutex<State>,}
struct State { config: Config, last_bytes: Vec<u8>, generation: Generation,}
struct Generation { token: CancellationToken, accept_handles: Vec<JoinHandle<()>>,}pathis the-cpath exactly as given (defaultconfig.toml). It never changes for the life of the process; every reload re-reads the same path.stateis atokio::sync::Mutexbecausereloadandshutdownhold it across.awaitpoints: while the old generation drains and while the new one binds. Holding it that long is deliberate. It serialises reloads against each other and against shutdown, so two generations never try to bind at once and a signal that arrives mid-reload waits for the reload to finish.last_bytesis the byte fingerprint of the last file contentreloadacted on, which is not always the content the running generation came from (see The byte fingerprint).Generationis deliberately small: one token that every task of the generation observes, and theJoinHandleof each accept task (one per successfully bound inbound). Connection tasks, sub-stream tasks and balancer probes are not inaccept_handles. They are reached only through the token.
Built and BuiltInbound
Section titled “Built and BuiltInbound”pub struct Built { router: Arc<Router>, inbounds: Vec<BuiltInbound>, balancers: Vec<(Arc<Balancer>, Duration, Duration)>, resolver: Resolver,}
struct BuiltInbound { tag: CompactString, bind: BindSpec, kind: InboundKind,}
pub fn build(cfg: &Config) -> io::Result<Built>build is the whole of validation. It is shared by Instance::start, Instance::reload and the --test dry run (main.rs → test_config), so a config that --test accepts is one that start and reload accept too, up to the point of binding. It reads every file the config references (certificates, keys, the DNS CA file, geodata) and constructs a new Resolver, new outbounds, new balancers and a new Router. Nothing in a Built is shared with the previous generation, so a reload also starts with an empty DNS cache and with every balancer member marked healthy (Member::new initialises healthy to true).
A Built owns the balancers’ probe parameters rather than starting the probes itself, so that the probes can be started against the token of the generation that will own them.
BindSpec and Listener
Section titled “BindSpec and Listener”build_inbound (in app/src/inbound/mod.rs) lowers each inbound into an InboundKind plus a BindSpec that says what the runtime must bind for it. Deciding this at build time is what lets --test refuse an inbound that a start would refuse, for example a Unix listen with a port, or a TUN inbound with a listen.
pub enum BindSpec { Tcp { host: String, port: u16 }, Udp { host: String, port: u16 }, Unix(PathBuf), Tun(DeviceSpec),}
pub enum InboundKind { Stream(Arc<StreamInbound>), Hysteria2(Hy2Inbound<()>), Tun(TunInbound<()>),}enum Listener { Stream(StreamListener), Udp(std::net::UdpSocket), Tun(OwnedFd),}
async fn bind_inbound(ib: &BuiltInbound) -> io::Result<Listener>BindSpec’s Display implementation is what appears in the log lines: 127.0.0.1:1080, udp 0.0.0.0:443, unix:/run/etemenanki/in.sock, tun auto or tun ete0.
spawn_scoped
Section titled “spawn_scoped”pub fn spawn_scoped<F>(token: CancellationToken, fut: F) -> JoinHandle<()>where F: Future + Send + 'static, F::Output: Send,spawn_scoped wraps fut in a tokio::select! against token.cancelled(). When the token fires, the select completes and fut is dropped, together with everything it owns: the client socket, the runtime’s buffers, the outbound stream, the permits. This is the single mechanism that ends every per-connection task of a stream inbound. The stream accept loop spawns each accepted socket with it, and serve_socket spawns each byte stream a transport yields (for example every gRPC stream of one connection) with it too.
The Hysteria 2 and TUN inbounds do not use spawn_scoped for their per-connection work. They keep it in a JoinSet owned by their run future, so dropping that future aborts every connection. Their run future returns when the token fires, or earlier when the endpoint closes or the device goes away.
The accept tasks
Section titled “The accept tasks”pub async fn run_stream_inbound( inbound: Arc<StreamInbound>, tag: CompactString, listener: StreamListener, router: Arc<Router>, token: CancellationToken,)
pub async fn run_hysteria_inbound( inbound: Hy2Inbound<()>, tag: CompactString, socket: std::net::UdpSocket, router: Arc<Router>, token: CancellationToken,)pub async fn run_tun_inbound( inbound: TunInbound<()>, tag: CompactString, fd: OwnedFd, router: Arc<Router>, token: CancellationToken,)Each accept task has the same contract, and reload depends on it: the task returns only after the handle it was given has been released. For a TCP listener that means the listener has been dropped. For a Unix listener the socket file has also been unlinked. For Hysteria 2 the UDP port can be bound again, and for TUN the device descriptor has closed. The release step for each kind:
impl StreamListener { fn release(self)}impl<T> Hy2Inbound<T> { pub async fn shutdown(&self)}impl<T> TunInbound<T> { pub async fn shutdown(&self)}Startup
Section titled “Startup”main initialises tracing first. init_tracing loads the config once, only to read [log].level (defaulting to info, also when the file does not load); RUST_LOG, when it is set to a valid filter, takes precedence. The filter is installed once and never replaced, which is why [log].level is not reloaded.
With --test, main calls test_config, which is config::load followed by instance::build, and exits. It prints Configuration OK. on success, and otherwise logs configuration invalid: … and exits with status 1. It binds nothing, so errors that only binding can reveal (a port in use, a non-socket file at a Unix path, missing TUN privileges) do not show up.
Otherwise main calls Instance::start:
impl Instance { pub async fn start(path: PathBuf) -> io::Result<Self> pub async fn reload(&self) pub async fn shutdown(&self)}start reads the file, runs config::parse_bytes and build, then calls spawn_generation(built, true). Any error is returned, main logs failed to start: …, and the process exits with status 1. On success the read bytes become the first last_bytes. main then wraps the Instance in an Arc, starts the watcher, and waits for a signal.
Spawning a generation
Section titled “Spawning a generation”async fn spawn_generation(built: Built, strict: bool) -> io::Result<Generation>spawn_generation creates a fresh CancellationToken, starts each balancer’s probes under it, then binds the inbounds in config order, one at a time, and spawns an accept task for each one that binds.
flowchart TB
start["new CancellationToken"] --> probes["Balancer::spawn_probe per balancer"]
probes --> next{"next inbound?"}
next -- "no" --> done["Ok(Generation)"]
next -- "yes" --> bind["bind_inbound"]
bind -- "Ok" --> pair["match InboundKind with Listener"]
pair --> spawn["tokio::spawn the accept task, push its JoinHandle"]
spawn --> next
bind -- "Err" --> strict{"strict?"}
strict -- "yes (startup)" --> abort["token.cancel(), return Err"]
strict -- "no (reload)" --> log["log the bind failure"]
log --> next
Balancer probes
Section titled “Balancer probes”pub fn spawn_probe( self: &Arc<Self>, token: CancellationToken, interval: Duration, timeout: Duration, resolver: Resolver,)spawn_probe starts one task per member. Each task probes at once, records the result, and then waits in a select! between token.cancelled() and a sleep(interval), so a probe loop ends at the first wait after its generation is cancelled. The interval and timeout default to DEFAULT_PROBE_INTERVAL (30 s) and DEFAULT_PROBE_TIMEOUT (5 s) when the balancer does not set probe_interval or probe_timeout. Probes are not in accept_handles, and nothing waits for them.
What bind_inbound binds
Section titled “What bind_inbound binds”BindSpec |
Bind call | Listener |
Accept task |
|---|---|---|---|
Tcp |
std::net::TcpListener::bind, then set_nonblocking(true) and tokio::net::TcpListener::from_std |
Stream(StreamListener::Tcp) |
run_stream_inbound |
Unix |
stale-socket check, then std::os::unix::net::UnixListener::bind, set_nonblocking(true) and tokio::net::UnixListener::from_std |
Stream(StreamListener::Unix) (keeps the path) |
run_stream_inbound |
Udp |
std::net::UdpSocket::bind |
Udp |
run_hysteria_inbound, which turns the socket into a quinn endpoint inside Hy2Inbound::run |
Tun |
etemenanki_protocols::tun::open(spec), which creates, addresses and routes the interface, then logs inbound <tag> owns tun device <name> |
Tun(OwnedFd) |
run_tun_inbound |
On success spawn_generation logs inbound <tag> listening on <bind>. build_inbound pairs every kind with its bind, so a mismatched (InboundKind, Listener) pair cannot occur. The match still has a fallback arm that logs inbound <tag>: listener does not match its kind and skips the inbound rather than panicking.
Stale-socket takeover
Section titled “Stale-socket takeover”A Unix listen path can already exist when the inbound binds: the process crashed, a failed start left it behind, or the previous generation’s socket is still there. bind_inbound inspects the path with std::fs::symlink_metadata (so a symlink is judged as itself, not as its target):
| What is at the path | Result |
|---|---|
Nothing (NotFound) |
Bind. |
A socket (FileTypeExt::is_socket) |
Remove it, then bind. The socket is treated as stale and taken over. |
| Anything else | Fail with ErrorKind::AlreadyExists and <path> exists and is not a socket. |
Metadata error other than NotFound |
Fail with that error. |
A regular file at the path therefore stops a start (timestamp removed):
ERROR etemenanki_app: failed to start: inbound u bind unix:/run/etemenanki/in.sock failed: /run/etemenanki/in.sock exists and is not a socketThe check does not ask whether another process is still listening on the socket. A socket file at the configured path is always considered this inbound’s to take.
Strict and lenient binding
Section titled “Strict and lenient binding”Startup (strict = true) |
Reload (strict = false) |
|
|---|---|---|
| First bind failure | Cancels the token, returns Err with inbound <tag> bind <bind> failed: <error> |
Logs inbound <tag> bind <bind> failed: <error> at ERROR |
| Remaining inbounds | Not attempted | Still bound and started |
| Result | failed to start: …, exit status 1 |
Ok(Generation) with fewer accept handles |
Startup is all-or-nothing, so a process that is running is running its whole configuration. Reload is lenient because by the time the new generation binds, the old one is already gone: failing the whole generation would take every inbound down to punish one. With strict = false, spawn_generation cannot return Err, so the Err arm after it in reload (which logs reload: failed to start new generation: …) is defensive.
When a strict start fails partway, the inbounds bound before the failing one have already had their accept tasks spawned. The token cancels them, and the process exits right after. A Unix socket file they created can be left on disk, and the next start takes it over.
The watcher
Section titled “The watcher”fn spawn_watcher( path: PathBuf, instance: Arc<Instance>,) -> notify::Result<notify::RecommendedWatcher>The watcher uses notify 8 (notify::recommended_watcher, which is inotify on Linux) and watches the config file’s parent directory, non-recursively. It watches the directory rather than the file because many editors save by writing a new file and renaming it over the old one. A watch on the old inode would see the rename once and then go quiet. When the path has no parent component (-c config.toml), the directory is ..
The notify callback runs on notify’s own thread. It sends () into an unbounded tokio::sync::mpsc channel for every event that is Ok, and drops error events. It does not look at which file the event names or what kind of event it is: any event in the directory is a reason to try. That includes a change to a certificate or geodata file that lives next to the config, and, because the watch mask that notify 8.2 installs on Linux includes IN_OPEN, a file in the directory merely being opened.
A task spawned on the runtime consumes the channel:
- Wait for one event.
- Drain everything already queued with
try_recv. - Sleep 200 ms (
Duration::from_millis(200), not a named constant). - Drain again, collapsing an editor’s burst of create, write, rename and attribute events into one attempt.
instance.reload().await, then go back to 1.
Because the task awaits reload, attempts never overlap. Events that arrive during a long reload (a Hysteria 2 drain can take seconds) queue up and cause one more attempt afterwards, which the byte check usually turns into a no-op.
main keeps the returned RecommendedWatcher alive in _watcher for the life of the process; dropping it would stop the watch. If creating the watcher or the watch fails, main logs config hot-reload disabled: <error> and runs on without hot reload. That failure is not fatal.
Reload step by step
Section titled “Reload step by step”sequenceDiagram
participant W as watcher task
participant R as Instance::reload
participant F as config file
participant G0 as old generation
participant G1 as new generation
W->>R: reload().await
R->>R: lock state (tokio Mutex)
R->>F: std::fs::read(path)
F-->>R: bytes
alt read error or bytes == last_bytes
R-->>W: return, generation untouched
else changed bytes
R->>R: config::parse_bytes, then build
alt parse or build error
R->>R: last_bytes = bytes
R-->>W: return, old generation keeps serving
else built
R->>R: log "config reload: " + config::diff
R->>G0: token.cancel()
G0->>G0: accept tasks release listener, socket file, UDP port, TUN fd
R->>G0: await every accept JoinHandle
R->>G1: spawn_generation(built, false)
G1-->>R: Generation (failed binds logged, skipped)
R->>R: store generation, config and last_bytes
end
end
In detail:
- Lock.
reloadtakesstate.lock().awaitand holds it to the end. - Read.
std::fs::read(&self.path). On error it logsreload: cannot read <path>: <error>and returns without touchinglast_bytes. A config that is briefly missing during an editor’s save therefore costs nothing: when the file reappears with the same content, the next attempt sees identical bytes. - Byte dedup. If
bytes == state.last_bytes, it returns silently. This absorbs every watcher event that did not change the config, for example events for other files in the directory. - Parse.
config::parse_bytes: UTF-8 check, thentoml::from_strintoConfig. The config structs usedeny_unknown_fields, so a misspelt key fails here rather than falling back to a default. The per-protocolsettingstables are kept as opaque TOML values at this stage, so a misspelt key insidesettingsfails in step 5 instead. On error it logsreload: parse failed, keeping current config: <error>, setslast_bytes = bytes, and returns. - Build.
build(&new_cfg). On error it logsreload: build failed, keeping current config: <error>, setslast_bytes = bytes, and returns. Build reads referenced files, so an unreadable certificate fails here. - Log the diff.
config reload: <diff>atINFO(see The diff line). - Cancel the old generation.
state.generation.token.cancel(). From this point on the running configuration is gone, whatever happens next. - Await the old accept tasks.
for h in accept_handles.drain(..) { let _ = h.await; }. All accept tasks saw the cancellation at the same moment and release concurrently, so the loop waits for the slowest one, not the sum of all of them. A panicked task’sJoinErroris ignored. - Spawn the new generation.
spawn_generation(built, false). - Commit. Store the new
Generation,config = new_cfgandlast_bytes = bytes.
Steps 4 and 5 are the transactional part: nothing is torn down until the new configuration has parsed and built completely. Steps 7 to 9 are not transactional. Binding happens after the old generation has released its handles, because the new generation usually needs the very same ports, paths and interface names.
Outcomes
Section titled “Outcomes”| Case | Log | last_bytes |
Running generation |
|---|---|---|---|
| File unreadable | ERROR reload: cannot read … |
unchanged | old, untouched |
Bytes identical to last_bytes |
none | unchanged | old, untouched |
| Parse error | ERROR reload: parse failed, keeping current config: … |
set to the new bytes | old, untouched |
| Build error | ERROR reload: build failed, keeping current config: … |
set to the new bytes | old, untouched |
| Success | INFO config reload: …, then one listening on or bind … failed line per inbound |
set to the new bytes | new, possibly with some inbounds down |
A failed parse, captured from a run of the debug binary (timestamps and colours removed):
INFO etemenanki_app::instance: inbound a listening on 127.0.0.1:18731ERROR etemenanki_app::instance: reload: parse failed, keeping current config: TOML parse error at line 12, column 12 |12 | garbage = [ | ^unclosed array, expected `]`The byte fingerprint
Section titled “The byte fingerprint”last_bytes records the last content reload made a decision about, and it is updated on a parse or build failure as well as on success. That choice has consequences that are easy to trip over:
- The same broken content is not retried. Touching the file, or saving it again unchanged, is a no-op after a failure. A build failure caused by something outside the file (a certificate that is not there yet) is therefore not retried when the certificate appears; the config bytes must change.
- A failed bind is not retried. After a successful reload in which one inbound failed to bind,
last_bytesholds the new content. Freeing the port does not bring the inbound back until the config bytes change again, for example by editing a comment. - Reverting a broken edit reloads. If the file is restored byte-for-byte to the running configuration after a failed parse, the bytes differ from
last_bytes(the broken content), so a full generation swap follows. The diff line readsno changes, and connections are still dropped. - Files referenced by the config are not watched as content. Replacing a certificate, key, CA file or geodata file triggers a watcher event (when it lives in the same directory) but no reload, because the config bytes are the same. The new file is read on the next reload that the config file itself causes.
The diff line
Section titled “The diff line”pub fn parse_bytes(bytes: &[u8]) -> io::Result<Config>
pub struct ConfigDiff { pub inbounds_added: Vec<CompactString>, pub inbounds_removed: Vec<CompactString>, pub inbounds_changed: Vec<CompactString>, pub outbounds_added: Vec<CompactString>, pub outbounds_removed: Vec<CompactString>, pub outbounds_changed: Vec<CompactString>, pub route_changed: bool, pub log_changed: bool,}
pub fn diff(old: &Config, new: &Config) -> ConfigDiffdiff matches inbounds and outbounds by tag and compares the parsed structs with PartialEq. Its Display form joins groups with ; : inbounds +[b] -[a] ~[c], outbounds ~[proxy], route changed, log changed, or no changes when nothing differs. Renaming a tag shows up as one addition and one removal.
The diff is for the log only. It does not decide what is rebuilt: every successful reload replaces everything. It also does not look at [dns] or [[balancer]], so a reload that changes only those logs config reload: no changes and still swaps the generation. Formatting-only edits (comments, whitespace) log no changes as well.
Tearing down the old generation
Section titled “Tearing down the old generation”Cancelling the token reaches each kind of task differently:
| Task | How it ends | What its accept task waits for before returning |
|---|---|---|
Stream accept loop (run_stream_inbound) |
Its select! takes the token.cancelled() branch and breaks. A loop sleeping ACCEPT_ERROR_BACKOFF after an accept error wakes on the token too (backoff_or_cancelled). |
StreamListener::release: drops the listener and, for a Unix listener, remove_files the path (a failure is logged at DEBUG). |
| Stream connections and transport sub-streams | Each spawn_scoped wrapper drops its future. |
Nothing. They are not awaited, and they finish on their own as the runtime polls them. |
Hysteria 2 (run_hysteria_inbound) |
Hy2Inbound::run breaks out of its accept loop and returns, dropping its JoinSet and with it every connection task. |
Hy2Inbound::shutdown: endpoint.close(CLOSE_CODE, b""), wait_idle bounded by DRAIN_TIMEOUT, drop the endpoint, then try UdpSocket::bind(local) every RELEASE_POLL until it succeeds or RELEASE_TIMEOUT passes. |
TUN (run_tun_inbound) |
TunInbound::run returns, dropping the IP stack and every flow. |
TunInbound::shutdown: poll every RELEASE_POLL until the device’s Weak handle no longer upgrades and the live TCP count is zero, bounded by RELEASE_TIMEOUT. |
| Balancer probes | Their select! takes the cancelled branch at the next wait. |
Not awaited. |
The Hysteria 2 wait exists because dropping a quinn endpoint does not release its UDP socket while connections are live. The socket belongs to the endpoint driver task, which drops it on a later poll. Without the explicit drain and the bind check, the new generation’s UdpSocket::bind would fail with “address in use”, and with lenient binding the inbound would stay down, visible only as one bind-failure line in the log. The shutdown checks the port itself, because a free port is the condition the next generation needs. If the port is still busy after RELEASE_TIMEOUT, it logs hysteria2: <addr> did not come free within 3s at WARN and returns anyway.
The TUN wait exists for the same reason at the interface level. Dropping the IpStack aborts the task that owns the device, but the abort lands only when the runtime schedules it, so the descriptor briefly outlives the drop. Recreating an interface with the same name before that fails. On timeout it logs tun: device fd or <n> tcp flows still open after 3s at WARN.
A TCP listener needs no such wait. Dropping it closes the listening socket at once, and the standard library sets SO_REUSEADDR on Unix, so connections of the old generation lingering in TIME_WAIT do not block the rebind.
Hy2Inbound::shutdown closes the old endpoint with QUIC application error code 0x100 (CLOSE_CODE). The upstream hysteria client reconnects to the new generation on its own, which a_reload_rebinds_the_udp_port relies on.
Generation lifecycle
Section titled “Generation lifecycle”stateDiagram-v2 [*] --> Built: parse_bytes and build succeed Built --> Binding: spawn_generation Binding --> Serving: all inbounds attempted Binding --> [*]: strict bind failure at startup Serving --> Cancelled: reload commits or shutdown Cancelled --> Released: every accept task returned Released --> [*]
A Built that is never spawned (because the process is only running --test) is dropped without having bound anything. Between Cancelled and the next generation’s Binding, no inbound is listening.
Operator-visible behaviour
Section titled “Operator-visible behaviour”- Every successful reload drops every connection, on every inbound, including inbounds whose configuration did not change and edits that only touch a comment. There is no connection hand-over between generations.
- There is a gap. Between cancelling the old generation and binding the new one, nothing listens. For TCP and Unix inbounds the gap is short. With a Hysteria 2 inbound holding live connections or a TUN inbound it can last up to the release timeouts above, and it applies to every inbound, because binding starts only after the slowest accept task has returned.
- A bind failure on reload leaves that inbound down. The other inbounds start, the old listener is already gone, and the inbound stays down until a later reload with different config bytes binds it.
- A broken config keeps the old generation running, with its connections intact, and logs why at
ERROR. [log].levelis not reloaded. The diff still reportslog changed.- The
-cpath is fixed for the life of the process.
The user guide’s hot-reload page describes the same behaviour from the operator’s side.
Shutdown
Section titled “Shutdown”async fn wait_for_shutdown()On Unix, wait_for_shutdown registers SIGTERM with tokio::signal::unix::signal and waits in a select! for either SIGTERM or tokio::signal::ctrl_c() (SIGINT). If SIGTERM cannot be registered, it waits for SIGINT alone. On other platforms it waits for Ctrl-C only.
When a signal arrives, main logs shutting down and calls Instance::shutdown, which takes the state lock (so an in-progress reload finishes first), cancels the current token and awaits every accept handle, exactly like steps 7 and 8 of a reload. main then returns ExitCode::SUCCESS. Dropping the runtime at the end of main drops every task that is still alive.
Going through the accept tasks rather than just exiting matters for the handles that outlive the process otherwise. The Unix socket file is unlinked by StreamListener::release. The TUN device and the routes the inbound installed disappear once the descriptor has closed, which TunInbound::shutdown waits for.
Invariants
Section titled “Invariants”| Invariant | Enforced by | Pinned by |
|---|---|---|
| A config that fails to parse or build never replaces the running generation. | reload returns before token.cancel() on any parse or build error. |
Not covered by an automated test. a_mistyped_key_is_rejected_rather_than_ignored in app/tests/unit/config.rs pins that typos fail at parse time. |
--test, start and reload accept the same configs up to binding. |
All three call the same config::parse_bytes and build. BindSpec is chosen inside build_inbound. |
a_unix_listen_takes_no_port, a_socket_listen_needs_a_port, hysteria2_refuses_a_unix_listen, tun_owns_its_interface_and_takes_no_listener in app/tests/unit/inbound.rs. |
| Every task of a generation ends with it. | spawn_scoped for stream connections and sub-streams, the JoinSets owned by Hy2Inbound::run and TunInbound::run, the token passed to Balancer::spawn_probe. |
a_reload_rebinds_the_udp_port in app/tests/integration/e2e_hysteria_inbound.rs reloads while the old generation holds a live client connection. |
| The old generation has released its ports, socket paths and devices before the new one binds. | reload awaits every accept handle, and each accept task returns only after release or shutdown. Hy2Inbound::shutdown waits on the port itself. |
a_reload_rebinds_the_udp_port (UDP port reused after reload). socks_over_a_unix_socket_relays_and_cleans_up in app/tests/integration/e2e_unix.rs and a_routed_connect_is_answered_while_the_app_runs in app/tests/integration/e2e_tun.rs pin the release on shutdown. |
| The old Unix listener’s cleanup cannot unlink the new generation’s socket. | Ordering: the old accept task’s remove_file has completed before spawn_generation binds the path again. |
Follows from the previous row. No dedicated test. |
Only a socket at a Unix listen path is taken over. |
symlink_metadata and is_socket in bind_inbound; anything else is AlreadyExists. |
Not covered by an automated test. |
| Reloads never overlap each other or a shutdown. | The watcher task awaits each reload, and reload and shutdown both hold the tokio::sync::Mutex<State> throughout. |
Structural. |
| A running process has bound its whole configuration at startup. | spawn_generation(built, true) aborts on the first bind failure. |
Not covered by an automated test. |
| Per-inbound caps start fresh with each generation. | The accept semaphores (MAX_HANDSHAKES_PER_INBOUND, MAX_LIVE_CONNECTIONS_PER_INBOUND), the Hysteria 2 connection semaphore and the TUN flow semaphore are created inside the accept task. |
Structural. |
Failure paths
Section titled “Failure paths”| Failure | Where | Log | Effect |
|---|---|---|---|
| Watcher cannot be created | spawn_watcher |
config hot-reload disabled: … |
Proxy runs without hot reload. |
| Watch error event | notify callback |
none | Event ignored. |
| Config unreadable on reload | reload step 2 |
reload: cannot read … |
Old generation kept, last_bytes unchanged. |
| Parse or build error on reload | reload steps 4 and 5 |
reload: parse failed, … / reload: build failed, … |
Old generation kept, bytes remembered. |
| Bind error on reload | spawn_generation, lenient |
inbound <tag> bind <bind> failed: … |
That inbound down, others started. |
| Bind error at startup | spawn_generation, strict |
failed to start: inbound <tag> bind <bind> failed: … |
Exit status 1. |
| Hysteria 2 endpoint cannot be created from the bound socket | Hy2Inbound::run |
hysteria2 inbound failed: … |
That inbound’s accept task ends; the inbound is down until the next generation. |
| TUN stack cannot start on the device | TunInbound::run |
tun inbound failed: … |
Same as above. |
| Hysteria 2 port or TUN device not released in time | Hy2Inbound::shutdown, TunInbound::shutdown |
WARN lines quoted above |
Teardown continues; the new bind may fail and is then logged as a bind failure. |
| Accept task panics | reload, shutdown |
none from reload |
JoinError ignored; teardown continues. |
Limits
Section titled “Limits”| Constant | Value | Defined in | Role here |
|---|---|---|---|
| watcher debounce | 200 ms | app/src/main.rs → spawn_watcher (literal) |
Settle time between the first event and reload, and so also the period of the watcher’s self-triggered re-reads. |
ACCEPT_ERROR_BACKOFF |
100 ms | app/src/serve.rs |
Sleep after an accept error; interrupted by cancellation. |
MAX_HANDSHAKES_PER_INBOUND |
2048 | app/src/serve.rs |
Per accept task, so per generation. |
MAX_LIVE_CONNECTIONS_PER_INBOUND |
65 536 | app/src/serve.rs |
Per accept task, so per generation. |
DRAIN_TIMEOUT |
3 s | protocols/src/hysteria/server/inbound.rs |
Bound on Endpoint::wait_idle during teardown. |
RELEASE_TIMEOUT |
3 s | protocols/src/hysteria/server/inbound.rs |
Bound on waiting for the UDP port to come free. |
RELEASE_POLL |
20 ms | protocols/src/hysteria/server/inbound.rs |
Interval between bind attempts. |
CLOSE_CODE |
0x100 |
protocols/src/hysteria/server/inbound.rs |
QUIC application error code sent to clients on teardown. |
RELEASE_TIMEOUT |
3 s | protocols/src/tun/inbound.rs |
Bound on waiting for the device and TCP flows to close. |
RELEASE_POLL |
20 ms | protocols/src/tun/inbound.rs |
Interval between checks. |
DEFAULT_PROBE_INTERVAL |
30 s | app/src/balancer.rs |
Probe period when probe_interval is unset. |
DEFAULT_PROBE_TIMEOUT |
5 s | app/src/balancer.rs |
Probe timeout when probe_timeout is unset. |
The worst-case teardown of one generation is therefore about 6 s for a Hysteria 2 inbound (drain plus port release) and about 3 s for a TUN inbound, and it runs concurrently across inbounds.
| Test | File | What it pins |
|---|---|---|
a_reload_rebinds_the_udp_port |
app/tests/integration/e2e_hysteria_inbound.rs |
A reload that changes the Hysteria 2 settings on the same UDP port, with a real client connected to the old generation, ends with the inbound serving again on that port. Skips when go is unavailable or the upstream hysteria build fails. |
socks_over_a_unix_socket_relays_and_cleans_up |
app/tests/integration/e2e_unix.rs |
A Unix listener relays, and after SIGTERM the socket file is gone. |
http_connect_over_a_unix_socket_relays |
app/tests/integration/e2e_unix.rs |
An HTTP CONNECT over a Unix listener. |
a_routed_connect_is_answered_while_the_app_runs |
app/tests/integration/e2e_tun.rs |
The TUN interface and its route exist while the app runs and are gone after SIGTERM. Needs CAP_NET_ADMIN, and skips without it. |
udp_flows_share_one_association_per_source, a_tcp_connection_becomes_a_stream |
protocols/tests/pipeline/tun.rs |
Both end with Fake::stop, which mirrors the app’s teardown: cancel the token, await the run task, then TunInbound::shutdown. |
accept_error_backoff_classification |
app/tests/unit/serve.rs |
ConnectionAborted and Interrupted retry at once; other accept errors back off. |
a_mistyped_key_is_rejected_rather_than_ignored |
app/tests/unit/config.rs |
parse_bytes fails on unknown keys, and a misspelt key inside settings fails when the settings are deserialized. That is what makes a typo a failed reload instead of a silent change. |
a_unix_listen_takes_no_port, a_stream_protocol_builds_a_unix_stack, tun_owns_its_interface_and_takes_no_listener, tun_parses_addresses_and_routes |
app/tests/unit/inbound.rs |
What build_inbound refuses at build time (a Unix listen with a port, a TUN inbound with listen or port) and the BindSpec it produces for Unix and TUN inbounds. |
The integration tests run the real binary through spawn_app in app/tests/support/mod.rs, with a relative app.toml in the test directory, so the watcher takes the “no parent component” path and watches .. Proc::terminate sends SIGTERM and waits up to 10 s, which exercises wait_for_shutdown and Instance::shutdown; dropping a Proc sends SIGKILL instead.
The byte dedup, the remembered bytes after a parse or build failure, lenient binding on reload, stale-socket takeover and the watcher’s self-triggered re-reads have no automated test. They were checked against the debug binary for this page. When you change reload or bind_inbound, adding a test for the path you touch is the expected first step.