Process runtime and reload
Source files: 14 · checked against katana v3.0.1
katana/src/main.rskatana/src/runtime.rskatana/src/config.rskatana/src/manager/mod.rskatana/src/manager/node.rskatana/src/api/mod.rskatana/src/api/newv2board.rskatana/src/api/sspanel.rskatana/tests/unit/runtime.rskatana/tests/unit/e2e.rskatana/tests/support/mod.rskatana/tests/integration/xray_interop.rskatana/tests/integration/sniff.rskatana/tests/integration/hysteria_interop.rs
src/runtime.rs is the root of a katana process. It loads the config file once, builds the global outbound pool, spawns one NodeManager task per [[node]], and then sits in a loop that waits for either a shutdown signal or a change in the config directory. When the file changes, it compares the new config with the running one, builds everything the difference needs, and refuses the whole reload if any of it fails. Otherwise it fans the difference out in a fixed order: a new log filter, a new outbound pool for every node, removed nodes, and then one pass over the new [[node]] list in file order that reconfigures kept nodes and spawns added ones.
This page is for contributors who change that loop, the node identity rule or the StaticUpdate channel. It describes the runtime from the process’s point of view. What a node does with an update once it receives one is covered in depth on Node manager; the operator’s view of the same machinery is Hot reload.
Responsibilities
Section titled “Responsibilities”| Concern | Owner | Left to |
|---|---|---|
Command line, --test, exit code |
src/main.rs → main, src/runtime.rs → test_config |
clap for usage errors |
| Log subscriber and live level changes | src/runtime.rs → init_tracing, LogReload |
tracing_subscriber::reload |
| The outbound pool and its shared DNS resolver | src/runtime.rs → build_outbounds |
src/outbound → build_outbound for each entry |
| One task per node, with its channel and cancellation token | src/runtime.rs → build_node, spawn_built, spawn_node, NodeHandle |
NodeManager::run for everything inside the node |
| Config file watching and debounce | src/runtime.rs → setup_watcher and the loop in run |
the notify crate |
| Diffing a reload and dispatching it | src/runtime.rs → apply_reload, identity |
NodeManager::apply_static |
| SIGINT and SIGTERM | src/runtime.rs → wait_for_shutdown |
each node’s own teardown and final report |
The runtime never touches a listener, a user table or a traffic counter. Everything below the node boundary is reached only through two channels per node: a CancellationToken to stop it and an mpsc::Sender<StaticUpdate> to steer it.
Unlike the standalone app, katana has no generation swap of the whole process. Instead, a reload is built and validated before anything is applied: the new outbound pool, every node the reload adds, and the panel client of every node whose table changed. If any of them fails, the running process is left exactly as it was. After that, each node rebuilds its own listener in place when an update asks for it. The app’s approach is described on Generations and reload.
Entry point
Section titled “Entry point”src/main.rs parses two flags and hands off to the runtime:
#[derive(Parser)]#[command(version, about)]struct Args { #[arg(short, long, default_value = "config.toml")] config: PathBuf, #[arg(long)] test: bool,}
#[tokio::main]async fn main() -> ExitCodemain always calls runtime::init_tracing(&args.config) first, so the dry run logs through the same subscriber as a real run. With --test it calls runtime::test_config, prints Configuration OK to stdout on success or configuration error: <error> to stderr on failure, and returns. Otherwise it awaits runtime::run(args.config, reload). The runtime is the default multi-threaded Tokio runtime from #[tokio::main].
| Exit code | When |
|---|---|
0 |
--test passed, run returned after a shutdown signal, or clap printed --help or --version |
1 |
--test failed; or at startup the config did not load, the outbound pool did not build, the file has no [[node]], or no node could be spawned |
2 |
A usage error reported by clap (its default) |
A node that is spawned but cannot come up, for example because its panel is unreachable, does not change the exit code: it keeps retrying in its own task; see Node task lifecycle.
Key types
Section titled “Key types”Type aliases
Section titled “Type aliases”pub type LogReload = Box<dyn Fn(&str) + Send + Sync>;
type Pool = Arc<HashMap<CompactString, Arc<Outbound>>>;
type NodeId = (String, String, u32, String, String);LogReloadis the setterinit_tracingreturns. The runtime calls it with a level string when[log].levelchanges.Poolis the global outbound pool, keyed by tag. It is anArcso that one reload can hand the same map to every node without copying it; each node keeps its own clone of theArcand compiles its router against it.NodeIdis the identity a[[node]]is matched by across reloads; see Node identity.
Identity helpers
Section titled “Identity helpers”fn identity(cfg: &NodeConfig) -> NodeId
fn display_id(id: &NodeId) -> Stringidentity builds the tuple (panel_type lowercased, api.host, api.node_id, api.key, api::panel_node_type(cfg)). The fifth element is the node type the panel finds the node by, described under Node identity; it is an empty string for sspanel.
display_id renders an identity for logs. When the node type is non-empty it prints panel_type@host#node_id/node_type, for example newv2board@https://panel.example.com#1/v2ray; for sspanel it prints panel_type@host#node_id, for example sspanel@https://panel.example.com#1. It deliberately drops the fourth element: the identity carries api.key, and display_id is the only way the runtime prints an identity, so the key never reaches a reload log line.
NodeHandle
Section titled “NodeHandle”struct NodeHandle { id: NodeId, cfg: NodeConfig, task: JoinHandle<()>, static_tx: mpsc::Sender<StaticUpdate>, shutdown: CancellationToken,}One handle per spawned node, held in a Vec<NodeHandle> owned by run. It is the runtime’s whole view of a node:
| Field | Use |
|---|---|
id |
Matching against the next config’s identities |
cfg |
The last [[node]] table sent to the node, compared with != to decide whether a StaticUpdate::Config is needed |
task |
Awaited when the node is removed and at shutdown, so the node’s final traffic report finishes before the runtime moves on |
static_tx |
The sending half of the node’s update channel, capacity 16 |
shutdown |
A child of the root token; cancelling it stops this node only |
StaticUpdate
Section titled “StaticUpdate”The update type lives in src/manager/mod.rs:
pub enum StaticUpdate { Config(Box<NodeConfig>), Outbounds(Arc<HashMap<CompactString, Arc<Outbound>>>),}Config carries a changed [[node]] table with the same identity. Outbounds carries a rebuilt pool. The runtime resolves identity changes itself (remove plus add), so a node never receives a Config for another panel node. A Config can still change [node.api] fields the node’s panel client was built from, such as api.timeout; the node then builds a new client itself (see How a node applies a static update).
The node side of the contract
Section titled “The node side of the contract”impl NodeManager { pub fn new( api: PanelClient, cfg: NodeConfig, pool: Arc<HashMap<CompactString, Arc<Outbound>>>, ) -> io::Result<Arc<Self>>
pub async fn run( self: Arc<Self>, shutdown: CancellationToken, mut static_rx: mpsc::Receiver<StaticUpdate>, )
async fn bootstrap( &self, shutdown: &CancellationToken, static_rx: &mut mpsc::Receiver<StaticUpdate>, ) -> bool
async fn serve( &self, shutdown: &CancellationToken, static_rx: &mut mpsc::Receiver<StaticUpdate>, )
async fn apply_static(&self, u: StaticUpdate)}NodeManager::new compiles the node’s router against the pool, with direct as the default tag, and fails on a route that names an unknown outbound tag. It binds nothing. The manager holds its panel client as Mutex<Arc<PanelClient>>, so that apply_static can swap in a new one. run is the node task’s body: bootstrap until the node is up, then serve, then teardown and the final reports. apply_static is what both bootstrap and serve call for each message from static_rx.
Other runtime functions
Section titled “Other runtime functions”pub async fn run(config_path: PathBuf, reload: LogReload) -> ExitCode
fn spawn_node(cfg: &NodeConfig, pool: &Pool, root: &CancellationToken) -> Option<NodeHandle>
fn build_node(cfg: &NodeConfig, pool: &Pool) -> anyhow::Result<Arc<NodeManager>>
fn spawn_built(nm: Arc<NodeManager>, cfg: &NodeConfig, root: &CancellationToken) -> NodeHandle
async fn apply_reload( cfg: &mut Config, pool: &mut Pool, handles: &mut Vec<NodeHandle>, new: Config, reload: &LogReload, root: &CancellationToken,)
fn setup_watcher( config_path: &Path, ev_tx: mpsc::UnboundedSender<()>,) -> notify::Result<RecommendedWatcher>
pub fn build_outbounds(cfg: &Config) -> io::Result<HashMap<CompactString, Arc<Outbound>>>
pub fn test_config(path: &Path) -> anyhow::Result<()>
pub fn init_tracing(config_path: &Path) -> LogReload
async fn wait_for_shutdown()The panel client the runtime builds for a node comes from src/api/mod.rs:
impl PanelClient { pub fn new(cfg: &NodeConfig) -> anyhow::Result<Self>}
pub fn panel_node_type(cfg: &NodeConfig) -> StringStartup
Section titled “Startup”run owns three pieces of mutable state for the life of the process: cfg: Config (the last applied config), pool: Pool and handles: Vec<NodeHandle>. Everything else is local to one step.
-
Load.
config::load(&config_path)reads the file and parses it withtomlintoConfig. Every config struct is#[serde(deny_unknown_fields)], so a misspelt key fails here. On errorrunlogsfailed to load config: …and returnsExitCode::FAILURE. -
Build the pool.
build_outbounds(&cfg)builds the resolver and every outbound. On errorrunlogsfailed to build outbounds: …and fails. -
Require nodes. An empty
cfg.nodeslogsconfig defines no [[node]] entriesand fails. The check comes after the pool build, so a bad[[outbound]]is reported first. -
Spawn nodes.
runcreates the rootCancellationTokenand callsspawn_nodefor each[[node]]in file order. A node that cannot be built is logged and skipped. If no handle results,runlogsno nodes could be startedand fails. -
Watch the file. An unbounded
mpsc::unbounded_channel::<()>()is created and its sender handed tosetup_watcher. If the watcher cannot be set up,runlogsconfig watcher disabled (no live reload): …and continues without it. The returnedRecommendedWatcheris held in a local_watcheruntilrunreturns; dropping it would stop the watch. -
Arm signals. A detached task awaits
wait_for_shutdown()and then cancels the root token. -
Loop.
runenters the select loop described under The select loop.
The outbound pool
Section titled “The outbound pool”build_outbounds produces one pool for the whole process. It first builds a single Resolver from [dns] through Resolver::from_spec(ResolverSpec { … }), reading dns.ca_file from disk if it is set. Every outbound in the pool shares that resolver, and so its cache, because the outbounds resolve an overlapping set of names.
It then seeds four reserved tags before it looks at the file:
| Tag | Handler |
|---|---|
direct |
Outbound::Direct(FreedomConnector::new(resolver.clone(), AddressFamilyStrategy::Auto)) |
block |
Outbound::Block |
freedom |
A second, separate Outbound::Direct built the same way as direct (Xray alias) |
blackhole |
Outbound::Block (Xray alias) |
Each [[outbound]] is then built with build_outbound(entry, &resolver) and inserted under its tag. A tag that is already present, whether reserved or used by an earlier entry, fails the whole pool with duplicate/reserved outbound tag <tag>. The check is an exact HashMap lookup, so it is case-sensitive: Direct is accepted as an ordinary tag.
Protocol-specific construction is described on Inbound and outbound, and the resolver on DNS.
Spawning a node
Section titled “Spawning a node”Spawning a node has two halves. build_node(cfg, pool) does everything that can fail and binds nothing:
PanelClient::new(cfg)picks the client bypanel_type, compared case-insensitively:sspanelbuildsPanelClient::Sspanel,newv2boardorv2boardbuildsPanelClient::NewV2board, anything else fails withunknown panel_type "<value>", where the value is shown lowercased. Both client constructors also reject an unknownapi.node_typewithunknown node_type "<value>"and build theirreqwest::Clientwith the panel timeout.NodeManager::new(api, cfg.clone(), pool.clone())compiles the router. A failure becomesbuild router: <error>.
spawn_built(nm, cfg, root) cannot fail:
mpsc::channel(16)creates the update channel, androot.child_token()creates the node’s own token.tokio::spawn(nm.run(shutdown.clone(), static_rx))starts the task, and the handle is returned withidentity(cfg)and a clone of the config.
At startup, spawn_node calls build_node and then spawn_built. When build_node fails, it logs node <node_id>: <error> (for a router, node <node_id>: build router: <error>) and returns None, so one bad [[node]] never aborts the process.
A reload does not use spawn_node. It calls build_node for every node it adds during its first stage, against the new pool if one was built and the running pool otherwise, and spawns those pre-built managers with spawn_built only once the whole reload has been accepted; see apply_reload.
Node task lifecycle
Section titled “Node task lifecycle”From the runtime’s side, a node task goes through the states below. The runtime observes none of them directly: it only holds the JoinHandle, and it looks at the handle only to await it.
stateDiagram-v2 [*] --> Bootstrap: tokio spawn Bootstrap --> Serving: node up Bootstrap --> Waiting: attempt failed Waiting --> Bootstrap: backoff elapsed or StaticUpdate Config Serving --> Serving: poll tick or StaticUpdate Bootstrap --> Draining: token cancelled Waiting --> Draining: token cancelled Serving --> Draining: token cancelled Draining --> Exited: tear down, report traffic, report illegal Exited --> [*]
-
Bootstrap.
NodeManager::bootstrapruns one attempt,try_bootstrap, inside a biasedtokio::select!against the token, so a cancelled node stops at once, even in the middle of a panel request. Each attempt first callsforget_etagson the panel client, so the panel answers in full instead of with304 Not Modifiedfor a request an earlier attempt already made. The attempt then asks for node info (from the panel, or, for a Hysteria 2 node whose[node.hysteria].portis set, from the config), then the user list, then callsbring_up, then fetches the audit rules unlesscontroller.disable_get_ruleis set. Each panel request is bounded by the panel timeout (api.timeout,5seconds when unset, fromApiConfig::timeout_secs). An attempt fails with one of these reasons:Reason Cause panel returned no node infoThe node info request was answered 304 Not Modifiednode_info failed: <error>The node info request failed panel returned port 0The node info has port 0. For a panel port of0only an SSPanel node gets this line: the newV2board client rejects aserver_portof0itself, so there the attempt fails withnode_info failed: newV2board: server port must be > 0. A newV2boardserver_portthat truncates to0as au16, such as65536, passes that check and reaches this onepanel returned no user listThe user list request was answered 304 Not Modifieduser_list failed: <error>The user list request failed initial start failed: <error>bring_upfailed, for example because the port is still held by another processThe user list is required: a node does not come up without one. An empty list is not a failure, though; the node comes up with no users and binds nothing, as the reconcile ladder does for a user-less node. A failed rule fetch does not fail the attempt.
-
Waiting. After a failed attempt, the node logs
node <node_id>: <reason>; retrying in <n>sand waits. The first wait isBOOTSTRAP_RETRY_MIN(1s); the wait doubles after each attempt, up toBOOTSTRAP_RETRY_MAX(60s), and every wait is capped at the poll period, so a node that polls every few seconds also retries that often. One timer covers the whole wait, so updates that arrive meanwhile do not push the retry back. While it waits, the node applies everyStaticUpdatefrom its channel; with nothing up, an update only stores what it carries. AStaticUpdate::Configends the wait and starts the next attempt at once, because the edit may be the fix. The next wait still doubles. -
Serving. A
tokio::select!over the token, the poll interval andstatic_rx. After eachStaticUpdatethe node rereadscontroller.update_periodicand, if it changed, replaces the interval and consumes its immediate first tick, so the next poll comes one new period later. -
Draining. Once the token fires, in any state,
runcallstear_down(which awaitsTransportManager::shutdown), thenreport_trafficandreport_illegal. Every relay has stopped by then, so the counters it reports are final. A node stopped before it came up normally has no listener and no traffic, but an attempt cut short after its bind is still taken down this way.
A node task never ends on its own. It ends only when its token is cancelled, by a reload that removes the node or by shutdown. A node whose panel is unreachable stays between Bootstrap and Waiting and logs one line per attempt; the rest of the process is not affected. The node’s side of this is described on Node manager.
The config watcher
Section titled “The config watcher”setup_watcher watches the config file’s parent directory, not the file. Editors commonly save by writing a temporary file and renaming it over the original, which replaces the inode a file watch would be attached to; a directory watch sees both kinds of save. When the path has no parent component (a bare config.toml), it watches .. The watch is RecursiveMode::NonRecursive.
notify delivers events on a std::sync::mpsc channel from its own thread, so the runtime bridges them into Tokio with a plain OS thread:
flowchart LR N["notify RecommendedWatcher"] -- "notify Result of Event" --> S["std mpsc channel"] S --> T["bridge thread"] T -- "unit tick" --> U["tokio unbounded mpsc"] U --> L["select loop in run"]
- The bridge forwards one
()per message, whatever the message is. It does not look at the event kind or the path, and a watcher error is forwarded as a tick like any event. Any event on an entry of that directory therefore starts a reload attempt: a certificate being rewritten, another file being created, and on Linux even a file being opened, becausenotify’s inotify backend (notify 8.2) subscribes to open events. When the config file itself did not change, that attempt parses the same file, finds no difference and does nothing: a certificate rewritten at the same path is not picked up by the reload. - The thread ends when the watcher is dropped (its
recvfails) or when the Tokio receiver is gone (itssendfails). Both happen whenrunreturns. - The Tokio channel is unbounded, but it only carries
()and the loop drains it on every reload.
Debounce
Section titled “Debounce”When the loop receives a tick, it sleeps a fixed 500 ms (Duration::from_millis(500)), then empties the channel with try_recv until it is empty, then reloads. This is a fixed window from the first event, not a sliding one: a burst from one save becomes one reload.
Ticks that arrive while config::load or apply_reload is running are not drained; they wait in the channel and start another reload as soon as the current one finishes. A save made during a slow reload is therefore never lost, at the cost of one extra reload.
On Linux this rule makes the loop feed itself. config::load opens the config file inside the watched directory after the drain, the open produces an event, and that tick starts the next reload. The startup loads happen before the watcher exists, so a freshly started process is quiet, but after the first event in the directory katana re-reads its config file about every 500 ms for the rest of its life. Each of these passes parses the file, finds nothing changed and applies nothing, so the cost is one file read and parse per cycle. It also means that a reload error is logged again on every cycle, about twice a second, until the file is corrected: a rejected reload never replaces the running cfg, so the next pass finds the same difference and fails the same way. This holds for config reload failed, keeping current: … for a file that does not parse, reload: bad outbounds, keeping current config: … for a pool that does not build, and reload: node <display_id>: <error>; keeping current config for a node that apply_reload refuses. Under strace, the config file is opened twice at startup (by init_tracing and by run), not again until a file in the directory is written, and from then on about every 500 ms.
The select loop
Section titled “The select loop”loop { tokio::select! { _ = root.cancelled() => break, Some(()) = ev_rx.recv() => { /* debounce, load, apply_reload */ } }}- If
config::loadfails, the loop logsconfig reload failed, keeping current: …and keeps everything. Because the next reload compares against the running config, not against the rejected one, any later event in the directory retries the same file. - If the watcher was never set up, its sender was dropped inside
setup_watcher,ev_rx.recv()returnsNone, the pattern does not match, and the branch is disabled; the loop then waits on the token alone. - The debounce sleep and
apply_reloadrun inside the branch body, not as select arms. A signal that arrives during a reload cancels the root token (and so every node’s token) immediately, but the loop only observes it once the reload returns.
apply_reload
Section titled “apply_reload”apply_reload takes the running cfg, pool and handles by mutable reference and applies a freshly parsed Config over them. Its order is fixed:
sequenceDiagram
participant L as run loop
participant R as apply_reload
participant B as build_outbounds
participant K as kept node
participant X as removed node
participant N as new node
L->>R: apply_reload(new)
opt outbounds differ
R->>B: build_outbounds(new)
B-->>R: new pool or error
Note over R: error means log and return, nothing applied
end
loop each new node table, in file order
R->>R: build_node for an added node, PanelClient new for a changed kept node
Note over R: error means log and return, nothing applied
end
opt log level differs
R->>R: reload(level)
end
opt new pool built
R->>K: StaticUpdate Outbounds
R->>X: StaticUpdate Outbounds
end
R->>X: shutdown.cancel()
X-->>R: task ends after final report
loop each new node table, in file order
alt identity matches a handle and the table differs
R->>K: StaticUpdate Config
else identity matches no handle
R->>N: spawn_built with the pre-built manager
end
end
R-->>L: cfg replaced by new
-
Validate the pool. Only if
new.outbounds != cfg.outboundsdoesapply_reloadcallbuild_outbounds(&new). A failure logsreload: bad outbounds, keeping current config: …and returns before any side effect, so the log level, the pool, the nodes andcfgall stay as they were. -
Build the nodes. The runtime walks the new
[[node]]list once, in file order, and checks each table against the running handles:- Kept node, table unchanged. Nothing to build.
- Kept node, table changed.
PanelClient::new(node_cfg)is built and dropped, so an edit the node could not build a client for, such as an unknownapi.node_typeon sspanel, is refused here. The node’s[node.route]is not compiled; see step 6. - Added node.
build_node(node_cfg, pool)builds the panel client and the manager, against the new pool if step 1 built one and the running pool otherwise. The manager is kept for step 6. Nothing is bound or spawned. - A second table with an identity this pass already built. Nothing to build; step 6 treats it as a reconfiguration of the first.
The first failure logs
reload: node <display_id>: <error>; keeping current configand returns. Nothing is applied: not the log level, not the pool, no removal, andcfgis not committed. Managers already built are dropped; none of them had bound anything. Because an identity change is a remove plus an add, this check is what keeps an edit whose replacement does not build from removing the running node. -
Swap the log level. If
new.log.level != cfg.log.level, the runtime callsreload(level), where a removed key becomes"info", and logsreload: log level → <level>. See Logging. -
Broadcast the pool. If a new pool was built,
*poolis replaced andStaticUpdate::Outbounds(pool.clone())is sent to every handle, including nodes that step 5 is about to remove. Such a node may rebuild its listener once before it sees its token cancelled. The runtime logsreload: outbound pool rebuilt. -
Remove nodes. The identities of the new config are collected into a
HashSet<NodeId>.handlesis drained; each handle whose identity is missing logsreload: removing node <display_id>, has its token cancelled and itsJoinHandleawaited. Removals run one after another, and each waits for that node’s final traffic and audit reports. -
Reconfigure and add, in one pass. The runtime walks the new
[[node]]list again, in file order, and handles each table as it reaches it, so reconfigurations and additions are interleaved:- Kept node. If the identity matches a handle, the runtime compares the handle’s
cfgwith the new table. If they differ, it stores the new table in the handle, sendsStaticUpdate::Config(Box::new(node_cfg.clone()))and logsreload: reconfigured node <display_id>. - Added node. If the identity matches no handle, the manager built for it in step 2 is spawned with
spawn_built, which cannot fail, and the runtime logsreload: added node <display_id>.
The runtime does not compile a kept node’s route. A
[node.route]edit that names an unknown outbound tag passes step 2 and is sent; the node refuses the edit itself (see How a node applies a static update). The handle’scfgalready holds the refused table by then, so a later reload that leaves the table as it is sends nothing: the refused edit is not sent again until the table changes again. - Kept node. If the identity matches a handle, the runtime compares the handle’s
-
Commit.
*cfg = new. Only a reload that passed steps 1 and 2 gets here.
What each step compares
Section titled “What each step compares”| Part of the config | Compared by | Sent as |
|---|---|---|
[[outbound]] list |
Vec<OutboundConfig> != |
StaticUpdate::Outbounds to every node |
[log].level |
Option<String> != |
A call to LogReload |
[dns] |
Not compared | Nothing |
[[node]] identity |
NodeId equality |
Remove and spawn; the added node is built in step 2 |
The rest of a [[node]] |
NodeConfig != against NodeHandle.cfg |
StaticUpdate::Config to that node; its panel client is built in step 2 |
The outbound list is compared as a whole Vec, so reordering [[outbound]] tables without changing any of them still counts as a change: it rebuilds the pool and every node’s listener.
[dns] is read only inside build_outbounds, which a reload calls only when the outbound list changed. An edit to [dns] alone is therefore not applied until the next [[outbound]] edit or a restart.
Ordering consequences
Section titled “Ordering consequences”- Removals come before additions. When an identity change turns one node into a remove and an add, the old node’s listener is shut down and its task awaited before the replacement is spawned, so the replacement can bind the same port. The replacement’s manager was built in step 2, but a built manager binds nothing until it is spawned. The same holds for a port that moves from a removed node to an added one in a single save.
- Updates are enqueued, not awaited.
static_tx.send(…).awaitreturns once the message is in the channel. The runtime does not wait for the node to apply it, so a node added in step 6 can bootstrap concurrently with a kept node that is still rebuilding for an update sent earlier in the same pass. Only removal is synchronous. - Backpressure on the channel.
sendwaits for capacity. A node takes messages from its channel while it serves and while it waits between bootstrap attempts, so the channel holds messages only while the node is busy with one piece of work, such as a bootstrap attempt in flight; the reload waits only if16updates pile up during it. Node tasks do not end on their own, so a handle’s receiver stays open until the runtime cancels the node. If asendfails anyway, the error is ignored. - Two updates can mean two rebuilds. When one save changes both the outbound list and a node’s table, that node receives
Outboundsfirst andConfigsecond, and each can rebuild its listener. A rename of an outbound together with the rules that use it relies on this order: the node first fails to compile its old rules against the new pool (and keeps its old router), then compiles the new rules whenConfigarrives. - An empty node list is accepted. Unlike startup,
apply_reloaddoes not require any[[node]]. A file with none removes every node and the process keeps running with no nodes.
Node identity
Section titled “Node identity”A [[node]] table in the new config is the same node as a running one exactly when their NodeId tuples are equal:
| Position | Source | Compared as |
|---|---|---|
| 0 | panel_type |
to_ascii_lowercase(), so NewV2board equals newv2board |
| 1 | api.host |
Exact string: a trailing / makes a different identity |
| 2 | api.node_id |
u32 |
| 3 | api.key |
Exact string |
| 4 | Panel node type, from api::panel_node_type(cfg) |
Exact string. For newv2board and v2board, newv2board::node_type_param: vless for a V2ray-family node (api.node_type V2ray, Vmess or Vless, any case) with api.enable_vless, otherwise api.node_type lowercased. Empty for sspanel |
The identity says which panel node a [[node]] table serves: which panel, which node on it, and with which credentials. The node type is part of that on newV2board, because UniProxy finds a node by its id and by the node type it is asked for: the same node_id asked for as vless and as vmess is two panel nodes, each with its own users and its own traffic. sspanel’s mod_mu finds a node by its id alone, so the type is left out there.
Changing any identity field removes the node and spawns a fresh one. On newV2board that includes a changed api.node_type (other than a change of case only) and, on a V2ray or Vmess node, a changed api.enable_vless; a Vless node is asked for as vless either way, and the flag means nothing to the other node types. Every other field, including api.timeout and, on sspanel, api.node_type and api.enable_vless, reaches the running node as a StaticUpdate::Config. The consequences:
- The removed node sends its final report with the old panel connection, and the new node starts with an empty
NodeTraffic, so no counter is carried from one panel node to another. - Order in the file does not matter: matching is by identity, not by position.
- Two
[[node]]tables with the same identity are not rejected. At startup both are spawned; at reload both match the first handle. The config is ambiguous and the behaviour follows fromfind.
The panel clients copy [node.api] fields when they are constructed: newv2board::Client::new keeps api.node_type, api.enable_vless, api.speed_limit, api.rule_list_path and the timeout; sspanel::Client::new also keeps api.vless_flow and api.disable_custom_config. The node does not keep a client past an edit to those fields: when panel_type or anything in [node.api] changes, apply_static builds a new PanelClient from the new table and swaps it in. The identity therefore only has to cover which panel node is meant, not every field the client was built from.
How a node applies a static update
Section titled “How a node applies a static update”NodeManager::apply_static in src/manager/node.rs is the receiving end. It runs in the node’s own task, between poll cycles or between bootstrap attempts, so a static update and a panel poll never interleave.
StaticUpdate::Config(new_cfg) is taken whole or not at all. apply_static builds what the edit needs before it stores anything:
- Client. If
panel_typeor anything inapidiffers from the running config,PanelClient::new(&new_cfg)builds a new client, andinherit_routescopies the routes the old newV2board client last read into it, so the audit rules derived from them keep applying until the new client reads the node config itself. sspanel asks for its rules on every refresh, so it has nothing to inherit. - Router. If
[node.route]differs,build_routercompiles the new route against the node’s current pool, withdirectas the default tag. - Refuse or store. If either build fails, the node logs
node <node_id>: config edit refused, keeping the running one: <error>and returns. Nothing is stored: the node keeps its config, client, router and listener. Otherwise it stores the new config, then the new router (loggingnode <node_id>: router rebuilt), then the new client. - Resync. If the node is bootstrapped and the edit needs a listener rebuild or brought a new client, the node runs one
poll_cycle(rebuild)at once, which also refreshes the audit rules and reports traffic. A node that is still bootstrapping has nothing to apply the edit to: its next attempt, which aConfigstarts at once, reads with the new client and builds from the new config.
| Changed | Built first | Then |
|---|---|---|
[node.route], controller.listen_ip, anything in controller.cert, controller.disable_sniffing, api.enable_vless, anything in [node.hysteria] |
The router, if the route changed; a client, if an api field changed |
poll_cycle(true): the reconcile ladder takes its transport-rebuild branch, a full listener rebuild against the panel’s answer |
A case-only change of panel_type, or any other [node.api] field, for example api.speed_limit, api.rule_list_path or api.timeout |
A client | poll_cycle(false): the new client holds no ETags, so node info and users are read in full, and the reconcile ladder drops connections only if that answer requires it |
Anything else, such as controller.update_periodic or controller.disable_get_rule |
Nothing | Stored; read where it is used |
controller.cert is compared as config values (mode, cert_file, key_file, reject_unknown_sni), not by file contents: rewriting the certificate at the same path is not a change.
A listener rebuild ends every connection on the node, including connections still in their handshake. Only a transport rebuild guarantees that, which is why a route edit forces one instead of retiring user scopes. NodeTraffic outlives the rebuild, so users present before and after keep their counters. If the panel’s answer has no users, the ladder tears the listener down instead of rebuilding it.
StaticUpdate::Outbounds(pool) replaces the node’s pool, then rebuild_router recompiles the router and apply_route_change does a full listener rebuild, whether or not the node’s rules use the outbound that changed. Every connection on every node drops on any [[outbound]] edit. This is the only path that still uses rebuild_router and apply_route_change.
If rebuild_router fails (for example, a rule names a tag the new pool no longer has), it logs node <id>: route rebuild failed, keeping current: … and keeps the previous router, and the listener is still rebuilt. The runtime compiles routes against a new pool only for the nodes a reload adds, in step 2; katana --test checks every node.
Logging
Section titled “Logging”init_tracing installs the global subscriber once, before --test or run:
pub fn init_tracing(config_path: &Path) -> LogReload- The initial filter is
EnvFilter::try_from_default_env(), which readsRUST_LOG. IfRUST_LOGis unset or does not parse,init_tracingloads the config file itself and uses[log].level, or"info"if the key is absent or the file does not load. That value goes throughEnvFilter::new, which drops invalid directives with a message on stderr instead of failing. - The filter is wrapped in
tracing_subscriber::reload::Layer::new, and the registry gets that layer plusfmt::layer(), which writes to stdout. - The returned
LogReloadclosure parses the new level withEnvFilter::try_new. On success it callshandle.reload(f), ignoring the result; on failure it logsinvalid log level "<level>": <error>and keeps the current filter.
A level applied by a reload replaces the filter entirely, including one that came from RUST_LOG at startup. Removing [log].level in a reload sets info, not whatever RUST_LOG said.
Dry run: test_config
Section titled “Dry run: test_config”--test runs test_config(path), which builds everything a real start builds except the node tasks. It binds nothing and sends no panel request.
pub fn test_config(path: &Path) -> anyhow::Result<()>| Order | Check | Example error, as printed by --test |
|---|---|---|
| 1 | config::load: UTF-8, TOML syntax, unknown keys |
configuration error: config parse error: TOML parse error at line 1, column 1 … |
| 2 | build_outbounds: DNS backend and CA file, every outbound, reserved and duplicate tags |
configuration error: duplicate/reserved outbound tag direct |
| 3 | At least one [[node]] |
configuration error: config defines no [[node]] entries |
| 4 | Per node: PanelClient::new (panel_type, api.node_type) |
configuration error: unknown panel_type "xboard" |
| 5 | Per node: build_router against the pool, default tag direct |
configuration error: route references unknown outbound tag: nope |
| 6 | Per Hysteria 2 node: inbound::validate_hysteria on [node.hysteria] and controller.cert |
prefixed node <node_id>: |
The Hysteria 2 check is the only protocol check, because a Hysteria 2 node’s parameters can be local; every other node type gets its port, transport and TLS flags from the panel, which the dry run does not contact. test_config stops at the first error. This is stricter than a real start: run logs a [[node]] whose panel client or router fails to build, skips it and serves the others, while --test fails on it. It is also stricter than a reload, which compiles routers only for the nodes it adds and does not run the Hysteria 2 check. An error from a file read carries only the OS message, for example configuration error: No such file or directory (os error 2) for a missing dns.ca_file or config file.
Because init_tracing runs first, an invalid [log].level does not fail --test: the dropped directive is reported on stderr and the run still prints Configuration OK.
Signals and shutdown
Section titled “Signals and shutdown”wait_for_shutdown returns on the first SIGINT (tokio::signal::ctrl_c) or, on Unix, SIGTERM (SignalKind::terminate()). If the SIGTERM handler cannot be installed, it waits for SIGINT alone. There is no SIGHUP handler; reload is driven only by the watcher.
The detached signal task cancels the root token. Because every node’s token is a child_token() of the root, the cancellation reaches all nodes at once. The loop then breaks, logs shutting down, cancels the root again (a no-op) and awaits every JoinHandle in handles in order. Each node drains as described in Node task lifecycle, so shutdown takes as long as the slowest node’s teardown and final panel reports. run then returns ExitCode::SUCCESS.
Tokio keeps its signal handlers installed once they are registered, so a second SIGINT or SIGTERM during the drain does not end the process early.
Invariants
Section titled “Invariants”| Invariant | Enforced by | Pinned by |
|---|---|---|
| A reload whose outbound pool, added node or changed panel client fails to build changes nothing | apply_reload builds the pool, then every added node and every changed node’s client, and returns before any side effect |
a_reload_with_a_node_that_does_not_build_changes_nothing in tests/unit/runtime.rs reloads a node-type typo together with a route edit and asserts that the running node was not replaced, that neither the handle’s cfg nor the runtime’s cfg changed, and that an open connection still relays. The build_outbounds errors are reproducible with katana --test |
| A removed node’s port is free before a new node binds | Removal awaits each removed node’s JoinHandle before any spawn_built |
No test; the order of statements in apply_reload |
| A node’s final counters are reported when it stops | NodeManager::run calls tear_down, then report_traffic and report_illegal once bootstrap or serve returns, whether or not the node came up; the runtime awaits the task |
Exercised by vmess_traffic_is_metered_and_reported in tests/unit/e2e.rs, which cancels the token and joins the task before its last assertion; the test also accepts totals from the periodic reports, so it does not isolate the final flush |
| A route edit drops every connection on the node | apply_static → poll_cycle(true) → reconcile → rebuild |
route_change_drops_connections in tests/unit/e2e.rs sends a StaticUpdate::Config over a capacity-16 channel and asserts the open VMess connection ends |
A node never receives a Config for another panel node |
identity covers panel, host, node id, key and panel node type; mismatches are removed and spawned |
an_sspanel_node_is_its_panel_node_id and a_newv2board_node_is_also_the_type_it_asks_for pin which edits change the identity; a_newv2board_type_edit_respawns_the_node turns VLESS off on a newV2board V2ray node and asserts a new channel and that the node no longer serves VLESS |
| An edit to a panel client field takes effect in place | apply_static builds a new PanelClient when panel_type or api changes, then runs one poll |
an_sspanel_api_edit_takes_effect_in_place turns VLESS off on an sspanel node and asserts the same channel and that the node no longer serves VLESS; a_client_edit_takes_effect_without_dropping_connections sets api.rule_list_path and asserts that new flows are refused while an open connection keeps relaying |
| The panel key never appears in reload log lines | display_id omits the key element, and it is the only formatter the runtime uses for identities |
a_node_is_logged_by_its_panel_node_not_its_key asserts the exact rendering for a newV2board and an sspanel node |
| A node that cannot come up keeps retrying, and still stops when cancelled | bootstrap retries try_bootstrap with backoff and watches the token in a biased select! |
a_node_comes_up_once_the_panel_answers, a_node_whose_port_is_taken_comes_up_once_it_is_free and a_node_that_never_bootstraps_still_stops in tests/unit/e2e.rs |
One bad [[node]] at startup never stops the others |
spawn_node returns None and logs; startup requires at least one handle, not all |
No test |
| No reload edit is lost | Ticks during a reload stay queued and start another reload, because the debounce drains only once, before config::load |
No test |
| Reserved tags cannot be redefined | build_outbounds seeds them before inserting entries and rejects existing keys |
No test; reproducible with katana --test |
Failure paths
Section titled “Failure paths”| Failure | Where | Result |
|---|---|---|
| Config does not load at startup | run |
failed to load config: …, exit 1 |
| Pool does not build at startup | run |
failed to build outbounds: …, exit 1 |
No [[node]] at startup |
run |
config defines no [[node]] entries, exit 1 |
| A node does not build at startup | spawn_node |
node <node_id>: …; skipped, the others are started |
| Every node fails to build at startup | run |
no nodes could be started, exit 1 |
| Watcher cannot be set up | run |
config watcher disabled (no live reload): …; runs without reload |
| Config does not load during a reload | select loop | config reload failed, keeping current: …; retried on the next event |
| New pool does not build | apply_reload |
reload: bad outbounds, keeping current config: …; nothing applied |
| Invalid new log level | LogReload closure |
invalid log level …; old filter kept, rest of the reload applied. reload: log level → … is still logged, because the closure returns nothing, and the invalid value is committed, so the same value is not retried |
| An added node, or a changed node’s panel client, does not build | apply_reload, step 2 |
reload: node <display_id>: <error>; keeping current config; nothing applied |
| A reload is refused | select loop | cfg is not committed, so every later pass of the self-feeding cycle refuses it again and logs the error about twice a second until the file is corrected; see Debounce |
| Node route does not compile against a new pool | NodeManager::rebuild_router, for StaticUpdate::Outbounds |
Old router kept; listener still rebuilt |
| Node config edit does not build (panel client or route) | NodeManager::apply_static |
node <node_id>: config edit refused, keeping the running one: …; nothing stored. The runtime has already recorded the table, so the edit is not sent again until the table changes again |
| Node bootstrap attempt fails | NodeManager::bootstrap |
node <node_id>: <reason>; retrying in <n>s; retried with backoff until the node is up or cancelled; see Node task lifecycle |
| Update sent to an ended node | static_tx.send |
Error ignored |
Limits
Section titled “Limits”The bootstrap backoff has named constants in src/manager/node.rs; the runtime’s other numbers are literals in src/runtime.rs and src/manager/node.rs.
| Value | Where | Meaning |
|---|---|---|
500 ms |
Duration::from_millis(500) in run |
Debounce window after the first event |
16 |
mpsc::channel(16) in spawn_built |
Queued StaticUpdates per node before send waits |
1 s |
BOOTSTRAP_RETRY_MIN |
Wait after the first failed bootstrap attempt; doubles after each attempt |
60 s |
BOOTSTRAP_RETRY_MAX |
Longest wait between bootstrap attempts; the poll period caps it when shorter |
| Unbounded | mpsc::unbounded_channel::<()>() in run |
Watcher ticks, drained on every reload |
"info" |
init_tracing and apply_reload |
Log level when [log].level is absent |
5 s |
ApiConfig::timeout_secs |
Panel request timeout when api.timeout is 0 |
60 s, minimum 1 s |
default_update_periodic, NodeManager::poll_period |
Poll period |
src/runtime.rs mounts tests/unit/runtime.rs as its unit-test module. Its behaviour is covered from four directions:
- Reload decisions, in
tests/unit/runtime.rs.- Identity, without any I/O.
an_sspanel_node_is_its_panel_node_idasserts that on sspanel onlypanel_type(other than its case),api.host,api.node_idandapi.keychange the identity, and thatapi.node_type,api.enable_vless,api.timeoutand the other[node.api]fields do not.a_newv2board_node_is_also_the_type_it_asks_forasserts that on newV2boardapi.node_typeand, on a V2ray node,api.enable_vlesschange it too, while a case-onlynode_typeedit,api.timeout,api.speed_limitandapi.rule_list_pathdo not.a_node_is_logged_by_its_panel_node_not_its_keypins thedisplay_idrendering for both panels. - Reloads against a live node. A
Runtimeharness holdscfg,pool, the root token andhandlesthe wayrundoes, starts one node withspawn_nodeagainst a fake panel fromtests/unit/e2e.rs, and callsapply_reloaddirectly. The tests tell a respawn from an in-place edit by comparing the handle’sstatic_txwithsame_channel. They area_newv2board_type_edit_respawns_the_node,a_reload_with_a_node_that_does_not_build_changes_nothing,an_sspanel_api_edit_takes_effect_in_placeanda_client_edit_takes_effect_without_dropping_connections, described under Invariants.
- Identity, without any I/O.
- Node side of the contract, in
tests/unit/e2e.rs. The test helperspawn_nodemirrors the runtime’s own: it builds aPanelClient::NewV2board, callsNodeManager::new, createsmpsc::channel(16)and spawnsnm.run(shutdown, static_rx)against a fake panel.route_change_drops_connectionssends aStaticUpdate::Configwith a different[node.route]and asserts that a live connection ends.vmess_traffic_is_metered_and_reportedcancels the node’s token, awaits the task, then checks the totals the fake panel received.a_node_comes_up_once_the_panel_answershas the fake panel fail the node config request twice with a500and asserts that the node comes up and relays.a_node_whose_port_is_taken_comes_up_once_it_is_freeholds the node’s port while a panel that answers repeated requests with304 Not Modifiedserves it, releases the port after the second node config request, and asserts that the node comes up.a_node_that_never_bootstraps_still_stopshas the panel fail every node config request and asserts that the task ends within2seconds of the token being cancelled.
- Whole-process startup, in
tests/integration/.spawn_katanaintests/support/mod.rsruns the real binary with-c katana.tomlfrom a test directory, which exercisesinit_tracing,run,build_outbounds,spawn_nodeand the watcher on a bare relative path. The interop testsvmess_tcp_plain,vless_ws_tls,trojan_tcp_tlsand the rest oftests/integration/xray_interop.rs, both tests intests/integration/sniff.rs, anda_real_client_proxies_through_a_katana_hysteria_nodeintests/integration/hysteria_interop.rsall start this way. They build their peer (Xray or the Hysteria client) withgo, and printSKIP: …and pass when the toolchain or the peer’s source tree is missing. None of them edits the config file.a_real_client_proxies_through_a_katana_hysteria_nodewritesclient.yamlinto the watched directory after the node has bound its port, which starts the self-feeding reload cycle described under Debounce; every one of those reloads finds nothing changed. The watcher, the debounce and the select loop’s reload branch are therefore exercised with a changed config only by hand;apply_reloaditself is driven by the unit tests above. - The dry run, by hand.
test_configcalls the same builders a real start calls (config::load,build_outbounds,PanelClient::new,build_router), so a small config andkatana --test -c <file>reproduce each row of the dry-run table above.
When you change apply_reload, the step order described above is the contract to preserve, and tests/unit/runtime.rs with its Runtime harness is where a test for a new reload path belongs. How tests are organised and run is described on Testing.