Skip to content

Node manager

Source files: 19 · checked against katana v3.0.1
  • katana/src/manager/node.rs
  • katana/src/manager/mod.rs
  • katana/src/manager/transport.rs
  • katana/src/manager/proxy.rs
  • katana/src/main.rs
  • katana/src/runtime.rs
  • katana/src/serve.rs
  • katana/src/api/mod.rs
  • katana/src/api/newv2board.rs
  • katana/src/api/sspanel.rs
  • katana/src/config.rs
  • katana/src/traffic.rs
  • katana/src/rule.rs
  • katana/src/router.rs
  • katana/tests/unit/e2e.rs
  • katana/tests/unit/runtime.rs
  • katana/tests/unit/traffic.rs
  • katana/tests/unit/rule.rs
  • katana/tests/integration/sniff.rs

Every [[node]] entry in a katana config becomes one NodeManager running in one tokio task. The manager is the node’s control loop: it asks the panel (or, for a locally described Hysteria 2 node, the config file) what the node should look like, keeps a listener that matches, and sends the panel the node’s traffic and audit hits. It also receives the config-file updates the process root fans out on a hot reload.

This page is for contributors who change src/manager/. It covers the manager’s state and locks, the run loop from the bootstrap retries to exit, the reconcile ladder that decides how much of a node to rebuild, the static-update path, the report paths and the small helpers in src/manager/mod.rs. The listener subtree it owns (TransportManager and ProxyManager) is introduced here only as far as the manager drives it; serving, admission and traffic accounting have their own pages.

NodeManager owns:

  • Desired state. The last node description and user list it applied (Cur), which every later change is diffed against.
  • The listener subtree. At most one TransportManager at a time, built by bring_up and dropped by tear_down.
  • The node’s shared registries. One NodeTraffic (per-user counters) and one RuleManager (audit rules and hits). Both outlive every listener rebuild, which is what keeps byte counts continuous across a rebuild.
  • The compiled router. The node’s [node.route] compiled against the global outbound pool, recompiled when either changes.
  • Reporting. One traffic report and one illegal-access report per poll cycle, and a final flush on exit.

It deliberately does not own:

  • the panel HTTP protocol (the PanelClient it holds, see the panel clients page); it only swaps in a new client when a config edit changes the settings the client was built from;
  • node identity and respawn on a reload (src/runtime.rs decides that before it sends an update);
  • per-connection serving, handshake deadlines and admission (src/serve.rs, src/connector.rs).
flowchart TB
  runtime["runtime::spawn_built"]
  nm["NodeManager (one task per node)"]
  api["PanelClient"]
  tm["TransportManager (one listener generation)"]
  pm["ProxyManager (user tables, pre-auth permits)"]
  traffic["NodeTraffic (Arc, shared)"]
  rules["RuleManager (Arc, shared)"]
  runtime -->|"spawns run(), holds static_tx"| nm
  nm -->|"node_info, user_list, node_rule, reports"| api
  nm -->|"bring_up / tear_down"| tm
  tm --> pm
  nm --- traffic
  nm --- rules
  pm --- traffic
src/manager/node.rs
pub struct NodeManager {
api: Mutex<Arc<PanelClient>>,
cfg: Mutex<NodeConfig>,
pool: Mutex<Arc<HashMap<CompactString, Arc<Outbound>>>>,
router: Mutex<Arc<Router<Outbound>>>,
traffic: Arc<NodeTraffic>,
rules: Arc<RuleManager>,
transport: tokio::sync::Mutex<Option<TransportManager>>,
cur: tokio::sync::Mutex<Cur>,
}
impl NodeManager {
pub fn new(
api: PanelClient,
cfg: NodeConfig,
pool: Arc<HashMap<CompactString, Arc<Outbound>>>,
) -> io::Result<Arc<Self>>;
fn api(&self) -> Arc<PanelClient>;
pub async fn run(
self: Arc<Self>,
shutdown: CancellationToken,
mut static_rx: mpsc::Receiver<StaticUpdate>,
);
}

new compiles the initial router with build_router(&cfg.route, &pool, &CompactString::from("direct")), so any router build error fails the constructor: a route or default that names an unknown outbound tag, an invalid cidr or port entry, or geodata that cannot be loaded. At start-up, runtime::spawn_node logs that as node <id>: build router: <error> and skips the node; the process still starts if at least one other node spawned. On a reload, the runtime builds every node the reload adds, router included, before it touches any running node, and refuses the whole reload if one does not build.

Mutex here is parking_lot::Mutex. The fields fall into two lock families:

Field Lock Written by Read by
api parking_lot::Mutex new, apply_static (StaticUpdate::Config) node_info, try_bootstrap, poll_cycle, refresh_rules, reports, all through api()
cfg parking_lot::Mutex apply_static (StaticUpdate::Config) almost every method, for one field at a time
pool parking_lot::Mutex apply_static (StaticUpdate::Outbounds) rebuild_router
router parking_lot::Mutex apply_static (StaticUpdate::Config), rebuild_router bring_up (cloned into the new Dispatcher)
traffic none (Arc, interior locks) bring_up, reconcile, ProxyManager::refresh report_traffic
rules none (Arc, interior locks) refresh_rules report_illegal, every flow’s audit check
transport tokio::sync::Mutex bring_up, tear_down reconcile (listener check, user refresh)
cur tokio::sync::Mutex set_cur poll_cycle, reconcile, apply_static, refresh_rules, report_illegal

The poll loop and the static updates run in the same task, one after the other, so none of these locks is ever contended. They exist only to carry mutable state across &self methods on the shared Arc<NodeManager>. Two rules keep that true if you add code:

  • A parking_lot guard is never held across an .await. Every use is a short expression such as self.cfg.lock().controller.listen_ip.clone(), or a block that copies the fields it needs and drops the guard. api() clones the Arc<PanelClient> out of its lock, so a panel request runs without the guard, and a request already in flight when apply_static swaps the client finishes on the old one. The compiler backs this up: tokio::spawn requires the run future to be Send, and a parking_lot guard is not Send unless lock_api’s send_guard feature is on, which nothing in katana’s dependency graph enables.
  • The two tokio mutexes are never held at the same time. Each method takes one, copies what it needs out, and releases it before taking the other.
src/manager/node.rs
#[derive(Default)]
struct Cur {
node: Option<NodeInfo>,
users: Vec<UserInfo>,
node_tag: CompactString,
}

Cur is the last applied desired state, not the last fetched one. node is None only before bootstrap finishes; apply_static reads that to tell whether a node is up yet. node_tag is the audit and report tag of that state, the key under which refresh_rules stores rules and report_illegal drains hits.

src/manager/mod.rs
pub enum StaticUpdate {
Config(Box<NodeConfig>),
Outbounds(Arc<HashMap<CompactString, Arc<Outbound>>>),
}

The runtime sends these over a bounded mpsc channel of capacity 16, created when it spawns the node (runtime::spawn_built). Config carries this node’s new [[node]] block and is sent only when the block changed and the node’s identity did not. Identity is the tuple (panel_type lowercased, api.host, api.node_id, api.key, panel node type); a change to any of those makes the runtime cancel the node and spawn a new one instead. The panel node type is the node_type a newV2board request names (vless for a V2ray-family node, api.node_type V2ray, Vmess or Vless, with enable_vless, otherwise the lowercased api.node_type), because UniProxy finds a node by its id and that type. It is empty for SSPanel, which finds a node by id alone. A Config may therefore change [node.api] fields the panel client was built from, such as timeout, speed_limit or rule_list_path; the node rebuilds its client for those (see StaticUpdate::Config). Outbounds carries the rebuilt global pool and is sent to every node the runtime holds a handle for whenever the [[outbound]] list changed.

src/manager/transport.rs
impl TransportManager {
pub async fn start(
node: &NodeInfo,
cert: &CertConfig,
listen_ip: &str,
enable_vless: bool,
sniff: bool,
users: &[UserInfo],
traffic: Arc<NodeTraffic>,
rules: Arc<RuleManager>,
router: Arc<Router<Outbound>>,
node_tag: CompactString,
hysteria: &HysteriaConfig,
) -> io::Result<Self>;
pub fn proxy(&self) -> &Arc<ProxyManager>;
pub async fn shutdown(self);
}
src/manager/proxy.rs
impl ProxyManager {
pub fn refresh(&self, node: &NodeInfo, users: &[UserInfo], enable_vless: bool);
pub fn retire_all(&self);
}

TransportManager::start builds everything that can fail (transport, protocol tables or Hysteria authenticator and server) and binds the socket before it commits the staged user set, so a bad description never half-binds. ProxyManager::refresh replaces only the user tables of a live listener. The manager calls nothing else on the subtree.

src/manager/node.rs
if self.bootstrap(&shutdown, &mut static_rx).await {
self.serve(&shutdown, &mut static_rx).await;
}
self.tear_down().await;
self.report_traffic().await;
self.report_illegal().await;

run has three phases. bootstrap brings the node up, retrying until it is up or the token is cancelled, and returns whether it came up. serve is the steady-state loop and runs only after a successful bootstrap. The exit steps run in both cases, including for a node that was cancelled before it ever came up.

sequenceDiagram
  participant RT as runtime
  participant NM as NodeManager::run
  participant P as PanelClient
  participant TM as TransportManager
  RT->>NM: tokio::spawn(nm.run(shutdown, static_rx))
  loop bootstrap, until up or shutdown
    NM->>P: forget_etags(), node_info(), user_list()
    NM->>TM: bring_up(node, users)
    NM->>NM: set_cur(Some(node), users, tag)
    NM->>P: node_rule() unless disable_get_rule
    Note over NM: on failure, wait and apply static updates
  end
  loop serve, until shutdown
    alt interval tick
      NM->>NM: poll_cycle(false)
    else static update
      NM->>NM: apply_static(update), maybe reset interval
    end
  end
  NM->>TM: tear_down()
  NM->>P: report_user_traffic(), report_illegal()

bootstrap calls try_bootstrap until one attempt succeeds. One attempt runs these steps and stops at the first that fails:

  1. Fresh reads. api().forget_etags() clears the client’s ETag cache. Nothing a failed attempt read was applied, and a panel answers a request it has answered before with 304 Not Modified, which would leave the attempt with nothing to start from.
  2. Node description. node_info returns the panel’s answer, except for a Hysteria 2 node whose [node.hysteria].port is non-zero (next section).
  3. Port check. A description with port == 0 is refused.
  4. Users. api().user_list() must return a list. An empty list is accepted.
  5. Tag. node_tag(&node, &listen_ip).
  6. Listener. bring_up(&node, &users). With an empty user list it commits an empty registry and binds nothing, so the node comes up dark and the first poll that sees users builds the listener.
  7. State. set_cur(Some(node), users, tag).
  8. Rules. refresh_rules(), unless controller.disable_get_rule is set.

An attempt fails with one of these reasons:

Condition Reason
node_info returns Ok(None) panel returned no node info
node_info returns Err node_info failed: <error>
the description has port 0 panel returned port 0
user_list returns Ok(None) panel returned no user list
user_list returns Err user_list failed: <error>
bring_up fails (build or bind) initial start failed: <error>

The newV2board client rejects a server_port of 0 itself, with the error newV2board: server port must be > 0, so on that panel a zero port fails the attempt as node_info failed. The port-0 row is reached when the client passes a zero port through, for example SSPanel with an offset_port_node of "0".

After a failure, bootstrap logs node <id>: <reason>; retrying in <n>s at error and waits. The wait starts at BOOTSTRAP_RETRY_MIN (1 s) and doubles after each failure up to BOOTSTRAP_RETRY_MAX (60 s), and it is never longer than the poll period (poll_period()), so a node that syncs every few seconds also retries every few seconds. With the default update_periodic of 60 s the waits are 1, 2, 4, 8, 16, 32 and then 60 seconds.

Each wait is one tokio::time::sleep, created once, so updates that arrive during the wait cannot push the retry back. While it waits, the node applies static updates as they arrive, which also keeps the runtime’s bounded channel from filling:

  • An Outbounds update stores the pool and recompiles the router. There is no listener to rebuild yet.
  • A Config update is applied or refused (see StaticUpdate::Config) and ends the wait at once, because the edit may be the fix. With the node not up, it only stores the new config, client and router, and the next attempt reads with the new client and builds from the new config.

Both the attempt and the wait sit in a biased select! whose first branch is the cancellation token, so a node that is cancelled while it is still coming up stops at once, dropping any panel request in flight.

While a node is in this loop it serves nothing on its port. The process and the other nodes are not affected.

node_info reads api.node_type, [node.hysteria] and api.speed_limit from the current cfg on every call. When the node type parses as NodeType::Hysteria2 and hysteria.port != 0, it builds the NodeInfo itself and never asks the panel:

NodeInfo field Value
node_type NodeType::Hysteria2
port hysteria.port
speed_limit api.speed_limit in Mbps, times 1_000_000.0 / 8.0, cast to u64 (the cast saturates, so a negative value becomes 0, unlimited)
transport, enable_tls Transport::Tcp, true: a QUIC listener has no stream transport and its TLS is part of the handshake
obfs_type, obfs_password hysteria.obfs, hysteria.obfs_password, empty when unset
every other field empty or false

With hysteria.port == 0 a Hysteria 2 node asks the panel like any other node type. The settings a panel cannot express (credential source, UDP relay, masquerade) come from [node.hysteria] in both cases, read by bring_up.

Because this path cannot fail, a locally described Hysteria 2 node never falls back to the last applied description in poll_cycle.

src/manager/node.rs
fn poll_period(&self) -> Duration {
Duration::from_secs(self.cfg.lock().controller.update_periodic.max(1))
}

update_periodic defaults to 60 seconds (default_update_periodic in src/config.rs) and is clamped to at least 1 second. serve builds a tokio::time::interval from it and immediately awaits the first tick, which an interval completes at once. The first poll_cycle therefore runs one full period after bootstrap, not straight after it. The interval uses tokio’s default MissedTickBehavior::Burst: when a cycle takes longer than the period, for example because several panel requests each ran into their timeout, the missed ticks fire back to back.

src/manager/node.rs
loop {
tokio::select! {
_ = shutdown.cancelled() => break,
_ = interval.tick() => self.poll_cycle(false).await,
Some(update) = static_rx.recv() => {
self.apply_static(update).await;
let new_period = self.poll_period();
if new_period != period {
period = new_period;
interval = tokio::time::interval(period);
interval.tick().await;
}
}
}
}
  • Polls and static updates are strictly serialized. A static update that arrives during a poll waits for the poll to finish, and the other way round.
  • After each static update the loop recomputes the period. When update_periodic changed, it replaces the interval and again consumes the immediate tick, so the next poll comes one new period after the update.
  • If every sender is dropped, recv() yields None, the Some(update) pattern fails and select! disables that branch for that iteration. The loop keeps polling.
  • Cancellation is observed only between iterations. A cancel that arrives during poll_cycle or apply_static (which may itself run a poll cycle) takes effect when that call returns. When several branches are ready at once, tokio::select! picks one at random, so a static update already queued when the token is cancelled may or may not be applied before the loop exits. The panel requests inside a cycle are bounded by the HTTP client timeout, api.timeout seconds or 5 when it is 0 (ApiConfig::timeout_secs).

The listener goes down first, so no relay is still adding bytes when the counters are read, and the final report carries everything up to the moment of shutdown instead of dropping up to one poll period of traffic. The flush is a single attempt: if the panel rejects it, the failure is logged and the bytes go with the task. For a node cancelled during bootstrap there is usually nothing to tear down or report; the exit steps still run so that a listener generation an attempt had already bound is shut down in order (see rebuild, bring_up and tear_down). The runtime awaits each node’s task after cancelling it, both on process shutdown and when a reload removes a node, so the flush completes before the process exits.

src/manager/node.rs
async fn poll_cycle(&self, rebuild: bool)

The timer calls it with rebuild = false. apply_static calls it once, straight after a config edit, with rebuild = true when the edit changes something the listener is built from (see StaticUpdate::Config). Each call runs these steps in order, whatever the earlier ones returned:

Step Call On Ok(None) On Err
1 node_info() reuse cur.node log node <id>: node_info: <error> at warn, reuse cur.node
2 api().user_list() reuse cur.users log node <id>: user_list: <error> at warn, reuse cur.users
3 reconcile(node, users, rebuild)
4 refresh_rules(), unless disable_get_rule keep the rules log at warn, keep the rules
5 report_traffic() restore residuals, log at warn
6 report_illegal() log at warn

Ok(None) is the panel clients’ “not modified” answer (HTTP 304 against a cached ETag). Treating it and a failed request the same way, as “nothing new”, is what lets a flaky panel leave a running node alone: the fallback description is exactly the one already applied, so reconcile finds nothing to change.

A description whose port is 0 is also treated as a miss, because a listener could not serve it. The cycle logs node <id>: refreshed port is 0, keeping the last one at error, keeps cur.node, and still reconciles the user list the panel returned, so user changes keep applying while the panel’s description is unusable. As at bootstrap, a newV2board panel that returns server_port 0 does not reach this check: its client returns Err, so the cycle takes the step 1 fallback instead.

The fallback reads cur.node with .expect("node_info present after bootstrap"). It holds because serve runs only after bootstrap returned true, which happens only after try_bootstrap ran set_cur(Some(node), …), and every later set_cur also passes Some. apply_static runs a poll cycle only when cur.node is already Some.

src/manager/node.rs
async fn reconcile(&self, node: NodeInfo, users: Vec<UserInfo>, rebuild: bool)

reconcile takes the first branch that applies, in the order that drops the fewest connections for the change at hand. rebuild forces the full-rebuild branch, for a local edit that the panel’s answer cannot show:

stateDiagram-v2
  [*] --> UsersCheck
  UsersCheck --> TearDown: users is empty
  UsersCheck --> ListenerCheck: users present
  ListenerCheck --> FullRebuild: rebuild forced, or no transport
  ListenerCheck --> FullRebuild: transport_eq or protocol_eq is false
  ListenerCheck --> UserCheck: same listener and protocol
  UserCheck --> Refresh: user_set_differs or speed_limit changed
  UserCheck --> Unchanged: nothing differs
  TearDown --> SetCur
  FullRebuild --> SetCur
  Refresh --> SetCur
  Unchanged --> SetCur
  SetCur --> [*]
Branch Trigger Action Effect on connections
User-less users.is_empty() tear_down(), then traffic.commit(traffic.prepare(Vec::new())) every connection drops; every counter moves to draining
Full rebuild rebuild is true, or no TransportManager, or !node.transport_eq(prev), or !node.protocol_eq(prev) rebuild(&node, &users) every connection drops
In-place refresh user_set_differs(&cur.users, &users), or node.speed_limit differs from the previous node’s transport.proxy().refresh(&node, &users, enable_vless) unchanged users keep their connections; departed and rebound users are retired
No change none of the above nothing none

Every branch ends with set_cur(Some(node), users, tag), where tag is computed from the new description and the current controller.listen_ip.

A few details matter when you change this code:

  • “No transport” is the retry path. When a rebuild fails, transport stays None. The next cycle takes the full-rebuild branch again because of !has_transport, which is how a node that went dark after a failed rebuild, or after a user-less period, comes back without any explicit retry loop.
  • enable_vless comes from the config. The in-place refresh passes cfg.api.enable_vless, not node.enable_vless, to ProxyManager::refresh, matching what bring_up passes to TransportManager::start.
  • speed_limit is a user-bucket change. Neither equality below includes it, because the node limit only feeds each user’s rate (determine_rate). A change creates fresh counters for the users whose effective rate changed; their live connections stay open, and flows they open afterwards run at the new rate.
  • The refresh is transactional. ProxyManager::refresh builds the replacement tables before publishing anything. If that build fails it logs proxy refresh build failed, keeping current: <error> and leaves both the tables and the registry untouched. When it succeeds it commits the registry first and swaps the tables second, so a newly added user is never authenticated by the new table while the registry still refuses their flows.

NodeInfo::transport_eq and NodeInfo::protocol_eq in src/api/mod.rs split the description into the two layers a rebuild cares about:

NodeInfo field Compared by A change means
port, transport, host, path, service_name, authority, enable_tls, header, headers, enable_reality, accept_proxy_protocol, obfs_type, obfs_password transport_eq full rebuild
node_type, enable_vless, vless_flow, cypher_method, server_key protocol_eq full rebuild
speed_limit neither in-place refresh

The obfuscation fields are in transport_eq because they are part of the wire: without the rebuild the listener would keep the old key and lock out every client that took the new one.

src/manager/node.rs
async fn rebuild(&self, node: &NodeInfo, users: &[UserInfo]);
async fn bring_up(&self, node: &NodeInfo, users: &[UserInfo]) -> io::Result<()>;
async fn tear_down(&self);
  • rebuild is tear_down followed by bring_up. A bring_up error is logged as node <id>: rebuild failed: <error> and the node stays dark until the next cycle retries.
  • bring_up with no users commits an empty registry and returns Ok(()) without binding. Otherwise it copies controller.cert, controller.listen_ip, api.enable_vless, !controller.disable_sniffing and [node.hysteria] out of cfg, clones the current router, takes the transport lock, calls TransportManager::start, stores the result and logs node <id>: listening on <listen_ip>:<port>. Because the lock is taken before start, no .await lies between a finished start and the store. A bootstrap attempt that shutdown cuts short after its bind has therefore already stored its generation, and run’s tear_down shuts it down in order, including the wait for a QUIC endpoint to release its UDP socket.
  • tear_down takes the TransportManager out of its slot and awaits TransportManager::shutdown, which cancels the accept scope and waits for every task under it, retires every user lease, and for a Hysteria 2 node awaits the QUIC endpoint until it has released its UDP socket. The last step is what lets rebuild bind the same UDP port immediately afterwards.

TransportManager also implements Drop as a backstop: it cancels the scope and retires every user if a generation is ever dropped without shutdown. It cannot wait for the tasks to finish, so every path in the manager uses shutdown.

The bind happens in a fresh TransportManager after the old one is gone, so a full rebuild has a short window in which the port is closed. A rebuild whose bind fails (for example because another process took the port in that window) leaves the node dark until a later cycle succeeds.

stateDiagram-v2
  [*] --> Bootstrap
  Bootstrap --> Bootstrap: attempt failed, wait and retry
  Bootstrap --> Dark: no users
  Bootstrap --> Serving: bring_up succeeded
  Bootstrap --> Stopped: shutdown
  Serving --> Serving: in-place refresh or no change
  Serving --> Serving: full rebuild succeeded
  Serving --> Dark: users empty, or rebuild failed
  Dark --> Serving: next cycle rebuilds with users
  Serving --> Stopped: shutdown
  Dark --> Stopped: shutdown
  Stopped --> [*]

Dark means the task is alive but holds no TransportManager. Bootstrap loops until the node is up or stopped; the task never ends while its token is live.

src/manager/node.rs
async fn apply_static(&self, u: StaticUpdate)

A config edit is taken whole or not at all. apply_static first builds what the edit needs, then stores it, then applies it with one poll cycle:

flowchart TB
  start["StaticUpdate::Config(new)"]
  client{"panel_type or api changed?"}
  newclient["PanelClient::new(new)"]
  route{"route changed?"}
  router["build_router(new route, pool)"]
  refused["refused: keep the running config"]
  store["store cfg, router, client"]
  resync{"bootstrapped, and rebuild or new client?"}
  poll["poll_cycle(rebuild)"]
  stored["stored only, read later"]
  start --> client
  client -->|yes| newclient
  client -->|no| route
  newclient -->|Err| refused
  newclient -->|Ok| route
  route -->|yes| router
  route -->|no| store
  router -->|Err| refused
  router -->|Ok| store
  store --> resync
  resync -->|yes| poll
  resync -->|no| stored
  1. Client. When panel_type or any [node.api] field differs, PanelClient::new builds a client from the new block. The client copies settings at construction (the node type, VLESS, the speed override, the rule file, the timeout), so without a new one the node would keep answering to the old values. The new client calls inherit_routes on the old one: a newV2board client takes over the routes the old one last read, so the audit rules derived from their block entries hold until the new client reads the node config itself.
  2. Router. When route differs, build_router compiles the new [node.route] against the current pool with "direct" as the fallback tag.
  3. Refusal. If either build fails, the manager logs node <id>: config edit refused, keeping the running one: <error> at error and returns. Nothing from the edit is stored, including fields that needed no build.
  4. Store. The new cfg replaces the old one. A new router replaces router and logs node <id>: router rebuilt; a new client replaces api.
  5. Apply. The edit needs a listener rebuild when any of these changed: route, controller.listen_ip, controller.cert, controller.disable_sniffing, api.enable_vless or anything in [node.hysteria]. The listener is built from these local settings, and the panel’s answer does not change when they do. If the node is bootstrapped (cur.node is Some) and the edit needs a rebuild or brought a new client, apply_static runs one poll_cycle(rebuild) at once.

That poll is a full cycle. A new client holds no ETags, so the panel answers in full and the node and users are read as the edited settings read them. reconcile then takes the full-rebuild branch when rebuild is set, and otherwise the narrowest branch the fresh answer needs. The poll also refreshes the rules and sends the traffic and illegal-access reports, as a timer poll does. It does not reset the poll interval.

For a node that is still in bootstrap, nothing is polled: the edit is stored, or refused as in step 3 if it does not build. Either way the wait ends at once and the next attempt reads with the new client and builds from the new config.

Changed field What happens Takes effect
route (the whole [node.route] table) new router, then poll_cycle(true): full rebuild now
controller.listen_ip, controller.cert, controller.disable_sniffing poll_cycle(true): full rebuild now
any [node.hysteria] field poll_cycle(true): full rebuild now
api.enable_vless new client, then poll_cycle(true): full rebuild. On newV2board with a V2ray or Vmess node this changes the identity instead, and the runtime respawns the node now
api.speed_limit, api.rule_list_path new client, then poll_cycle(false): an in-place refresh, unless the fresh answer changes the transport or protocol; the same poll refreshes the rules now
api.timeout and the other [node.api] fields outside the identity (vless_flow, device_limit, disable_custom_config, node_type on SSPanel, a case-only node_type change on newV2board), and a case-only panel_type change new client, then poll_cycle(false): the reconcile ladder decides from the fresh answer now
controller.update_periodic stored; the select loop replaces the interval the next poll, one new period later
controller.disable_get_rule, controller.disable_upload_traffic stored the next poll cycle

A routing edit rebuilds the whole transport rather than cancelling authenticated connections, because the requirement is that a routing edit drops every connection on the node, including ones still in their handshake. Cancelling user leases alone would miss those. The rebuild keeps byte continuity: NodeTraffic persists, so prepare reuses the counters of users whose key, uid and rate did not change.

The new pool replaces pool, rebuild_router recompiles the router, and apply_route_change rebuilds the transport from cur.node and cur.users. Every node receives this update when the pool changed, so an [[outbound]] edit drops every connection on every node that is up. A node still in bootstrap has no cur.node yet, so it only keeps the new router for its next attempt.

runtime::apply_reload sends Outbounds to every running node before it removes, reconfigures or adds any node. Two consequences follow:

  • A node that the same reload removes can still rebuild once against the new pool. The runtime cancels its token right after queuing the update, and when both are ready the serve loop picks one at random, so whether that rebuild happens depends on timing.
  • A reload that changes both [[outbound]] and a node’s [[node]] block (in a way that triggers a rebuild) rebuilds that node twice in a row: once for Outbounds, then once for Config through its poll_cycle(true). The select loop handles them one after the other, possibly with a timer poll in between.
src/manager/node.rs
fn rebuild_router(&self)

Only the Outbounds path calls it. It compiles the current cfg.route against the new pool with "direct" as the fallback tag. On success it replaces router and logs node <id>: router rebuilt. On failure it logs node <id>: route rebuild failed, keeping current: <error> and keeps the old router. apply_route_change still runs, so the fresh listener binds with the router that was current before the pool changed.

Most router errors are caught before this point. The runtime builds the new pool before it sends Outbounds and rejects the whole reload if that fails, so a bad [[outbound]] never reaches a node. It also builds every node the reload adds, router included, before it touches a running node. apply_static refuses a route edit that does not compile. What is left for rebuild_router is the unchanged route of a node the reload keeps: a route or default that names a tag the new pool no longer has (for example after an [[outbound]] was removed), or geodata that can no longer be loaded.

src/manager/node.rs
async fn refresh_rules(&self)

api().node_rule() returning Ok(Some(rules)) replaces the rules stored under cur.node_tag through RuleManager::update, which skips the write when the new list has the same ids and pattern strings in the same order. Ok(None) keeps the current rules, and Err is logged as node <id>: node_rule: <error> at warn. Only the SSPanel client makes a request here and can answer Ok(None) or Err. The newV2board client always returns Ok(Some): it builds the list from api.rule_list_path and the block routes cached by its last node_info that returned a body (or inherited from the client it replaced), without contacting the panel.

src/manager/node.rs
async fn report_traffic(&self)

With controller.disable_upload_traffic set, it calls traffic.clear_residuals() and traffic.prune_draining() and returns. Nothing will ever be sent, so neither set may grow for the life of the process.

Otherwise:

  1. traffic.snapshot() returns one row per live counter, one per draining counter, and one per parked residual. A draining counter with no writer left (Arc::strong_count == 1) is emitted as a counterless row and removed from the draining set.
  2. Rows with up == 0 && down == 0 are skipped.
  3. The rest are summed per uid with saturating adds. A residual or draining row can coexist with a live counter for the same user, and the newV2board payload is keyed by uid, so two rows would let one overwrite the other.
  4. Rows with a counter go to a commit list, rows without one to a residual list.
  5. If no uid has bytes, nothing is sent.
  6. One api().report_user_traffic(&reports) call carries every uid.
  7. On success, every counter in the commit list gets commit_reported(up, down), which subtracts exactly the reported amounts so bytes added during the request stay for the next report.
  8. On failure, traffic.restore_residuals(residuals) parks the counterless rows again, and the log reads node <id>: report traffic: <error>. Live counters were not touched, so they still hold their bytes.

This is what makes a failed report lose nothing and a retry bill nothing twice. The registry side (prepare, commit, draining and residuals) is described on the traffic accounting page.

src/manager/node.rs
async fn report_illegal(&self)

It drains the hits recorded under cur.node_tag (RuleManager::drain) and, when there are any, sends them with api().report_illegal. The hits are drained before the request, so a failed report is logged as node <id>: report illegal: <error> and those hits are not sent again. The newV2board client’s report_illegal is a no-op, and the SSPanel client leaves out hits of local rules (rule id -1, from rule_list_path) and sends nothing when no panel rule was hit.

src/manager/mod.rs
pub(crate) fn user_key(by_email: bool, u: &UserInfo) -> Option<AuthKey>;
pub(crate) fn user_tag(by_email: bool, u: &UserInfo) -> Option<Arc<UserTag>>;
pub(crate) fn node_tag(node: &NodeInfo, listen_ip: &str) -> CompactString;
pub(crate) fn build_user_entries(
node: &NodeInfo,
users: &[UserInfo],
) -> (Vec<UserEntry>, Vec<UserInfo>);
pub(crate) fn user_set_differs(a: &[UserInfo], b: &[UserInfo]) -> bool;
Helper Behaviour
user_key With by_email, AuthKey::Name(traffic_email(u)): the panel email, or the uid as a string when the email is empty. Otherwise AuthKey::Uuid of u.uuid, or None when it does not parse. by_email comes from NodeType::keys_by_email, which is false for V2ray and true for Trojan, Shadowsocks and Hysteria2.
user_tag user_key wrapped as Arc<UserTag> with the user’s uid: the payload a protocol’s user table carries for each credential.
node_tag Type_listenip_port, where Type is V2ray, Trojan, Shadowsocks or Hysteria2. For example V2ray_0.0.0.0_443. Used as the audit rule key, the dispatcher tag and the handshake-failure log tag.
build_user_entries For every user with a key, one UserEntry { key, uid, rate } with rate = determine_rate(node.speed_limit, u.speed_limit), plus the list of those users. A user without a key, which can only happen on a V2ray node whose uuid does not parse, is skipped with skipping user <uid>: uuid is not a valid UUID at warn; the log names the uid and never the credential.
user_set_differs Order-independent set comparison on the whole UserInfo (it derives Eq and Hash). Different lengths differ; otherwise it checks that every entry of b is in a HashSet of a. Any field change, a password or speed limit included, counts as a difference.

determine_rate in src/traffic.rs returns the smaller of the two limits, where 0 means unlimited on that side: (0, 0) is 0, (0, u) is u, (n, 0) is n, (n, u) is n.min(u).

Both TransportManager::start and ProxyManager::refresh go through build_user_entries, and both pass the same by_email to user_tag. That is the single place the registry key and the protocol table’s key are derived from, so the two cannot disagree about how a node type keys its users.

Invariant Enforced by Pinned by
A user-set change keeps unchanged users’ live connections and drops departed users’ connections. reconcile routes user changes to ProxyManager::refresh; NodeTraffic::prepare keeps counters with the same key, uid and rate and lists departed keys for retirement. unchanged_user_survives_user_refresh in tests/unit/e2e.rs
A Hysteria 2 user refresh never rebinds the UDP socket. Tables::Hysteria replaces only the authenticator (Hy2Inbound::set_authenticator); admission refuses a retired user’s new flows. a_retired_user_stops_while_the_rest_keep_their_connections, repeated_user_refreshes_never_disturb_a_live_connection in tests/unit/e2e.rs
A routing edit drops every connection on the node. apply_static builds the new router and runs poll_cycle(true), whose reconcile takes the full-rebuild branch. route_change_drops_connections in tests/unit/e2e.rs
A node that cannot come up keeps trying and comes up once the cause clears. bootstrap retries try_bootstrap with a capped, doubling wait; each attempt calls forget_etags so a 304 cannot leave it without a node or users. a_node_comes_up_once_the_panel_answers, a_node_whose_port_is_taken_comes_up_once_it_is_free in tests/unit/e2e.rs
A node stops promptly while it is still retrying its bootstrap. bootstrap watches the token first in a biased select!, during each attempt and each wait. a_node_that_never_bootstraps_still_stops in tests/unit/e2e.rs
A config edit to a field the panel client was built from takes effect in the running node. apply_static builds a new PanelClient and runs poll_cycle, which reads without ETags and reconciles. an_sspanel_api_edit_takes_effect_in_place, a_client_edit_takes_effect_without_dropping_connections in tests/unit/runtime.rs
A reload with a node that does not build leaves the running node untouched. runtime::apply_reload builds added nodes and checks every changed node’s panel client before it applies anything; apply_static refuses an edit whose client or router does not build. a_reload_with_a_node_that_does_not_build_changes_nothing in tests/unit/runtime.rs
Every relayed byte is attributed to the authenticated uid and reported. bring_up shares one NodeTraffic with the dispatcher; report_traffic sums per uid. vmess_traffic_is_metered_and_reported, proxy_outbound_relays_and_meters, a_hysteria_node_relays_and_meters in tests/unit/e2e.rs
A panel-described Hysteria 2 node takes its port and obfuscation from the panel. node_info asks the panel when hysteria.port == 0; obfs_type and obfs_password are part of transport_eq. a_panel_described_hysteria_node_serves_obfuscated_traffic in tests/unit/e2e.rs
A credential the panel never issued does not proxy. The authenticator is built from build_user_entries’ valid users only. a_hysteria_node_refuses_an_unknown_credential in tests/unit/e2e.rs
A failed report loses no bytes and a retry bills none twice. commit_reported runs only on success and subtracts the reported amounts; residual rows are restored on failure. restored_residuals_are_retried, commit_reported_preserves_concurrent, rate_change_drains_old_counter_and_reports_once, a_departed_users_late_bytes_are_still_reported, rebound_credential_reports_the_old_uid_separately in tests/unit/traffic.rs
The effective user rate is the smaller non-zero limit. determine_rate determine_rate_min_nonzero in tests/unit/traffic.rs
An identical rule list does not replace the stored one; drained hits are cleared. RuleManager::update compares ids and pattern strings; RuleManager::drain empties the set. update_skips_identical_ruleset, detect_records_and_drains in tests/unit/rule.rs
disable_sniffing reaches the listener. bring_up passes !controller.disable_sniffing to TransportManager::start. a_sniffed_host_reaches_a_domain_rule, disable_sniffing_stops_the_domain_rule_matching in tests/integration/sniff.rs
cur.node is Some whenever poll_cycle runs. serve runs only after bootstrap returned true, which requires try_bootstrap’s set_cur(Some(node), …); apply_static polls only when cur.node is Some; every other set_cur passes Some. no test; poll_cycle would panic the node task otherwise
No listener generation is leaked. Every path goes through tear_down, which awaits TransportManager::shutdown; Drop cancels as a backstop. no dedicated test
Failure Where Result
Router build fails at spawn (unknown outbound tag, bad matcher, geodata) NodeManager::new at start-up the node is not spawned, and the process exits only if no node spawned at all; on a reload that adds the node, the whole reload is refused
Bootstrap node_info or user_list error or Ok(None), port 0, or bring_up error try_bootstrap attempt fails and is logged; retried after 1 s, doubling to 60 s and capped at the poll period; a Config edit retries at once
Bootstrap user list is empty try_bootstrap node comes up dark; the first poll with users builds the listener
Poll node_info or user_list error poll_cycle last applied value reused; nothing changes
Poll description with port 0 poll_cycle last applied description kept; the panel’s user list is still reconciled
Config edit whose client or router does not build apply_static edit refused whole; the node keeps running on its current config
Full rebuild fails rebuild node dark; retried every cycle through !has_transport
In-place refresh build fails ProxyManager::refresh tables and registry untouched
Router rebuild after an Outbounds update fails rebuild_router old router kept; the transport is still rebuilt
Traffic report fails report_traffic residuals restored, counters untouched; retried next cycle
Illegal report fails report_illegal hits dropped
Final flush fails run exit bytes lost with the task

Cancellation comes from the node’s CancellationToken, a child of the process root token. The runtime cancels it on SIGINT or SIGTERM, and when a reload removes the node or changes its identity. Bootstrap watches the token with a biased select!, both around each attempt and during each wait, so a node cancelled before it came up stops at once and drops any panel request in flight. Inside the serve loop, cancellation takes effect between iterations. The runtime then awaits the task, which covers the listener teardown and both final reports.

Name Value Where
update_periodic default 60 s default_update_periodic in src/config.rs
Minimum poll period 1 s NodeManager::poll_period (update_periodic.max(1))
Static update channel 16 updates mpsc::channel(16) in runtime::spawn_built; the runtime’s send().await waits when it is full; the node drains it in both bootstrap and serve
First bootstrap retry wait 1 s BOOTSTRAP_RETRY_MIN in src/manager/node.rs
Longest bootstrap retry wait 60 s, or the poll period if shorter BOOTSTRAP_RETRY_MAX in src/manager/node.rs; NodeManager::bootstrap caps each wait at poll_period()
Panel request timeout api.timeout s, 5 s when 0 ApiConfig::timeout_secs in src/config.rs
Router fallback tag "direct" NodeManager::new, apply_static, rebuild_router
Pre-auth streams per listener generation (stream nodes) 512 MAX_PREAUTH_STREAMS_PER_NODE in src/manager/proxy.rs
Handshake-failure warning threshold more than 10 per second HANDSHAKE_FAILURE_ALERT_PER_SEC in src/manager/proxy.rs

The last two belong to the ProxyManager a bring_up creates, so a full rebuild resets them; an in-place refresh does not.

The manager has no unit tests of its own. It is exercised end to end by tests/unit/e2e.rs, which src/main.rs compiles into the binary’s test harness as the e2e module, and by the reload tests in tests/unit/runtime.rs, which src/runtime.rs compiles as its tests module.

Each e2e test starts a hand-rolled fake newV2board panel on a loopback port that serves /UniProxy/config and /UniProxy/user and records every /UniProxy/push body. fake_panel, fake_panel_dynamic and fake_panel_with_config are thin wrappers over fake_panel_quirky, which takes a Quirks value and also returns how many times the node config was requested:

Quirks field Effect
config_failures answer that many /UniProxy/config requests with 500 Internal Server Error before the first real answer, like a panel that is not up yet
etags tag every answer with an ETag, and answer a request that sends the tag back with 304 Not Modified

fake_sspanel serves an SSPanel mod_mu panel with one V2ray node, described by the legacy server string, and one user. test_node_cfg sets update_periodic = 1, and the e2e spawn_node helper builds a real newV2board PanelClient and NodeManager from the config it is given, spawns run, and returns the static-update sender so the channel stays open. The runtime tests instead hold a node the way runtime::run does and pass edited configs through apply_reload.

Test What it proves about the manager
vmess_traffic_is_metered_and_reported bootstrap binds a VMess node; report_traffic delivers the relayed bytes under uid 1001
unchanged_user_survives_user_refresh adding a user keeps user A’s connection; removing A drops it; A’s bytes from both phases are reported
proxy_outbound_relays_and_meters the router built in new reaches a proxy outbound and metering still attributes the bytes
route_change_drops_connections StaticUpdate::Config with a different route drops a live connection
a_hysteria_node_relays_and_meters the local Hysteria 2 description (hysteria.port != 0) binds, relays and reports
a_hysteria_node_refuses_an_unknown_credential the Hysteria 2 authenticator holds only panel users
a_retired_user_stops_while_the_rest_keep_their_connections a Hysteria 2 refresh retires user B while A’s QUIC connection keeps working
a_panel_described_hysteria_node_serves_obfuscated_traffic with hysteria.port == 0 the port and Salamander key come from the panel
repeated_user_refreshes_never_disturb_a_live_connection several consecutive refreshes leave one QUIC connection intact
a_node_comes_up_once_the_panel_answers with the first two config requests failing, bootstrap retries and the node binds and relays after the third
a_node_whose_port_is_taken_comes_up_once_it_is_free a failed bind is retried, and against a panel that answers 304 via ETags each attempt still reads the node and users in full
a_node_that_never_bootstraps_still_stops a node retrying against a panel that always fails stops within 2 seconds of cancellation

The reload tests in tests/unit/runtime.rs that reach the manager:

Test What it proves about the manager
an_sspanel_node_is_its_panel_node_id, a_newv2board_node_is_also_the_type_it_asks_for which [[node]] edits change the identity (a respawn) and which reach the node as a Config update
a_newv2board_type_edit_respawns_the_node turning enable_vless off on a newV2board V2ray node respawns it, and the new node serves VMess
a_reload_with_a_node_that_does_not_build_changes_nothing an unknown node_type together with a route edit leaves the running node and its live connection untouched
an_sspanel_api_edit_takes_effect_in_place on SSPanel, turning enable_vless off reaches the running node, which rebuilds its client and listener and serves VMess
a_client_edit_takes_effect_without_dropping_connections a rule_list_path edit gets a new client whose rules refuse new flows to the target at once, while an open connection keeps relaying

Run them from the katana checkout:

Terminal window
cargo test --locked e2e::
cargo test --locked runtime::tests::

No test covers the port-0 checks, the Outbounds update, the per-uid merge in report_traffic or the exit flush in isolation. If you change any of them, add a test in tests/unit/e2e.rs. The port-0 checks in try_bootstrap and poll_cycle need a description whose port is 0 after parsing. The newV2board client checks server_port for 0 as an i64 and only then casts it to u16, so through the fake panel only a value that truncates to 0, such as 65536, reaches them.