Hot reload
katana changes its running state from two sources without a restart. When you save the config file, katana reloads it and applies the difference. When the panel changes a node or its users, katana picks the change up at the next poll. Most changes apply within seconds and some drop the node’s open connections. Only a [dns] edit on its own, and renewed certificate or geodata files, wait for a restart or for a later rebuild.
This page is for operators who edit a live node’s config or who need to predict what a panel change does to connected users. It describes katana v3.0.1. For the config keys themselves, see Config file.
| Source | When katana notices | Where it is decided |
|---|---|---|
| The config file | About 500 ms after the file is written | File changes |
| The panel | At the node’s next poll, every update_periodic seconds (default 60) |
Panel changes |
katana has no reload signal and no reload command. It does not handle SIGHUP, so the signal’s default action applies: it ends the process at once, without the final traffic reports that a SIGTERM or SIGINT shutdown sends.
File changes
Section titled “File changes”How katana notices an edit
Section titled “How katana notices an edit”At startup katana watches the directory that contains the config file, not the file itself. With a relative path such as the default -c config.toml, that is the working directory. Editors often save by writing a new file and renaming it over the old one, and a watch on the directory sees both styles of save. The watch is not recursive.
Any event in that directory starts a reload, including writes to unrelated files such as certificates. Opening a file there counts too, so cat or katana --test on a file in that directory also starts one. katana then:
- waits 500 ms and discards every event that arrived meanwhile, so a burst of writes from one save becomes one reload;
- reads and parses the config file again;
- compares the result with the running config and applies only what differs.
A reload that finds the parsed config unchanged applies nothing, so a save that only changes comments or whitespace is free. A reload that katana refuses leaves the running config as it was, so the next event compares the same file again and refuses it again (see step 2).
katana’s own read of the config file is an event in the watched directory as well. Once the first event has arrived, katana therefore reads and compares the file again about every 500 ms for as long as it runs, whether or not you save anything. For an unchanged file that costs one read and one comparison. It also means that:
- a file that does not parse logs the
config reload failedline below about twice a second until you fix it; - a rejected outbound pool logs
reload: bad outboundsat the same rate; - a file with a node that does not build logs
reload: node <name>: <error>; keeping current configat the same rate.
This does not happen when the config path is a symlink into another directory, because the reads then land in the target’s directory.
If the file does not parse, katana logs the error and keeps the running config. Nothing is applied, not even the parts of the file that are valid:
ERROR katana::runtime: config reload failed, keeping current: config parse error: TOML parse error at line 14, column 1 |14 | listen_ipp = "0.0.0.0" | ^^^^^^^^^^unknown field `listen_ipp`, expected one of `listen_ip`, `send_ip`, `update_periodic`, `disable_upload_traffic`, `disable_get_rule`, `disable_sniffing`, `cert`The same config reload failed, keeping current line, with a different cause, appears when the file is missing or is not UTF-8. katana applies the file as soon as it parses again: saving the fix is itself an event, and the re-reads described above pick it up as well.
What a reload applies, in order
Section titled “What a reload applies, in order”flowchart TB
E["Event in the config directory"] --> W["Wait 500 ms, drop queued events"]
W --> P{"File parses?"}
P -- no --> K["Keep the running config"]
P -- yes --> O{"outbound list changed?"}
O -- yes --> B{"New outbound pool builds?"}
B -- no --> K
B -- yes --> C{"Added nodes and changed panel clients build?"}
O -- no --> C
C -- no --> K
C -- yes --> L["Apply log level"]
L --> S["Hand the new pool to every node"]
S --> R["Stop removed nodes"]
R --> U["Update changed nodes"]
U --> N["Start new nodes"]
-
Outbound pool. If any
[[outbound]]entry changed, katana builds the complete new pool first: the built-indirect,block,freedomandblackholehandlers, every[[outbound]]entry, and the[dns]resolver they share. If the pool fails to build, katana rejects the whole reload and logsreload: bad outbounds, keeping current config: …. Log level, nodes and everything else stay as they were. -
Node check. Before it applies anything, katana builds every node the file adds, including a node whose identity changed: its panel client, and its router compiled against the new pool if step 1 built one. For each running node whose
[[node]]table changed, katana builds the new panel client. If any of these fails, katana rejects the whole reload and logs which node stopped it:ERROR katana::runtime: reload: node newv2board@https://panel.example.com#1/v2rayy: unknown node_type "V2rayy"; keeping current configA new node whose rules do not compile logs
reload: node <name>: build router: …; keeping current config. Nothing is applied: not the log level, not the pool, and no removal, edit or addition. Every node keeps running as it was. -
Log level. If
[log].levelchanged, katana swaps its log filter and logsreload: log level → <level>. Removing the key setsinfo. The new filter replaces the one taken fromRUST_LOGat startup. -
Pool swap. If a new pool was built, katana hands it to every node and logs
reload: outbound pool rebuilt. Each node recompiles its router against the new pool and rebuilds its listener, whether or not its rules use the outbound that changed. -
Removed nodes. Each running node whose identity no longer appears in the file is stopped: its listener closes, every connection on it ends, and it sends one last traffic report (unless
disable_upload_trafficis on) and, on SSPanel, one last audit report to its panel. katana waits for these reports before it continues. Each request is bounded by the panel timeout (api.timeout, default 5 seconds), so a slow panel can delay the rest of the reload by up to twice that. If the final traffic report fails, katana does not retry it. -
Changed nodes. Each node whose identity is unchanged but whose
[[node]]table differs receives its new settings, and katana logsreload: reconfigured node <name>. The node applies them itself; see What each change does. -
New nodes. Each node built in step 2 is started. It contacts its panel, fetches its users and binds its port, and keeps trying until it is up (see A node that cannot come up).
katana builds new nodes before it removes anything, and starts them after the removals. When a save changes a node’s identity, or replaces one node with another that gets the same port, the old node releases the port before the new one binds it.
Node identity
Section titled “Node identity”katana matches each [[node]] in the new file to a running node by its identity, a combination of five values:
| Value | Compared as |
|---|---|
panel_type |
Case-insensitive: NewV2board and newv2board are the same. NewV2board and V2board differ, although both use the same panel client |
api.host |
Exact text: https://panel.example.com and https://panel.example.com/ differ |
api.node_id |
Number |
api.key |
Exact text |
| The node type katana asks the panel for | NewV2board and V2board only: vless when api.node_type is a V2ray value (V2ray, Vmess or Vless) and api.enable_vless = true, otherwise api.node_type in lowercase. SSPanel has no such value |
These values decide which panel node the entry serves. The UniProxy API of NewV2board and V2board finds a node by its ID and the node type it is asked for, so the same ID asked as vless and as vmess is two panel nodes, each with its own users and traffic. SSPanel finds a node by its ID alone.
Changing any of the five values makes katana stop the old node, with its final report sent under the old values, and start a fresh node under the new ones. The fresh node begins with empty traffic counters, so one panel node’s traffic is never reported to another, and it fetches everything from the panel again. On NewV2board with enable_vless off, for example, changing node_type from V2ray to Vmess starts a fresh node because the lowercased value changes, but V2ray to v2ray does not. Setting enable_vless on a Trojan node leaves the identity unchanged.
Everything else in a [[node]] table is a setting of the same node. The order of [[node]] tables in the file does not matter, so you can reorder them freely.
Log lines name a node by panel type, host, node ID and, on NewV2board and V2board, the node type, never by its key: reload: removing node newv2board@https://panel.example.com#1/v2ray, or reload: removing node sspanel@https://panel.example.com#1 on SSPanel. This page writes that name as <name>.
A node that cannot come up
Section titled “A node that cannot come up”A node keeps trying until it is up, whether katana started it or a reload added it. Each attempt forgets the ETags of earlier attempts, fetches the node settings, fetches the user list and binds the port. If a step fails, the node logs an error and waits:
ERROR katana::manager::node: node 1: node_info failed: …; retrying in 1s| Reason in the log line | What failed |
|---|---|
node_info failed: … |
The request for the node settings |
panel returned no node info |
The panel answered 304 Not Modified instead of the node settings |
panel returned port 0 |
The node settings name port 0. SSPanel only: on NewV2board and V2board a port of 0 shows as node_info failed: newV2board: server port must be > 0 |
user_list failed: … |
The request for the user list |
panel returned no user list |
The panel answered 304 Not Modified instead of the user list |
initial start failed: … |
Building or binding the listener, for example because the port is taken or a certificate file is missing |
The first wait is 1 second. It doubles after each failure, up to 60 seconds, and it is never longer than update_periodic (at least 1 second). With the default update_periodic = 60 the waits are 1, 2, 4, 8, 16, 32, 60, 60 … seconds.
While the node waits:
- an edit to its own
[[node]]table is stored, and the next attempt starts at once with the new settings; - a new outbound pool is stored for the next attempt, which does not start early;
SIGINTorSIGTERM, or removing the node from the file, stops the retries.
What each change does
Section titled “What each change does”“Listener rebuild” means katana closes the node’s listener, ends every connection on it, including connections still in their handshake, and binds the port again with the new settings. Traffic counts survive a rebuild, so bytes already relayed are still reported.
| Change in the file | When it takes effect | Open connections |
|---|---|---|
[log].level |
At once | Kept |
Any [[outbound]] added, removed or edited |
At once, after the new pool builds | Dropped on every node |
[dns], with no [[outbound]] change |
Not applied and not checked. It takes effect with the next [[outbound]] edit, or at a restart. A broken [dns] makes that next [[outbound]] edit fail as a whole |
Kept |
[[node]] added |
At once | None yet |
[[node]] removed |
At once, after a final report | Dropped on that node |
An identity value: panel_type other than its letter case, api.host, api.node_id, api.key, and on NewV2board and V2board the node type katana asks for |
At once: the node is stopped and started again | Dropped on that node |
api.enable_vless |
NewV2board and V2board with node_type V2ray or Vmess: an identity change, so the node is stopped and started again. Otherwise, including SSPanel: at once, through a new panel client and a listener rebuild |
Dropped on that node |
api.node_type |
NewV2board and V2board: an identity change when the node type katana asks for changes; a change of letter case only behaves like the next row. SSPanel: at once, through a new panel client and an immediate poll, and the reconcile ladder rebuilds the listener if the protocol changed | NewV2board and V2board: dropped, unless only the letter case changed. SSPanel: kept unless the protocol changed |
api.vless_flow, api.disable_custom_config, api.speed_limit, api.rule_list_path, api.timeout, or the letter case of panel_type |
At once, through a new panel client. A speed_limit change refreshes the user tables in place, and flows already open keep the old rate. A rule_list_path change applies the new audit rules to new flows at once |
Kept, unless the node settings read again change the protocol or transport |
controller.listen_ip |
At once: listener rebuild | Dropped on that node |
Any key in [node.controller.cert] |
At once: listener rebuild, which reads the certificate and key files again | Dropped on that node |
controller.disable_sniffing |
At once: listener rebuild | Dropped on that node |
Any key or rule in [node.route] |
At once: the node compiles the new router, then rebuilds its listener. A route that does not compile is refused, see below | Dropped on that node |
Any key in [node.hysteria]: port, obfs, obfs_password, credential, udp, udp_idle_timeout, [node.hysteria.masquerade] |
At once: listener rebuild. Setting or clearing port switches a Hysteria 2 node between the local and the panel description. obfs and obfs_password count only while port is set; with port = 0 the panel’s obfuscation settings are used |
Dropped on that node |
controller.update_periodic |
At once: the poll timer restarts, and the next poll comes one new period later | Kept |
controller.disable_upload_traffic |
At the next traffic report | Kept |
controller.disable_get_rule |
At the next poll. Turning it on keeps the audit rules already loaded | Kept |
controller.send_ip, api.device_limit |
katana accepts these keys but does not use them. As a key in [node.api], a device_limit edit still builds a new panel client |
Kept |
A node takes its own edit whole or not at all. It first builds what the edit needs: a new panel client if anything in [node.api] changed, or the letter case of panel_type, and a new router if [node.route] changed. On NewV2board the new client keeps the routes that the audit rules are derived from until it reads the node settings itself. The new client holds no ETags, so it reads the node settings and the user list in full. A change to listen_ip, the certificate settings, disable_sniffing, enable_vless, [node.hysteria] or [node.route] also forces a listener rebuild.
When a node gets a new panel client or needs a listener rebuild, and it is already up, it runs one poll cycle at once: it fetches the node settings and users, applies them through the reconcile ladder, refreshes its audit rules, and sends its traffic and audit reports. The regular poll timer is not reset. A node that is still trying to come up stores the edit and uses it at its next attempt. The other keys, such as update_periodic, disable_upload_traffic and disable_get_rule, are only stored and read where they are used.
The route rebuild failed, keeping current line comes from an outbound-pool change instead: a new [[outbound]] list that a node’s unchanged rules no longer compile against, for example because you removed an [[outbound]] that the rules still reference. The node keeps routing with its old rules and the old outbound, and it still rebuilds its listener. katana --test rejects this mistake too.
Renaming an outbound and updating the rules that use it in the same save works. The node receives the new pool first, fails to compile its old rules against it, logs the route rebuild failed line once, and then compiles the new rules when its own settings arrive and logs router rebuilt. Its listener is rebuilt twice.
Files katana reads, not watches
Section titled “Files katana reads, not watches”katana reads some files named in the config when it builds something, and does not notice when their contents change. Saving the config with the same path does not count as a change.
| File | Read again when |
|---|---|
cert_file, key_file |
The node’s listener is rebuilt |
[node.route] geoip, geosite |
The node’s router is compiled, if a rule uses a geoip or geosite category: at node start, on a [node.route] edit, and on any [[outbound]] edit |
[dns].ca_file |
The outbound pool is built: at startup and on an [[outbound]] edit |
api.rule_list_path |
While disable_get_rule is off. newV2board: at every poll. SSPanel: at each poll where the panel sends its rule list; a 304 Not Modified answer skips the re-read. An edit to [node.api] reads it again at once, because the new panel client asks the panel in full |
A renewed certificate is the common case. katana keeps serving the old certificate until the listener is rebuilt, and nothing in the config file has to change for a renewal, so plan for a restart after each renewal. See Renew a certificate.
Panel changes
Section titled “Panel changes”Every update_periodic seconds each node polls its panel for the node’s settings and its user list, then applies the difference. The same poll also refreshes the audit rules and sends the traffic and audit reports.
If the panel request fails, the node logs a warning such as node 1: node_info: … and keeps the last settings or users it received. A failed poll never tears anything down. If the panel returns port 0, the node logs node 1: refreshed port is 0, keeping the last one. It keeps the last node settings it received and still applies the user list from that poll.
The reconcile ladder
Section titled “The reconcile ladder”The node takes the first step that matches, which is the step that drops the fewest connections while still applying the change:
flowchart TB
P["Poll: node settings and users"] --> Z{"User list empty?"}
Z -- yes --> T["Close the listener"]
Z -- no --> X{"No listener, or transport or protocol changed?"}
X -- yes --> RB["Listener rebuild"]
X -- no --> US{"Users or node speed limit changed?"}
US -- yes --> RF["Refresh user tables in place"]
US -- no --> NO["Nothing to do"]
| Panel change | Effect | Open connections |
|---|---|---|
| The user list becomes empty | The listener closes and the port is released | All end |
| Users come back after an empty list | The listener is built again | None to keep |
| A transport setting: port, network, host, path, service name, authority, TLS on or off, headers, REALITY, PROXY protocol, Hysteria obfuscation type or password | Listener rebuild | All end |
| A protocol setting: node type, VLESS on or off, VLESS flow, cipher, server key | Listener rebuild | All end |
| Users added, removed or edited, or the node speed limit changed | User tables replaced in place | Kept, except for removed users |
| Nothing | Nothing | Kept |
A listener that failed to bind at an earlier rebuild counts as “no listener”, so the next poll tries again. A node that lost its listener because a certificate file was missing, or because its port was taken, recovers on its own at the first poll after you fix the cause.
What a user refresh keeps
Section titled “What a user refresh keeps”A refresh builds the new user tables next to the old ones and switches over only when they are complete. If the new tables cannot be built, the node logs proxy refresh build failed, keeping current: … and keeps serving the old users. It still records the new list as applied, so it does not try again until the panel’s user list or the node’s settings change once more.
When the switch succeeds:
- Unchanged users keep their connections, and their traffic counts continue.
- Removed users are retired. Their open connections end, and new connections with their credentials are refused. A credential that moved to a different user counts as removed from the old user.
- Users whose speed limit changed, directly or through the node speed limit, keep their connections. Flows that were already open keep the old rate. Flows opened after the refresh run at the new rate.
- New users can connect as soon as the switch is done.
Any change to a user’s record counts, including its password, UUID or speed limit. On a Hysteria 2 node the refresh replaces only the authenticator: the UDP socket stays bound, and the QUIC connections of the users who remain stay up.
For how the speed limits combine, see Speed limits. For what happens to the traffic of a removed user, see Traffic reporting.
Operating a live node
Section titled “Operating a live node”Check an edit before it goes live
Section titled “Check an edit before it goes live”A reload checks less than katana --test does. It rejects a file that does not parse, an outbound pool that does not build, and a file whose added nodes, or the panel clients of its changed nodes, do not build. A running node refuses a [node.route] edit that does not compile. Bad Hysteria settings are still applied: the listener rebuild fails with rebuild failed, and the node tries again at every poll. A node that cannot reach its panel or bind its port keeps retrying on its own. Check the file first, then put it in place in one step:
-
Copy the live config next to it and edit the copy:
Terminal window cp /etc/katana/config.toml /etc/katana/config.toml.newThe copy lives in the watched directory, so each save of it starts a reload of
config.toml. That reload finds nothing changed and does nothing. -
Check the copy:
Terminal window katana --test -c /etc/katana/config.toml.newA valid file prints
Configuration OKand exits with status0. Anything else printsconfiguration error: …and exits with status1. -
Move the copy over the live file:
Terminal window mv /etc/katana/config.toml.new /etc/katana/config.tomlA rename within one directory replaces the file in a single step, so katana never reads a half-written config.
-
Follow katana’s log and look for the
reload:lines. With the template unit from Deployment, for example, runjournalctl -u katana@xboard -f.
--test builds the outbound pool, each node’s panel client and router, and each Hysteria 2 node’s own settings. It does not contact the panel, bind a port, or read the certificate files of nodes whose TLS the panel controls, so a wrong cert_file path on such a node passes the check and fails at the listener rebuild. CLI lists the flags.
Save related edits together. Each save that touches [[outbound]], an identity value, enable_vless, listen_ip, the certificate settings, disable_sniffing, [node.hysteria] or [node.route] drops connections again.
Apply a [dns] change
Section titled “Apply a [dns] change”katana builds the [dns] resolver together with the outbound pool, so a [dns] edit on its own is not applied until the next [[outbound]] edit. To apply it now, save the file and then restart katana. A restart stops every node, sends each node’s final report, and starts again from the file. Deployment covers running katana as a service.
Renew a certificate
Section titled “Renew a certificate”katana reads the certificate and key when it builds a listener, so a renewed certificate is not served until the next rebuild. The dependable way is to restart katana after each renewal, for example from your ACME client’s deploy hook. Pointing cert_file and key_file at new paths also works, because changing them rebuilds the listener, but it drops the node’s connections all the same.
Log lines
Section titled “Log lines”In the reload: lines, <name> is <panel type>@<host>#<id>, followed by /<node type> on NewV2board and V2board.
| Line | Meaning |
|---|---|
config reload failed, keeping current: … |
The file did not parse or could not be read. Nothing was applied |
reload: bad outbounds, keeping current config: … |
The new outbound pool did not build. Nothing was applied |
reload: node <name>: <error>; keeping current config |
A node the file adds, or the new panel client of a changed node, did not build. Nothing was applied |
reload: log level → <level> |
[log].level changed. katana logs this line even when the new value is rejected with invalid log level |
reload: outbound pool rebuilt |
Every node received the new pool and rebuilds its router and listener |
reload: removing node <name> |
A node was removed, or its identity changed |
reload: added node <name> |
A new node, or a node under a new identity, was built and started. It still has to reach its panel and bind its port, and retries until it does |
reload: reconfigured node <name> |
A node received new settings. Check the table above for what they do |
node <id>: config edit refused, keeping the running one: … |
The node’s new router, or its new panel client, did not build. Nothing from that edit was applied |
node <id>: router rebuilt |
The node compiled its routing rules |
node <id>: route rebuild failed, keeping current: … |
After an outbound-pool change, the node’s rules did not compile against the new pool. The node routes with its previous rules and still rebuilds its listener |
node <id>: <reason>; retrying in <N>s |
An attempt to bring the node up failed. See A node that cannot come up |
node <id>: listening on <ip>:<port> |
A listener is up after a start or a rebuild |
node <id>: rebuild failed: … |
The new listener could not be built or bound. The node has no listener until the next poll |
node <id>: refreshed port is 0, keeping the last one |
The panel sent port 0. The node keeps its last node settings and applies the user list from that poll |
node <id>: build router: … |
At startup only: a node was not started because its rules did not compile. On a reload the same error appears in the reload: node <name>: …; keeping current config line |
node <id>: unknown panel_type "<value>" |
At startup only: a node was not started because its panel client could not be built. An unknown api.node_type logs node <id>: unknown node_type "<value>". On a reload these errors appear in the reload: node <name>: …; keeping current config line |
invalid log level "<level>": … |
The new [log].level is not a valid filter. The previous filter stays |
The lines from the reload and from the nodes it updates run concurrently, so their order in the log can vary.