Deployment and upgrades
This page takes katana from a binary that passes --test to a service you can leave running and upgrade without surprises. It covers which release build to install, where the files go, a systemd template unit that runs one katana process per config file, how to read the logs, an upgrade procedure that restarts one instance before the others, how to roll back, what to test after a restart, and how memory and file descriptors grow with load.
It is written for operators of katana v3.0.1 on x86_64 Linux with systemd. The commands use two example instances, xboard and sspanel, one per panel. If you have not run katana yet, start with the quick start; the general service behaviour that katana shares with etemenanki-app (exit codes, signals, log filters) is on Running in production.
What a production node looks like
Section titled “What a production node looks like”flowchart LR SD["systemd: katana@xboard"] --> P1["katana process"] SD2["systemd: katana@sspanel"] --> P2["katana process"] P1 --> C1["/etc/katana/config-xboard.toml"] P2 --> C2["/etc/katana/config-sspanel.toml"] P1 -- "stdout" --> J["journald"] P2 -- "stdout" --> J P1 -- "poll and report" --> PA["Panel A"] P2 -- "poll and report" --> PB["Panel B"]
- One binary,
/usr/local/bin/katana, shared by every instance. - One config file per instance. Each file may hold several
[[node]]entries. - One systemd template unit,
katana@.service. The instance name picks the file. - Logs go to standard output, and journald keeps them.
- katana writes no files. It needs read access to its config, certificates, geodata, rule list and
[dns] ca_fileif you set one, and nothing else on disk.
Choose a release build
Section titled “Choose a release build”Every katana release tag publishes two x86_64 Linux executables, each with a SHA-256 checksum file. VERSION is the tag, for example v3.0.1.
katana-VERSION-linux-x86_64-gnu |
katana-VERSION-linux-x86_64-musl |
|
|---|---|---|
| Linking | Dynamically linked against the host’s glibc. OpenSSL is compiled in. | Fully static: no program interpreter, no shared libraries. |
| Host requirement | A glibc at least as new as the build host’s, Ubuntu 22.04 (glibc 2.35). | Any x86_64 Linux, including Alpine. |
[dns] backend = "system" |
glibc’s getaddrinfo, which follows /etc/nsswitch.conf and its NSS modules. |
musl’s resolver, which reads /etc/hosts and /etc/resolv.conf directly. |
| Memory allocator | glibc’s malloc |
musl’s malloc |
The release workflow checks both properties it promises: it fails if the -gnu binary loads libssl or libcrypto dynamically, and it fails if the -musl binary has any dynamic dependency at all.
How to pick:
| Host | Build |
|---|---|
| Ubuntu 22.04 or later, Debian 12 or later | -gnu, or -musl if you prefer one binary for every host |
| Debian 11, Ubuntu 20.04, RHEL 9 and its rebuilds (glibc 2.34), Alpine | -musl |
Resolution depends on NSS modules (for example nss-resolve, LDAP or a custom hosts: line) |
-gnu |
| Not sure | -musl |
katana does not install its own memory allocator, so the two builds also differ in which allocator serves every buffer and user table. Memory figures from one build do not predict the other. Pick one build and keep it across upgrades, so that a before-and-after comparison measures the new version and not the new allocator.
To see whether a host can run the -gnu build, run the binary’s --version. A glibc that is too old fails right there with version `GLIBC_2.xx' not found, before any config is involved.
Download and verify
Section titled “Download and verify”Download the binary and its .sha256 file into the same directory, from the release page or with the GitHub CLI. OWNER is the account that hosts the katana repository; you need read access to it. The .sha256 file names the asset by its original file name, so run the check in the download directory, before you rename or move anything:
VERSION=v3.0.1BUILD=musl # or gnugh release download "$VERSION" --repo OWNER/katana \ --pattern "katana-$VERSION-linux-x86_64-$BUILD*"sha256sum -c "katana-$VERSION-linux-x86_64-$BUILD.sha256"katana-v3.0.1-linux-x86_64-musl: OKAnything other than OK means the download is incomplete or damaged: download it again and do not install it. The checksum comes from the same CI job as the binary, so it detects corruption, not a substituted file. The assets are plain executables, not archives. Install walks through the same download for a first installation.
Lay out the files
Section titled “Lay out the files”katana has no built-in file locations. It reads the config from -c and every other file from the path written in the config. This layout keeps everything for all instances under /etc/katana:
Directoryusr/local/bin/
- katana the binary every instance runs
- katana.prev the previous version, kept after an upgrade for rollback
Directoryetc/
Directorykatana/ owner root, group katana, mode 0750
- config-xboard.toml instance
katana@xboard, mode 0640 - config-sspanel.toml instance
katana@sspanel, mode 0640 Directorycerts/
- fullchain.pem certificate chain for TLS and Hysteria 2 nodes
- privkey.pem private key, mode 0640, group katana
- geoip.dat only if a route rule uses
geoip - geosite.dat only if a route rule uses
geosite - rules.txt local audit rules, only if you set
rule_list_path
- config-xboard.toml instance
Directorysystemd/
Directorysystem/
- katana@.service the template unit below
Directorykatana@xboard.service.d / optional per-instance overrides
- override.conf
Write every path in the config as an absolute path. katana opens a relative path such as geoip = "geoip.dat" relative to the process’s working directory, not relative to the config file, and the error does not name the file:
configuration error: No such file or directory (os error 2)A node that points at the layout above looks like this:
[log]level = "info"
[[node]]panel_type = "NewV2board"
[node.api]host = "https://panel.example.com"node_id = 1key = "replace-with-the-panel-key"node_type = "V2ray"rule_list_path = "/etc/katana/rules.txt"
[node.controller]listen_ip = "0.0.0.0"update_periodic = 60
[node.controller.cert]mode = "file"cert_file = "/etc/katana/certs/fullchain.pem"key_file = "/etc/katana/certs/privkey.pem"
[node.route]default = "direct"geoip = "/etc/katana/geoip.dat"geosite = "/etc/katana/geosite.dat"
[[node.route.rule]]outbound = "block"geoip = ["private"]Create the service user and set the ownership once:
sudo useradd --system --no-create-home --shell /usr/sbin/nologin katanasudo install -d -o root -g katana -m 0750 /etc/katana /etc/katana/certssudo chown root:katana /etc/katana/*.toml /etc/katana/certs/*.pemsudo chmod 0640 /etc/katana/*.toml /etc/katana/certs/*.pemThe config files hold the panel keys, so keep them unreadable to other users. katana only reads these files, so neither the directory nor the files need to be writable by katana.
Run katana under systemd
Section titled “Run katana under systemd”The template unit
Section titled “The template unit”Save this as /etc/systemd/system/katana@.service. The text after @ in the instance name selects the config file: katana@xboard runs /etc/katana/config-xboard.toml.
[Unit]Description=katana node agent (%i)After=network-online.targetWants=network-online.target
[Service]Type=execUser=katanaGroup=katanaWorkingDirectory=/etc/katanaEnvironment=NO_COLOR=1# Only for [dns] backend "tls" or "https" without ca_file:# Environment=SSL_CERT_FILE=/etc/ssl/certs/ca-certificates.crtExecStartPre=/usr/local/bin/katana --test -c /etc/katana/config-%i.tomlExecStart=/usr/local/bin/katana -c /etc/katana/config-%i.tomlRestart=on-failureRestartSec=5sLimitNOFILE=1048576
# Only if a panel assigns a node port below 1024.AmbientCapabilities=CAP_NET_BIND_SERVICECapabilityBoundingSet=CAP_NET_BIND_SERVICE
NoNewPrivileges=yesProtectSystem=strictProtectHome=yesPrivateTmp=yesPrivateDevices=yesProtectKernelTunables=yesProtectKernelModules=yesProtectControlGroups=yesRestrictAddressFamilies=AF_INET AF_INET6 AF_UNIX AF_NETLINKRestrictNamespaces=yesLockPersonality=yesSystemCallArchitectures=native
[Install]WantedBy=multi-user.targetThen start the instances and follow their logs:
sudo systemctl daemon-reloadsudo systemctl enable --now katana@xboard katana@sspaneljournalctl -u 'katana@*' -fWhat each setting does
Section titled “What each setting does”| Setting | Why |
|---|---|
After= and Wants=network-online.target |
Each node asks its panel for its settings right after start. A node whose first attempt fails retries by itself, 1 s later at first and then with a wait that doubles up to 60 s, or up to update_periodic if that is shorter. Starting after the network and DNS are up avoids a burst of first-attempt errors in the journal, and a delay of up to a minute: by the time the network is ready, a node’s wait for its next attempt may have grown that long. See Restarts and node failures. |
ExecStartPre=… --test |
Refuses to start a file that does not parse, has an unknown key, names an unknown panel_type, node_type, outbound or geodata category, points at a geodata file it cannot read, has invalid Hysteria 2 settings, or holds no [[node]]. A plain start would skip a node it cannot build and run the rest; the check makes a broken node fail the whole start, where you notice it. --test prints Configuration OK into the journal. |
Restart=on-failure |
Restarts katana after it exits with status 1, or after a signal other than SIGTERM, SIGINT, SIGHUP or SIGPIPE kills it, for example SIGKILL from the out-of-memory killer. It does not restart after systemctl stop or SIGHUP. A failed ExecStartPre= counts as a failure too: with RestartSec=5s, systemd’s default start limit (5 starts in 10 seconds) is never reached, so a broken config is retried every 5 seconds, logging the same error, and the instance comes up on its own once you fix the file. |
LimitNOFILE=1048576 |
Every client connection holds a descriptor, and usually a second one for its outbound. systemd’s default soft limit of 1024 runs out quickly. See File descriptors. |
AmbientCapabilities=CAP_NET_BIND_SERVICE |
Lets the unprivileged katana user bind ports below 1024. See Ports below 1024. |
Environment=NO_COLOR=1 |
katana writes ANSI colour codes even when standard output is not a terminal. NO_COLOR turns them off, so the journal shows plain text instead of sequences like [2m. |
WorkingDirectory=/etc/katana |
A safety net for relative paths in a config. Absolute paths do not need it. |
ProtectSystem=strict and the other hardening lines |
katana only reads files and opens network sockets, so the file system can be read-only for it and the rest of the system hidden. Remove a line if your setup needs what it takes away. |
Leave TimeoutStopSec= at its default of 90 seconds. On SIGTERM each node closes its listener and connections and then sends one last traffic report, plus one audit report on SSPanel. Each report can take up to api.timeout seconds (5 when unset), and the nodes do this in parallel. If systemd has to send SIGKILL first, the traffic of the last poll period is lost.
Ports below 1024
Section titled “Ports below 1024”The panel, not the config file, decides each node’s port. The one exception is a Hysteria 2 node that sets [node.hysteria].port, which listens on that port without asking the panel. A node assigned port 443 needs CAP_NET_BIND_SERVICE. Without it the node logs the error below and keeps retrying, while the process and every other node keep running:
ERROR katana::manager::node: node 1: initial start failed: Permission denied (os error 13); retrying in 1sThe retries cannot succeed on their own. A process receives its capabilities when it starts, so after you grant the right, restart the instance.
There are three ways to grant the right:
| Method | How | Note |
|---|---|---|
| Unit capability | AmbientCapabilities= and CapabilityBoundingSet=, as in the unit above |
Recommended. It belongs to the service, so it survives binary upgrades. |
| File capability | sudo setcap cap_net_bind_service=+ep /usr/local/bin/katana |
Stored on the file. The upgrade procedure below installs a new file, which has no capability, so every low-port node fails after the next restart. Avoid it. |
| System-wide | sysctl net.ipv4.ip_unprivileged_port_start=443 |
Lets every user on the host bind those ports. |
If every node uses a port of 1024 or above, delete the AmbientCapabilities= line and leave CapabilityBoundingSet= empty, so the process holds no capability at all.
Restarts and node failures
Section titled “Restarts and node failures”Restart=on-failure watches the process, not the nodes inside it. katana exits with status 1 only when it cannot load the config, cannot build the outbound pool, has no [[node]], or could not create any node at all. A node that cannot come up retries by itself instead: its panel is unreachable at start, the panel returns port 0 or no user list, or its port is taken. It logs an error for every failed attempt, and systemd still shows the unit as active (running):
ERROR katana::manager::node: node 2: node_info failed: …; retrying in 1sERROR katana::manager::node: node 3: initial start failed: Address already in use (os error 98); retrying in 4sThe first retry comes 1 s after the first failure. Each further wait doubles, up to 60 s, or up to update_periodic if that is shorter. Every attempt asks the panel afresh and builds the listener again, so the node comes up at its next attempt once the cause is gone: the panel answers, the port is free, the certificate is readable. Saving an edit to the node’s [[node]] entry starts the next attempt at once. Restart the instance only for a change that the process picks up at start, such as a unit setting or a capability.
After every start, check that each node logged its listening on line, as described in Smoke test.
Run several instances
Section titled “Run several instances”One katana process can serve many nodes: repeat [[node]] in one file. Split the nodes across instances when you want:
- to restart one group of nodes without dropping the others’ connections, for example one instance per panel;
- to upgrade one group first and keep the rest on the old binary until it has proved itself;
- separate log streams,
journalctl -u katana@xboardandjournalctl -u katana@sspanel.
Instances do not coordinate with each other. Keep these rules:
| Rule | Why |
|---|---|
| Serve each panel node from exactly one instance. | The panel gives the node one port, so the instance that starts second cannot bind it. It logs initial start failed: Address already in use (os error 98); retrying in …s for that node and keeps retrying, so it takes the port at its next attempt after the other instance releases it, for example when that instance stops. |
Keep ports unique across instances whose listen_ip values overlap. 0.0.0.0, the default, overlaps every IPv4 address. |
The panel assigns ports per node, and katana cannot see another instance’s ports. The second bind fails with Address already in use (os error 98), and that node keeps retrying. A Hysteria 2 node listens on UDP, so it can share a number with a TCP node. |
| Give each instance its own WireGuard key, or keep all WireGuard outbounds in one instance. | Each process opens its own tunnel. Two tunnels with the same private key to the same peer take the peer’s session from each other. |
| Expect a little more memory and DNS traffic. | Each process builds its own outbound pool, resolver cache and user tables. |
Every instance watches /etc/katana, because that is the directory of its config file. Saving one instance’s file makes every instance re-read its own file. An instance whose file did not change applies nothing and logs nothing.
For a setting that only one instance needs, use a drop-in instead of editing the template:
sudo systemctl edit katana@xboard[Service]Environment=RUST_LOG=info,katana=debugRestart the instance to apply it: sudo systemctl restart katana@xboard.
katana writes one line per event to standard output, and journald stores it under the unit’s name. Each line has a UTC timestamp, the level, the module and the message:
2026-09-24T08:00:01.512304Z INFO katana::manager::node: node 1: listening on 0.0.0.0:4432026-09-24T08:00:01.733918Z WARN katana::manager::node: node 2: report traffic: …2026-09-24T09:14:52.004816Z INFO katana::runtime: shutting downReading the journal
Section titled “Reading the journal”journalctl -u katana@xboard -f # follow one instancejournalctl -u 'katana@*' --since '10 min ago' # every instancejournalctl -u 'katana@*' -b -g 'WARN|ERROR' # problems since bootjournalctl -u katana@xboard -g 'listening on' -n 50 # which nodes are upChoosing the level
Section titled “Choosing the level”The log filter comes from the first of these that is set:
- the
RUST_LOGenvironment variable, if it parses; [log].levelin the config file;info.
Set RUST_LOG in a drop-in, as shown in Run several instances, when you want a level for one run without touching the config. Change [log].level in the config when you want it to last: katana applies a new [log].level at once, without a restart and without dropping connections, and from then on it replaces the RUST_LOG filter. Useful values:
| Value | What you get |
|---|---|
info |
Startup, listener, reload and shutdown lines, warnings and errors. The default, and the right level for production. |
warn |
Warnings and errors only. |
info,katana=debug |
katana’s own modules at debug, for example connections refused by a limit. |
debug |
Every module at debug, including the kernel’s per-connection errors. Verbose on a busy node. |
At debug, a busy node can exceed journald’s rate limit, and journald then drops lines and logs Suppressed N messages from katana@xboard.service. Raise the limit for the instance while you debug, with LogRateLimitIntervalSec=0 in its drop-in, and remove it afterwards. Running in production explains the filter syntax in full.
Lines worth watching
Section titled “Lines worth watching”| Line | Level | Meaning |
|---|---|---|
node N: listening on IP:PORT |
INFO | The node is serving. There is no such line while the panel lists no users for the node: katana binds nothing until the node has users. |
node N: node_info failed: …, node N: panel returned no node info, node N: panel returned port 0, node N: user_list failed: …, node N: panel returned no user list, node N: initial start failed: …, each ending in ; retrying in Ns |
ERROR | The node is not up yet and retries after N seconds. The wait doubles from 1 s up to 60 s, or up to update_periodic if that is shorter. |
node N: node_info: …, node N: user_list: … |
WARN | A poll could not reach the panel. The node keeps its last settings and users. |
node N: report traffic: … |
WARN | A traffic report failed. katana keeps the bytes and sends them with the next report. |
node N: rebuild failed: … |
ERROR | A running node lost its listener, for example to a missing certificate. It retries at every poll. |
accept error, backing off 100ms: Too many open files (os error 24) |
WARN | The descriptor limit is reached. Raise LimitNOFILE=. |
node TAG: N inbound handshake failures in the last 1s (possible handshake scan/DoS or misconfigured clients) |
WARN | More than 10 failed handshakes in one second on one node. TAG has the form V2ray_0.0.0.0_443. |
cannot read rule_list_path PATH: … |
WARN | The local audit rule list is missing or unreadable by katana. The node runs without the local rules until a later read succeeds. On SSPanel that read happens only at a poll where the panel sends its rule list. |
config reload failed, keeping current: …, reload: bad outbounds, keeping current config: …, reload: node …: …; keeping current config |
ERROR | An edit to the config file was rejected as a whole, for example because one node names an unknown node_type. None of the edit is applied, and every node keeps running as it was. |
node N: config edit refused, keeping the running one: … |
ERROR | This node’s edited entry did not build, for example because a route rule names an unknown outbound tag. The node keeps running with its previous settings. |
shutting down |
INFO | katana received SIGINT or SIGTERM. |
Upgrade
Section titled “Upgrade”A restart of an instance drops every connection on it, and its nodes are unreachable for a few seconds while the old process sends its last reports and the new one fetches its nodes from the panel. Clients reconnect on their own. The procedure below makes sure the new binary accepts every config before anything stops, keeps the old binary for a rollback, and exposes only one instance to the new version until it has passed a smoke test.
flowchart TB
D["Download and verify the checksum"] --> T{"--test passes for every config?"}
T -- no --> X["Stop here, the old binary keeps running"]
T -- yes --> K["Keep the old binary as katana.prev"]
K --> I["Install the new binary by rename"]
I --> C["Restart one instance"]
C --> W{"Nodes listening, client works, traffic reported?"}
W -- no --> R["Roll back"]
W -- yes --> A["Restart the other instances"]
A --> M["Watch memory for a day"]
-
Record a baseline. Note each instance’s memory, and run your smoke test against the old version, so that you have something to compare with:
Terminal window systemctl show -p MemoryCurrent katana@xboard katana@sspanel -
Download and verify the new release, as in Download and verify. Use the same build,
-gnuor-musl, as the running one. -
Stage the binary next to the old one under a temporary name, and check that it runs on this host:
Terminal window sudo install -m 0755 "katana-$VERSION-linux-x86_64-$BUILD" /usr/local/bin/katana.new/usr/local/bin/katana.new --version -
Test every config with the new binary, as the service user. A new version may reject a key the old one accepted, and katana rejects unknown keys:
Terminal window sudo -u katana sh -c 'for f in /etc/katana/config-*.toml; doprintf "%s: " "$f"/usr/local/bin/katana.new --test -c "$f" || echo FAILEDdone'The loop runs as
katanaas a whole, so the file list comes from a directory that only root and thekatanagroup can read. Every file must printConfiguration OK. If one fails, stop here and fix the config first; the running instances are untouched. Running askatanaalso catches files the service user cannot read, such as geodata or a Hysteria 2 node’s private key.--testdoes not read the certificates of other node types or the rule list; the smoke test covers those. -
Keep the current binary for a rollback:
Terminal window sudo cp -p /usr/local/bin/katana /usr/local/bin/katana.prev -
Install the new binary by renaming it over the old one:
Terminal window sudo mv -f /usr/local/bin/katana.new /usr/local/bin/katanaDo not
cponto/usr/local/bin/katana. Linux refuses to open a running executable for writing, so the copy fails withText file busy. A rename only replaces the directory entry: running processes keep the old file, and the next start uses the new one. -
Restart one instance and watch it come up:
Terminal window sudo systemctl restart katana@xboardjournalctl -u katana@xboard -fExpect
Configuration OKfrom the start check, then onenode N: listening on …line for every node that has users, and noERRORlines. -
Smoke-test that instance for at least one poll period,
update_periodicseconds (60 by default), so that the first traffic report has been sent. Follow Smoke test and compare with the baseline. -
Restart the remaining instances, one at a time or together:
Terminal window sudo systemctl restart katana@sspanel -
Keep watching memory and the journal for the next day. See Memory for what to expect.
Roll back
Section titled “Roll back”To go back to the previous version, check that it still accepts the configs, rename it into place and restart the instances that run the new one:
sudo -u katana sh -c ' for f in /etc/katana/config-*.toml; do printf "%s: " "$f" /usr/local/bin/katana.prev --test -c "$f" || echo FAILED done'sudo mv -f /usr/local/bin/katana.prev /usr/local/bin/katanasudo systemctl restart katana@xboardThe --test step matters if you added a key during the upgrade that only the new version knows. The old version rejects the file with unknown field and, with the unit above, refuses to start; remove the key first. A rollback drops the instance’s connections once more, like any restart.
The rename consumes katana.prev. Copy it first (sudo cp -p /usr/local/bin/katana.prev /usr/local/bin/katana.prev2) if you want to keep it around.
Smoke test
Section titled “Smoke test”A passing --test shows that the file is valid. It does not contact the panel, bind a port, read the rule list, or read the certificate of any node other than a Hysteria 2 node. Only a real start shows that a node works.
To read only the lines of the current run, and not those of the process before the restart, filter on the run’s systemd invocation ID:
id=$(systemctl show -p InvocationID --value katana@xboard)journalctl _SYSTEMD_INVOCATION_ID="$id" -g 'listening on|WARN|ERROR'These checks, in order, take a few minutes:
| Check | How | Expected |
|---|---|---|
| Every node is up | The current run’s journal, as shown above | One listening on line per node that has users, with the port the panel assigned, and no ERROR line |
| The ports are bound | sudo ss -ltnp | grep katana (TCP) and sudo ss -lunp | grep katana (Hysteria 2) |
Each node’s port, owned by katana |
| A client connects | Use a real client with a test user’s subscription from the panel, and open an HTTPS site through it | The page loads. For TLS nodes, the client accepts the certificate |
| Routing works | Through the client, visit a destination one of your rules sends to block or to a named outbound |
Blocked, or leaving through the expected exit |
| UDP works | If the node relays UDP, run a DNS lookup or another UDP application through the client | An answer comes back |
| Traffic is reported | Download a file of known size through the client, then wait update_periodic seconds |
The test user’s usage in the panel grows by about that size, times the node’s traffic rate |
| No new warnings | journalctl -u katana@xboard --since '15 min ago' -g 'WARN|ERROR' |
Nothing new compared with the baseline |
A few notes on these checks:
- katana logs nothing when a traffic report succeeds, only when one fails. The panel is where you confirm it.
- Some panels apply reports in a background queue, so the panel’s figure can trail the report by a little. Xboard, for example, hands each report to a queue job, which needs its queue worker running.
- Run the same client tests against the old version before the upgrade. Comparing two runs finds differences faster than judging one run by itself.
- If a TLS node has no
listening online, check its certificate paths and permissions. For every node type except Hysteria 2, the listener build is the first time katana reads them, and a failure shows up asinitial start failed: …; retrying in …s, not in--test. The node reads the files again at every retry, so it comes up once they are in place and readable.
Certificate renewal
Section titled “Certificate renewal”katana reads a node’s certificate and key when it builds the node’s listener, so a renewed certificate is not served until the listener is rebuilt. Nothing in the config file changes at a renewal, so no rebuild follows by itself. Restart the instances that use it from your ACME client’s deploy hook, after the new files are in place and readable by the katana group:
install -m 0640 -o root -g katana fullchain.pem /etc/katana/certs/fullchain.peminstall -m 0640 -o root -g katana privkey.pem /etc/katana/certs/privkey.pemsystemctl restart katana@xboard katana@sspanelThe restart drops the instances’ connections, like any restart. Because of the caution under Upgrade, do not let a renewal run in the middle of an upgrade.
Resources
Section titled “Resources”Memory
Section titled “Memory”Plan memory per node, and per user on each node:
- Users. Each node builds its own table of its users, with each user’s key material and traffic counters. Memory grows with the number of users, multiplied by the number of nodes that serve them. A user who can use ten nodes is held ten times.
- Connections. Each open connection holds its buffers and, for a stream node, a reference to the user table that was current when it connected. When the panel’s user list changes, the node builds a new table for new connections, and the old table stays in memory until the last connection that started under it has closed. On a node with many users and long-lived connections, several generations of the table can be alive at once.
- Instances. Each instance holds its own tables, resolver cache and outbounds.
Memory therefore rises after a start, as clients connect and user lists change, before it levels off. Judge a new version by the level it settles at over hours, not by the first minutes, and compare it with the baseline from the same build.
systemctl status katana@xboard | grep Memorysystemctl show -p MemoryCurrent katana@xboardIf you set a hard limit with MemoryMax= in a drop-in, leave generous headroom above the level you observed. When the kernel kills katana at the limit, every connection on the instance drops, and the traffic counted since the last report is lost, because katana keeps its counters in memory. Restart=on-failure then starts it again.
File descriptors
Section titled “File descriptors”Every accepted TCP connection holds one descriptor, and every flow that reaches an outbound usually holds another. A gRPC connection can carry many streams, each with its own outbound. A direct UDP flow can hold one socket per address family. A Hysteria 2 node shares one UDP socket for all its clients, but its streams open outbound sockets like any other node.
katana does not raise its own descriptor limit, so the limit systemd gives it is the one it has. When it runs out, the node stops accepting: it logs accept error, backing off 100ms: Too many open files (os error 24) and tries again every 100 ms until descriptors come free. LimitNOFILE=1048576 covers several nodes at their connection limits. Check what a running instance has:
pid=$(systemctl show -p MainPID --value katana@xboard)sudo ls /proc/$pid/fd | wc -lgrep 'open files' /proc/$pid/limitsPer-node limits
Section titled “Per-node limits”These limits are built in and cannot be configured. They protect a node, not the host, so size LimitNOFILE= and memory for your real load rather than for these numbers.
| Limit | Value | When it is reached |
|---|---|---|
| Live connections per node (stream nodes) | 65,536 sockets | A new socket is closed at once and counted as a failed handshake. |
| Sockets in their TLS, WebSocket or HTTP/2 handshake per node | 2,048 | The node stops accepting until one finishes or has held its place for 10 s. |
| Streams in their protocol handshake per node | 512 | The new stream is dropped and counted as a failed handshake. |
| Time to finish the protocol handshake (stream nodes) | 10 s | The connection is closed and counted as a failed handshake. |
| Connection with no traffic in either direction (stream nodes) | 360 s | The connection is closed. The protocol’s own idle timeout of 300 s usually closes a quiet flow first. |
| Hysteria 2 QUIC connections per node | 4,096 | A new connection is refused. |
| Hysteria 2 relayed streams and UDP associations per node, across all connections | 65,536 | A new stream is rejected; a new association’s packets are dropped. |
Troubleshooting
Section titled “Troubleshooting”| Symptom | Cause | Fix |
|---|---|---|
The unit is active (running) but a node has no listening on line |
The node’s panel request failed, the panel returned port 0 or no user list, or its bind failed. Or the node has no users yet. |
Read the node’s latest ERROR line, which ends in ; retrying in Ns. Once you fix the cause, the node comes up at its next retry; restart the instance only after a unit or capability change. With no users, the node binds at the first poll after users appear. |
initial start failed: Permission denied (os error 13); retrying in …s |
Port below 1024 without CAP_NET_BIND_SERVICE, often after an upgrade removed a setcap capability. |
Use the unit’s AmbientCapabilities=, then restart the instance. |
initial start failed: Address already in use (os error 98); retrying in …s |
Another instance or program holds the port. --test cannot detect this. |
Keep ports unique across instances. The node takes the port at its next retry after the other holder releases it. |
The unit fails at ExecStartPre with configuration error: No such file or directory (os error 2) |
The config file is missing, or a file it names (geodata or [dns] ca_file) does not exist, often because of a relative path. |
Check the instance name against the file name, and use absolute paths. |
configuration error: Permission denied (os error 13) from the start check, or configuration error: node N: Permission denied (os error 13) for a Hysteria 2 node’s certificate or key |
The katana user cannot read the config or a file it names. |
Group katana, mode 0640. |
version `GLIBC_2.xx' not found |
The -gnu build on a host with an older glibc. |
Use the -musl build. |
cp: cannot create regular file '/usr/local/bin/katana': Text file busy |
Copying over the running binary. | Install by rename, as in the upgrade steps. |
Sequences like [2m in the journal |
Colour codes. | Environment=NO_COLOR=1. |
journalctl -p warning shows nothing |
journald stores katana’s lines as info. |
Filter with -g 'WARN|ERROR'. |
| The service stopped and systemd did not restart it | It received SIGHUP, or was stopped on purpose. |
Remove any ExecReload= that sends SIGHUP. |
accept error, backing off 100ms: Too many open files (os error 24) |
Descriptor limit too low. | Raise LimitNOFILE=. |
More node-level errors, and what the panel’s answers mean, are on Troubleshooting.