Skip to content

Deployment and upgrades

This page takes katana from a binary that passes --test to a service you can leave running and upgrade without surprises. It covers which release build to install, where the files go, a systemd template unit that runs one katana process per config file, how to read the logs, an upgrade procedure that restarts one instance before the others, how to roll back, what to test after a restart, and how memory and file descriptors grow with load.

It is written for operators of katana v3.0.1 on x86_64 Linux with systemd. The commands use two example instances, xboard and sspanel, one per panel. If you have not run katana yet, start with the quick start; the general service behaviour that katana shares with etemenanki-app (exit codes, signals, log filters) is on Running in production.

flowchart LR
  SD["systemd: katana@xboard"] --> P1["katana process"]
  SD2["systemd: katana@sspanel"] --> P2["katana process"]
  P1 --> C1["/etc/katana/config-xboard.toml"]
  P2 --> C2["/etc/katana/config-sspanel.toml"]
  P1 -- "stdout" --> J["journald"]
  P2 -- "stdout" --> J
  P1 -- "poll and report" --> PA["Panel A"]
  P2 -- "poll and report" --> PB["Panel B"]
  • One binary, /usr/local/bin/katana, shared by every instance.
  • One config file per instance. Each file may hold several [[node]] entries.
  • One systemd template unit, katana@.service. The instance name picks the file.
  • Logs go to standard output, and journald keeps them.
  • katana writes no files. It needs read access to its config, certificates, geodata, rule list and [dns] ca_file if you set one, and nothing else on disk.

Every katana release tag publishes two x86_64 Linux executables, each with a SHA-256 checksum file. VERSION is the tag, for example v3.0.1.

katana-VERSION-linux-x86_64-gnu katana-VERSION-linux-x86_64-musl
Linking Dynamically linked against the host’s glibc. OpenSSL is compiled in. Fully static: no program interpreter, no shared libraries.
Host requirement A glibc at least as new as the build host’s, Ubuntu 22.04 (glibc 2.35). Any x86_64 Linux, including Alpine.
[dns] backend = "system" glibc’s getaddrinfo, which follows /etc/nsswitch.conf and its NSS modules. musl’s resolver, which reads /etc/hosts and /etc/resolv.conf directly.
Memory allocator glibc’s malloc musl’s malloc

The release workflow checks both properties it promises: it fails if the -gnu binary loads libssl or libcrypto dynamically, and it fails if the -musl binary has any dynamic dependency at all.

How to pick:

Host Build
Ubuntu 22.04 or later, Debian 12 or later -gnu, or -musl if you prefer one binary for every host
Debian 11, Ubuntu 20.04, RHEL 9 and its rebuilds (glibc 2.34), Alpine -musl
Resolution depends on NSS modules (for example nss-resolve, LDAP or a custom hosts: line) -gnu
Not sure -musl

katana does not install its own memory allocator, so the two builds also differ in which allocator serves every buffer and user table. Memory figures from one build do not predict the other. Pick one build and keep it across upgrades, so that a before-and-after comparison measures the new version and not the new allocator.

To see whether a host can run the -gnu build, run the binary’s --version. A glibc that is too old fails right there with version `GLIBC_2.xx' not found, before any config is involved.

Download the binary and its .sha256 file into the same directory, from the release page or with the GitHub CLI. OWNER is the account that hosts the katana repository; you need read access to it. The .sha256 file names the asset by its original file name, so run the check in the download directory, before you rename or move anything:

Terminal window
VERSION=v3.0.1
BUILD=musl # or gnu
gh release download "$VERSION" --repo OWNER/katana \
--pattern "katana-$VERSION-linux-x86_64-$BUILD*"
sha256sum -c "katana-$VERSION-linux-x86_64-$BUILD.sha256"
katana-v3.0.1-linux-x86_64-musl: OK

Anything other than OK means the download is incomplete or damaged: download it again and do not install it. The checksum comes from the same CI job as the binary, so it detects corruption, not a substituted file. The assets are plain executables, not archives. Install walks through the same download for a first installation.

katana has no built-in file locations. It reads the config from -c and every other file from the path written in the config. This layout keeps everything for all instances under /etc/katana:

  • Directoryusr/local/bin/
    • katana the binary every instance runs
    • katana.prev the previous version, kept after an upgrade for rollback
  • Directoryetc/
    • Directorykatana/ owner root, group katana, mode 0750
      • config-xboard.toml instance katana@xboard, mode 0640
      • config-sspanel.toml instance katana@sspanel, mode 0640
      • Directorycerts/
        • fullchain.pem certificate chain for TLS and Hysteria 2 nodes
        • privkey.pem private key, mode 0640, group katana
      • geoip.dat only if a route rule uses geoip
      • geosite.dat only if a route rule uses geosite
      • rules.txt local audit rules, only if you set rule_list_path
    • Directorysystemd/

Write every path in the config as an absolute path. katana opens a relative path such as geoip = "geoip.dat" relative to the process’s working directory, not relative to the config file, and the error does not name the file:

configuration error: No such file or directory (os error 2)

A node that points at the layout above looks like this:

/etc/katana/config-xboard.toml
[log]
level = "info"
[[node]]
panel_type = "NewV2board"
[node.api]
host = "https://panel.example.com"
node_id = 1
key = "replace-with-the-panel-key"
node_type = "V2ray"
rule_list_path = "/etc/katana/rules.txt"
[node.controller]
listen_ip = "0.0.0.0"
update_periodic = 60
[node.controller.cert]
mode = "file"
cert_file = "/etc/katana/certs/fullchain.pem"
key_file = "/etc/katana/certs/privkey.pem"
[node.route]
default = "direct"
geoip = "/etc/katana/geoip.dat"
geosite = "/etc/katana/geosite.dat"
[[node.route.rule]]
outbound = "block"
geoip = ["private"]

Create the service user and set the ownership once:

Terminal window
sudo useradd --system --no-create-home --shell /usr/sbin/nologin katana
sudo install -d -o root -g katana -m 0750 /etc/katana /etc/katana/certs
sudo chown root:katana /etc/katana/*.toml /etc/katana/certs/*.pem
sudo chmod 0640 /etc/katana/*.toml /etc/katana/certs/*.pem

The config files hold the panel keys, so keep them unreadable to other users. katana only reads these files, so neither the directory nor the files need to be writable by katana.

Save this as /etc/systemd/system/katana@.service. The text after @ in the instance name selects the config file: katana@xboard runs /etc/katana/config-xboard.toml.

/etc/systemd/system/katana@.service
[Unit]
Description=katana node agent (%i)
After=network-online.target
Wants=network-online.target
[Service]
Type=exec
User=katana
Group=katana
WorkingDirectory=/etc/katana
Environment=NO_COLOR=1
# Only for [dns] backend "tls" or "https" without ca_file:
# Environment=SSL_CERT_FILE=/etc/ssl/certs/ca-certificates.crt
ExecStartPre=/usr/local/bin/katana --test -c /etc/katana/config-%i.toml
ExecStart=/usr/local/bin/katana -c /etc/katana/config-%i.toml
Restart=on-failure
RestartSec=5s
LimitNOFILE=1048576
# Only if a panel assigns a node port below 1024.
AmbientCapabilities=CAP_NET_BIND_SERVICE
CapabilityBoundingSet=CAP_NET_BIND_SERVICE
NoNewPrivileges=yes
ProtectSystem=strict
ProtectHome=yes
PrivateTmp=yes
PrivateDevices=yes
ProtectKernelTunables=yes
ProtectKernelModules=yes
ProtectControlGroups=yes
RestrictAddressFamilies=AF_INET AF_INET6 AF_UNIX AF_NETLINK
RestrictNamespaces=yes
LockPersonality=yes
SystemCallArchitectures=native
[Install]
WantedBy=multi-user.target

Then start the instances and follow their logs:

Terminal window
sudo systemctl daemon-reload
sudo systemctl enable --now katana@xboard katana@sspanel
journalctl -u 'katana@*' -f
Setting Why
After= and Wants=network-online.target Each node asks its panel for its settings right after start. A node whose first attempt fails retries by itself, 1 s later at first and then with a wait that doubles up to 60 s, or up to update_periodic if that is shorter. Starting after the network and DNS are up avoids a burst of first-attempt errors in the journal, and a delay of up to a minute: by the time the network is ready, a node’s wait for its next attempt may have grown that long. See Restarts and node failures.
ExecStartPre=… --test Refuses to start a file that does not parse, has an unknown key, names an unknown panel_type, node_type, outbound or geodata category, points at a geodata file it cannot read, has invalid Hysteria 2 settings, or holds no [[node]]. A plain start would skip a node it cannot build and run the rest; the check makes a broken node fail the whole start, where you notice it. --test prints Configuration OK into the journal.
Restart=on-failure Restarts katana after it exits with status 1, or after a signal other than SIGTERM, SIGINT, SIGHUP or SIGPIPE kills it, for example SIGKILL from the out-of-memory killer. It does not restart after systemctl stop or SIGHUP. A failed ExecStartPre= counts as a failure too: with RestartSec=5s, systemd’s default start limit (5 starts in 10 seconds) is never reached, so a broken config is retried every 5 seconds, logging the same error, and the instance comes up on its own once you fix the file.
LimitNOFILE=1048576 Every client connection holds a descriptor, and usually a second one for its outbound. systemd’s default soft limit of 1024 runs out quickly. See File descriptors.
AmbientCapabilities=CAP_NET_BIND_SERVICE Lets the unprivileged katana user bind ports below 1024. See Ports below 1024.
Environment=NO_COLOR=1 katana writes ANSI colour codes even when standard output is not a terminal. NO_COLOR turns them off, so the journal shows plain text instead of sequences like [2m.
WorkingDirectory=/etc/katana A safety net for relative paths in a config. Absolute paths do not need it.
ProtectSystem=strict and the other hardening lines katana only reads files and opens network sockets, so the file system can be read-only for it and the rest of the system hidden. Remove a line if your setup needs what it takes away.

Leave TimeoutStopSec= at its default of 90 seconds. On SIGTERM each node closes its listener and connections and then sends one last traffic report, plus one audit report on SSPanel. Each report can take up to api.timeout seconds (5 when unset), and the nodes do this in parallel. If systemd has to send SIGKILL first, the traffic of the last poll period is lost.

The panel, not the config file, decides each node’s port. The one exception is a Hysteria 2 node that sets [node.hysteria].port, which listens on that port without asking the panel. A node assigned port 443 needs CAP_NET_BIND_SERVICE. Without it the node logs the error below and keeps retrying, while the process and every other node keep running:

ERROR katana::manager::node: node 1: initial start failed: Permission denied (os error 13); retrying in 1s

The retries cannot succeed on their own. A process receives its capabilities when it starts, so after you grant the right, restart the instance.

There are three ways to grant the right:

Method How Note
Unit capability AmbientCapabilities= and CapabilityBoundingSet=, as in the unit above Recommended. It belongs to the service, so it survives binary upgrades.
File capability sudo setcap cap_net_bind_service=+ep /usr/local/bin/katana Stored on the file. The upgrade procedure below installs a new file, which has no capability, so every low-port node fails after the next restart. Avoid it.
System-wide sysctl net.ipv4.ip_unprivileged_port_start=443 Lets every user on the host bind those ports.

If every node uses a port of 1024 or above, delete the AmbientCapabilities= line and leave CapabilityBoundingSet= empty, so the process holds no capability at all.

Restart=on-failure watches the process, not the nodes inside it. katana exits with status 1 only when it cannot load the config, cannot build the outbound pool, has no [[node]], or could not create any node at all. A node that cannot come up retries by itself instead: its panel is unreachable at start, the panel returns port 0 or no user list, or its port is taken. It logs an error for every failed attempt, and systemd still shows the unit as active (running):

ERROR katana::manager::node: node 2: node_info failed: …; retrying in 1s
ERROR katana::manager::node: node 3: initial start failed: Address already in use (os error 98); retrying in 4s

The first retry comes 1 s after the first failure. Each further wait doubles, up to 60 s, or up to update_periodic if that is shorter. Every attempt asks the panel afresh and builds the listener again, so the node comes up at its next attempt once the cause is gone: the panel answers, the port is free, the certificate is readable. Saving an edit to the node’s [[node]] entry starts the next attempt at once. Restart the instance only for a change that the process picks up at start, such as a unit setting or a capability.

After every start, check that each node logged its listening on line, as described in Smoke test.

One katana process can serve many nodes: repeat [[node]] in one file. Split the nodes across instances when you want:

  • to restart one group of nodes without dropping the others’ connections, for example one instance per panel;
  • to upgrade one group first and keep the rest on the old binary until it has proved itself;
  • separate log streams, journalctl -u katana@xboard and journalctl -u katana@sspanel.

Instances do not coordinate with each other. Keep these rules:

Rule Why
Serve each panel node from exactly one instance. The panel gives the node one port, so the instance that starts second cannot bind it. It logs initial start failed: Address already in use (os error 98); retrying in …s for that node and keeps retrying, so it takes the port at its next attempt after the other instance releases it, for example when that instance stops.
Keep ports unique across instances whose listen_ip values overlap. 0.0.0.0, the default, overlaps every IPv4 address. The panel assigns ports per node, and katana cannot see another instance’s ports. The second bind fails with Address already in use (os error 98), and that node keeps retrying. A Hysteria 2 node listens on UDP, so it can share a number with a TCP node.
Give each instance its own WireGuard key, or keep all WireGuard outbounds in one instance. Each process opens its own tunnel. Two tunnels with the same private key to the same peer take the peer’s session from each other.
Expect a little more memory and DNS traffic. Each process builds its own outbound pool, resolver cache and user tables.

Every instance watches /etc/katana, because that is the directory of its config file. Saving one instance’s file makes every instance re-read its own file. An instance whose file did not change applies nothing and logs nothing.

For a setting that only one instance needs, use a drop-in instead of editing the template:

Terminal window
sudo systemctl edit katana@xboard
/etc/systemd/system/katana@xboard.service.d/override.conf
[Service]
Environment=RUST_LOG=info,katana=debug

Restart the instance to apply it: sudo systemctl restart katana@xboard.

katana writes one line per event to standard output, and journald stores it under the unit’s name. Each line has a UTC timestamp, the level, the module and the message:

2026-09-24T08:00:01.512304Z INFO katana::manager::node: node 1: listening on 0.0.0.0:443
2026-09-24T08:00:01.733918Z WARN katana::manager::node: node 2: report traffic: …
2026-09-24T09:14:52.004816Z INFO katana::runtime: shutting down
Terminal window
journalctl -u katana@xboard -f # follow one instance
journalctl -u 'katana@*' --since '10 min ago' # every instance
journalctl -u 'katana@*' -b -g 'WARN|ERROR' # problems since boot
journalctl -u katana@xboard -g 'listening on' -n 50 # which nodes are up

The log filter comes from the first of these that is set:

  1. the RUST_LOG environment variable, if it parses;
  2. [log].level in the config file;
  3. info.

Set RUST_LOG in a drop-in, as shown in Run several instances, when you want a level for one run without touching the config. Change [log].level in the config when you want it to last: katana applies a new [log].level at once, without a restart and without dropping connections, and from then on it replaces the RUST_LOG filter. Useful values:

Value What you get
info Startup, listener, reload and shutdown lines, warnings and errors. The default, and the right level for production.
warn Warnings and errors only.
info,katana=debug katana’s own modules at debug, for example connections refused by a limit.
debug Every module at debug, including the kernel’s per-connection errors. Verbose on a busy node.

At debug, a busy node can exceed journald’s rate limit, and journald then drops lines and logs Suppressed N messages from katana@xboard.service. Raise the limit for the instance while you debug, with LogRateLimitIntervalSec=0 in its drop-in, and remove it afterwards. Running in production explains the filter syntax in full.

Line Level Meaning
node N: listening on IP:PORT INFO The node is serving. There is no such line while the panel lists no users for the node: katana binds nothing until the node has users.
node N: node_info failed: …, node N: panel returned no node info, node N: panel returned port 0, node N: user_list failed: …, node N: panel returned no user list, node N: initial start failed: …, each ending in ; retrying in Ns ERROR The node is not up yet and retries after N seconds. The wait doubles from 1 s up to 60 s, or up to update_periodic if that is shorter.
node N: node_info: …, node N: user_list: … WARN A poll could not reach the panel. The node keeps its last settings and users.
node N: report traffic: … WARN A traffic report failed. katana keeps the bytes and sends them with the next report.
node N: rebuild failed: … ERROR A running node lost its listener, for example to a missing certificate. It retries at every poll.
accept error, backing off 100ms: Too many open files (os error 24) WARN The descriptor limit is reached. Raise LimitNOFILE=.
node TAG: N inbound handshake failures in the last 1s (possible handshake scan/DoS or misconfigured clients) WARN More than 10 failed handshakes in one second on one node. TAG has the form V2ray_0.0.0.0_443.
cannot read rule_list_path PATH: … WARN The local audit rule list is missing or unreadable by katana. The node runs without the local rules until a later read succeeds. On SSPanel that read happens only at a poll where the panel sends its rule list.
config reload failed, keeping current: …, reload: bad outbounds, keeping current config: …, reload: node …: …; keeping current config ERROR An edit to the config file was rejected as a whole, for example because one node names an unknown node_type. None of the edit is applied, and every node keeps running as it was.
node N: config edit refused, keeping the running one: … ERROR This node’s edited entry did not build, for example because a route rule names an unknown outbound tag. The node keeps running with its previous settings.
shutting down INFO katana received SIGINT or SIGTERM.

A restart of an instance drops every connection on it, and its nodes are unreachable for a few seconds while the old process sends its last reports and the new one fetches its nodes from the panel. Clients reconnect on their own. The procedure below makes sure the new binary accepts every config before anything stops, keeps the old binary for a rollback, and exposes only one instance to the new version until it has passed a smoke test.

flowchart TB
  D["Download and verify the checksum"] --> T{"--test passes for every config?"}
  T -- no --> X["Stop here, the old binary keeps running"]
  T -- yes --> K["Keep the old binary as katana.prev"]
  K --> I["Install the new binary by rename"]
  I --> C["Restart one instance"]
  C --> W{"Nodes listening, client works, traffic reported?"}
  W -- no --> R["Roll back"]
  W -- yes --> A["Restart the other instances"]
  A --> M["Watch memory for a day"]
  1. Record a baseline. Note each instance’s memory, and run your smoke test against the old version, so that you have something to compare with:

    Terminal window
    systemctl show -p MemoryCurrent katana@xboard katana@sspanel
  2. Download and verify the new release, as in Download and verify. Use the same build, -gnu or -musl, as the running one.

  3. Stage the binary next to the old one under a temporary name, and check that it runs on this host:

    Terminal window
    sudo install -m 0755 "katana-$VERSION-linux-x86_64-$BUILD" /usr/local/bin/katana.new
    /usr/local/bin/katana.new --version
  4. Test every config with the new binary, as the service user. A new version may reject a key the old one accepted, and katana rejects unknown keys:

    Terminal window
    sudo -u katana sh -c '
    for f in /etc/katana/config-*.toml; do
    printf "%s: " "$f"
    /usr/local/bin/katana.new --test -c "$f" || echo FAILED
    done'

    The loop runs as katana as a whole, so the file list comes from a directory that only root and the katana group can read. Every file must print Configuration OK. If one fails, stop here and fix the config first; the running instances are untouched. Running as katana also catches files the service user cannot read, such as geodata or a Hysteria 2 node’s private key. --test does not read the certificates of other node types or the rule list; the smoke test covers those.

  5. Keep the current binary for a rollback:

    Terminal window
    sudo cp -p /usr/local/bin/katana /usr/local/bin/katana.prev
  6. Install the new binary by renaming it over the old one:

    Terminal window
    sudo mv -f /usr/local/bin/katana.new /usr/local/bin/katana

    Do not cp onto /usr/local/bin/katana. Linux refuses to open a running executable for writing, so the copy fails with Text file busy. A rename only replaces the directory entry: running processes keep the old file, and the next start uses the new one.

  7. Restart one instance and watch it come up:

    Terminal window
    sudo systemctl restart katana@xboard
    journalctl -u katana@xboard -f

    Expect Configuration OK from the start check, then one node N: listening on … line for every node that has users, and no ERROR lines.

  8. Smoke-test that instance for at least one poll period, update_periodic seconds (60 by default), so that the first traffic report has been sent. Follow Smoke test and compare with the baseline.

  9. Restart the remaining instances, one at a time or together:

    Terminal window
    sudo systemctl restart katana@sspanel
  10. Keep watching memory and the journal for the next day. See Memory for what to expect.

To go back to the previous version, check that it still accepts the configs, rename it into place and restart the instances that run the new one:

Terminal window
sudo -u katana sh -c '
for f in /etc/katana/config-*.toml; do
printf "%s: " "$f"
/usr/local/bin/katana.prev --test -c "$f" || echo FAILED
done'
sudo mv -f /usr/local/bin/katana.prev /usr/local/bin/katana
sudo systemctl restart katana@xboard

The --test step matters if you added a key during the upgrade that only the new version knows. The old version rejects the file with unknown field and, with the unit above, refuses to start; remove the key first. A rollback drops the instance’s connections once more, like any restart.

The rename consumes katana.prev. Copy it first (sudo cp -p /usr/local/bin/katana.prev /usr/local/bin/katana.prev2) if you want to keep it around.

A passing --test shows that the file is valid. It does not contact the panel, bind a port, read the rule list, or read the certificate of any node other than a Hysteria 2 node. Only a real start shows that a node works.

To read only the lines of the current run, and not those of the process before the restart, filter on the run’s systemd invocation ID:

Terminal window
id=$(systemctl show -p InvocationID --value katana@xboard)
journalctl _SYSTEMD_INVOCATION_ID="$id" -g 'listening on|WARN|ERROR'

These checks, in order, take a few minutes:

Check How Expected
Every node is up The current run’s journal, as shown above One listening on line per node that has users, with the port the panel assigned, and no ERROR line
The ports are bound sudo ss -ltnp | grep katana (TCP) and sudo ss -lunp | grep katana (Hysteria 2) Each node’s port, owned by katana
A client connects Use a real client with a test user’s subscription from the panel, and open an HTTPS site through it The page loads. For TLS nodes, the client accepts the certificate
Routing works Through the client, visit a destination one of your rules sends to block or to a named outbound Blocked, or leaving through the expected exit
UDP works If the node relays UDP, run a DNS lookup or another UDP application through the client An answer comes back
Traffic is reported Download a file of known size through the client, then wait update_periodic seconds The test user’s usage in the panel grows by about that size, times the node’s traffic rate
No new warnings journalctl -u katana@xboard --since '15 min ago' -g 'WARN|ERROR' Nothing new compared with the baseline

A few notes on these checks:

  • katana logs nothing when a traffic report succeeds, only when one fails. The panel is where you confirm it.
  • Some panels apply reports in a background queue, so the panel’s figure can trail the report by a little. Xboard, for example, hands each report to a queue job, which needs its queue worker running.
  • Run the same client tests against the old version before the upgrade. Comparing two runs finds differences faster than judging one run by itself.
  • If a TLS node has no listening on line, check its certificate paths and permissions. For every node type except Hysteria 2, the listener build is the first time katana reads them, and a failure shows up as initial start failed: …; retrying in …s, not in --test. The node reads the files again at every retry, so it comes up once they are in place and readable.

katana reads a node’s certificate and key when it builds the node’s listener, so a renewed certificate is not served until the listener is rebuilt. Nothing in the config file changes at a renewal, so no rebuild follows by itself. Restart the instances that use it from your ACME client’s deploy hook, after the new files are in place and readable by the katana group:

deploy hook
install -m 0640 -o root -g katana fullchain.pem /etc/katana/certs/fullchain.pem
install -m 0640 -o root -g katana privkey.pem /etc/katana/certs/privkey.pem
systemctl restart katana@xboard katana@sspanel

The restart drops the instances’ connections, like any restart. Because of the caution under Upgrade, do not let a renewal run in the middle of an upgrade.

Plan memory per node, and per user on each node:

  • Users. Each node builds its own table of its users, with each user’s key material and traffic counters. Memory grows with the number of users, multiplied by the number of nodes that serve them. A user who can use ten nodes is held ten times.
  • Connections. Each open connection holds its buffers and, for a stream node, a reference to the user table that was current when it connected. When the panel’s user list changes, the node builds a new table for new connections, and the old table stays in memory until the last connection that started under it has closed. On a node with many users and long-lived connections, several generations of the table can be alive at once.
  • Instances. Each instance holds its own tables, resolver cache and outbounds.

Memory therefore rises after a start, as clients connect and user lists change, before it levels off. Judge a new version by the level it settles at over hours, not by the first minutes, and compare it with the baseline from the same build.

Terminal window
systemctl status katana@xboard | grep Memory
systemctl show -p MemoryCurrent katana@xboard

If you set a hard limit with MemoryMax= in a drop-in, leave generous headroom above the level you observed. When the kernel kills katana at the limit, every connection on the instance drops, and the traffic counted since the last report is lost, because katana keeps its counters in memory. Restart=on-failure then starts it again.

Every accepted TCP connection holds one descriptor, and every flow that reaches an outbound usually holds another. A gRPC connection can carry many streams, each with its own outbound. A direct UDP flow can hold one socket per address family. A Hysteria 2 node shares one UDP socket for all its clients, but its streams open outbound sockets like any other node.

katana does not raise its own descriptor limit, so the limit systemd gives it is the one it has. When it runs out, the node stops accepting: it logs accept error, backing off 100ms: Too many open files (os error 24) and tries again every 100 ms until descriptors come free. LimitNOFILE=1048576 covers several nodes at their connection limits. Check what a running instance has:

Terminal window
pid=$(systemctl show -p MainPID --value katana@xboard)
sudo ls /proc/$pid/fd | wc -l
grep 'open files' /proc/$pid/limits

These limits are built in and cannot be configured. They protect a node, not the host, so size LimitNOFILE= and memory for your real load rather than for these numbers.

Limit Value When it is reached
Live connections per node (stream nodes) 65,536 sockets A new socket is closed at once and counted as a failed handshake.
Sockets in their TLS, WebSocket or HTTP/2 handshake per node 2,048 The node stops accepting until one finishes or has held its place for 10 s.
Streams in their protocol handshake per node 512 The new stream is dropped and counted as a failed handshake.
Time to finish the protocol handshake (stream nodes) 10 s The connection is closed and counted as a failed handshake.
Connection with no traffic in either direction (stream nodes) 360 s The connection is closed. The protocol’s own idle timeout of 300 s usually closes a quiet flow first.
Hysteria 2 QUIC connections per node 4,096 A new connection is refused.
Hysteria 2 relayed streams and UDP associations per node, across all connections 65,536 A new stream is rejected; a new association’s packets are dropped.
Symptom Cause Fix
The unit is active (running) but a node has no listening on line The node’s panel request failed, the panel returned port 0 or no user list, or its bind failed. Or the node has no users yet. Read the node’s latest ERROR line, which ends in ; retrying in Ns. Once you fix the cause, the node comes up at its next retry; restart the instance only after a unit or capability change. With no users, the node binds at the first poll after users appear.
initial start failed: Permission denied (os error 13); retrying in …s Port below 1024 without CAP_NET_BIND_SERVICE, often after an upgrade removed a setcap capability. Use the unit’s AmbientCapabilities=, then restart the instance.
initial start failed: Address already in use (os error 98); retrying in …s Another instance or program holds the port. --test cannot detect this. Keep ports unique across instances. The node takes the port at its next retry after the other holder releases it.
The unit fails at ExecStartPre with configuration error: No such file or directory (os error 2) The config file is missing, or a file it names (geodata or [dns] ca_file) does not exist, often because of a relative path. Check the instance name against the file name, and use absolute paths.
configuration error: Permission denied (os error 13) from the start check, or configuration error: node N: Permission denied (os error 13) for a Hysteria 2 node’s certificate or key The katana user cannot read the config or a file it names. Group katana, mode 0640.
version `GLIBC_2.xx' not found The -gnu build on a host with an older glibc. Use the -musl build.
cp: cannot create regular file '/usr/local/bin/katana': Text file busy Copying over the running binary. Install by rename, as in the upgrade steps.
Sequences like [2m in the journal Colour codes. Environment=NO_COLOR=1.
journalctl -p warning shows nothing journald stores katana’s lines as info. Filter with -g 'WARN|ERROR'.
The service stopped and systemd did not restart it It received SIGHUP, or was stopped on purpose. Remove any ExecReload= that sends SIGHUP.
accept error, backing off 100ms: Too many open files (os error 24) Descriptor limit too low. Raise LimitNOFILE=.

More node-level errors, and what the panel’s answers mean, are on Troubleshooting.