Known limitations
Source files: 45 · checked against Etemenanki 596916d · katana v3.0.1
Etemenanki/app/src/instance.rsEtemenanki/app/src/main.rsEtemenanki/app/src/outbound/freedom.rsEtemenanki/app/src/outbound/udp_fanout.rsEtemenanki/app/src/serve.rsEtemenanki/environment/src/dial/udp.rsEtemenanki/protocols/src/helpers/address_family.rsEtemenanki/concepts/src/core.rsEtemenanki/protocols/src/vless/codec.rsEtemenanki/protocols/src/vmess/codec.rsEtemenanki/protocols/src/trojan/codec.rsEtemenanki/protocols/src/mux/demux.rsEtemenanki/protocols/src/socks/server.rsEtemenanki/protocols/src/socks/handshake.rsEtemenanki/protocols/src/socks/protocol.rsEtemenanki/protocols/src/socks/udp_link.rsEtemenanki/protocols/src/vless/core.rsEtemenanki/protocols/src/trojan/core.rsEtemenanki/protocols/src/vmess/core.rsEtemenanki/app/tests/integration/e2e_xray_mux.rsEtemenanki/app/tests/integration/e2e_hysteria_inbound.rsEtemenanki/app/tests/support/mod.rsEtemenanki/protocols/tests/unit/vless/codec.rsEtemenanki/protocols/tests/unit/mux/demux.rsEtemenanki/protocols/tests/pipeline/socks.rsEtemenanki/protocols/tests/unit/socks/server.rsEtemenanki/protocols/tests/unit/socks/protocol.rskatana/src/runtime.rskatana/src/manager/node.rskatana/src/manager/mod.rskatana/src/api/mod.rskatana/src/api/newv2board.rskatana/src/api/sspanel.rskatana/src/config.rskatana/src/inbound.rskatana/src/router.rskatana/src/connector.rskatana/src/outbound/freedom.rskatana/src/rule.rskatana/src/traffic.rskatana/tests/unit/e2e.rskatana/tests/unit/traffic.rskatana/tests/unit/connector.rskatana/tests/unit/api/newv2board.rskatana/tests/unit/runtime.rs
This page collects the functional limitations of the releases the rest of this guide describes: Etemenanki 2.0.2 (etemenanki-protocols 2.0.2; the other kernel crates and etemenanki-app are still 2.0.0) and katana 3.0.1, which builds on etemenanki-protocols 2.0.1. Entries marked fixed describe the earlier releases they name, and stay for operators who still run those releases. Each entry is behaviour that an operator notices or a contributor trips over: a reload that drops connections, an edit that is silently not applied, a panel field that is read under the wrong name. Each one is confirmed in the code and, where a test exists, in the tests.
Read it before you report a bug that may already be known, and before you change the component an entry names: the cause section points at the type or function responsible. The page lists functional behaviour only.
At a glance
Section titled “At a glance”Every entry has the same four parts. Symptom is what you see. Cause names the component and the mechanism, in a sentence or two, followed by the code it comes from. Workaround is what to do with the release you run. Status says whether a later release changes the behaviour, and which test, if any, pins it.
Etemenanki 2.0.2
Section titled “Etemenanki 2.0.2”Every reload replaces the whole generation
Section titled “Every reload replaces the whole generation”Symptom. Saving the config file closes every open connection on every inbound, including inbounds whose settings did not change and saves that only touch a comment. For a moment no inbound listens at all. With a Hysteria 2 inbound that holds live connections, or a TUN inbound, that gap can last up to a few seconds.
Cause. Instance::reload in app/src/instance.rs never patches a running generation. It builds a complete Built from the file, cancels the old generation’s CancellationToken, awaits every accept handle until the old listeners have given back their ports, and only then binds the new generation with spawn_generation(built, false). Cancelling the token drops every per-connection task. There is no hand-over of connections between generations.
pub struct Instance { path: PathBuf, state: tokio::sync::Mutex<State>,}
struct State { config: Config, last_bytes: Vec<u8>, generation: Generation,}
struct Generation { token: CancellationToken, accept_handles: Vec<JoinHandle<()>>,}
pub fn build(cfg: &Config) -> io::Result<Built>async fn spawn_generation(built: Built, strict: bool) -> io::Result<Generation>
impl Instance { pub async fn reload(&self)}With strict set to false, a bind failure on reload is logged as inbound <tag> bind <addr> failed: <error> and the other inbounds start. The old listener for that inbound is already gone, so the inbound stays down until a later reload binds it (see the next entries for when that happens).
Workaround. Batch edits into one save, and make them when a reconnect is acceptable. Clients reconnect on their own; Hysteria 2 clients see the connection closed with application error code 0x100 and reconnect to the new generation.
Status. Open. This is how the generation model works in 2.0.0. a_reload_rebinds_the_udp_port in app/tests/integration/e2e_hysteria_inbound.rs drives a reload while a Hysteria 2 connection is live and checks that the inbound comes back on its port. No test asserts that stream connections are dropped. The full cycle is described in Generations and hot reload.
[log].level is not reloaded
Section titled “[log].level is not reloaded”Symptom. Changing [log].level and saving logs a config reload: line that includes log changed, swaps the generation (dropping connections, as above), and leaves the log verbosity where it was.
Cause. main in app/src/main.rs calls init_tracing once, before anything else. It reads [log].level (default info) unless RUST_LOG is set, builds an EnvFilter and installs the subscriber with .init(). No reload handle is kept, and Instance::reload does not touch tracing.
fn init_tracing(config_path: &Path)Workaround. Restart etemenanki-app to change the level, or set RUST_LOG in its environment, which takes precedence over the file.
Status. Open. katana does reload its level: apply_reload calls the LogReload setter that init_tracing in katana returns.
Files the config references are not watched
Section titled “Files the config references are not watched”Symptom. Three related effects:
- A renewed certificate or key, a replaced
geoiporgeositefile, or a new DNSca_fileis not used, although the config names it by the same path. - A reload that failed because a referenced file was missing is not retried when the file appears.
- An inbound that failed to bind on reload does not come back when its port becomes free.
Cause. spawn_watcher watches the config file’s parent directory, non-recursively, and turns every event in it into a reload attempt after a 200 ms debounce. Instance::reload then compares the file’s bytes with State::last_bytes and returns at once when they are equal. Referenced files are read only inside build, so they are re-read only when the config bytes change. last_bytes is also updated when parsing or building fails, and after a reload in which an inbound failed to bind, so the same bytes are never tried twice.
flowchart TB
ev["notify event in the config directory"] --> deb["debounce 200 ms"]
deb --> read["std::fs::read(path)"]
read --> same{"bytes == last_bytes?"}
same -- yes --> none["return, nothing happens"]
same -- no --> parse{"config::parse_bytes ok?"}
parse -- no --> keep["log ERROR, last_bytes = bytes, old generation keeps serving"]
parse -- yes --> build{"build ok? reads certs, geodata, CA file"}
build -- no --> keep
build -- yes --> swap["cancel old token, await accept handles"]
swap --> spawn["spawn_generation(built, false)"]
spawn --> done["last_bytes = bytes"]
A certificate that lives outside the config directory does not even produce a watcher event.
Workaround. After replacing a referenced file or freeing a port, change the config bytes, for example by editing a comment, or restart. Either way the whole generation is replaced and connections drop. Saving the file unchanged does nothing.
Status. Open. No test covers it.
freedom UDP sends each name to one address
Section titled “freedom UDP sends each name to one address”Symptom. On a host where one address family does not work (no IPv6 route, or no IPv6 socket), UDP through a freedom outbound to a domain that resolves to both families can fail for the whole association: every packet to that name is dropped. TCP to the same name works, because the TCP dial tries every address in turn.
Cause. ResolvingUdp::poll_target in app/src/outbound/freedom.rs resolves a name with destination_to_socketaddrs, which applies the outbound’s address_family policy with FamilySupport::both(), and keeps only addrs.first(). It does not consider which families bind_dual actually bound, and it caches the result, success or failure, in resolved for the life of the association. If that first address’s family has no socket, DualStackUdp::poll_send_to fails with udp: no local socket in the family of <peer>; if the family has no route, the kernel refuses the send. Either way the runtime reports SendFailed and the packet is dropped.
pub struct ResolvingUdp { socket: DualStackUdp, resolver: Resolver, strategy: AddressFamilyStrategy, resolved: HashMap<CompactString, Option<IpAddr>>, resolving: Option<(CompactString, ResolveFuture)>,}
impl ResolvingUdp { fn poll_target(&mut self, cx: &mut Context<'_>, to: &Destination) -> Poll<Option<SocketAddr>>}Workaround. On a single-stack host, set address_family = "ipv4_only" (or "ipv6_only") on the freedom outbound, so that the only candidates are addresses the host can reach. "prefer_ipv4" also helps when every name has an IPv4 address, because it puts IPv4 first.
Status. Open in etemenanki-app. No test covers the family choice. katana’s own direct outbound (src/outbound/freedom.rs in katana) builds a FamilySupport from the sockets it bound and resolves only to those families. The dialer side is described in Dialers and socket policy.
VLESS and VMess outbounds send a UDP association to its first destination
Section titled “VLESS and VMess outbounds send a UDP association to its first destination”Symptom. A UDP association that talks to several peers through one vless or vmess outbound reaches only the first peer. Every later packet routed to that outbound goes to the first peer, whatever address it was meant for, and every reply is reported as coming from the first peer. This affects etemenanki-app and katana alike.
Cause. The VLESS and VMess UDP requests name one target in the request header, and the client codecs carry no address per packet. seal_to ignores its to argument and open_from attributes every reply to the stored target. The UDP fan-out (FanOutLink in app/src/outbound/udp_fanout.rs, FanOut in katana’s src/connector.rs) keeps one sub-link per outbound, keyed by the outbound’s Arc pointer and capped at MAX_SUBS = 64, not one per destination, so the second peer reuses the first peer’s sub-link.
pub type OpenedFrom = (Opened, Option<Destination>);
pub trait ProxyCoreEncodeDatagram: ProxyCoreEncodeHandshake { fn seal_to( &mut self, plain: &[u8], to: &Destination, out: &mut Staging<'_>, ) -> Result<Option<()>, Self::Error>;
fn open_from(&mut self, wire: &mut [u8]) -> Result<OpenedFrom, Self::Error>;}pub struct VlessDatagram { header: BytesMut, target: Destination, replied: bool,}
impl ProxyCoreEncodeDatagram for VlessDatagram { fn seal_to( &mut self, plain: &[u8], _: &Destination, out: &mut Staging<'_>, ) -> io::Result<Option<()>>;
fn open_from(&mut self, wire: &mut [u8]) -> io::Result<OpenedFrom>;}VMessDatagram in protocols/src/vmess/codec.rs has the same shape: its seal_to also takes _: &Destination, and its open_from returns Some(self.target.clone()) for every data frame.
Workaround. Route UDP that must reach several peers through an outbound that addresses each packet: in etemenanki-app, trojan (TrojanDatagram::seal_to writes to into every packet) or socks; in katana, which has no trojan outbound, socks or wireguard. Alternatively, split the peers across outbounds with route rules.
Status. Open. The outbounds do not speak XUDP. datagram_codec_frames_to_the_fixed_target in protocols/tests/unit/vless/codec.rs pins the behaviour: a packet sealed toward another address is framed for the fixed target. The katana side is covered in Connector and UDP fan-out.
Mux carriers repeat downlink frames or scramble VMess uploads
Section titled “Mux carriers repeat downlink frames or scramble VMess uploads”Symptom. With an Xray client that enables mux:
- Over VLESS or Trojan, when the client sends anything after receiving data on a mux connection, the server sends the most recent downlink frames a second time. A multiplexed TCP stream receives duplicated bytes and is corrupted; a UDP sub-flow over XUDP receives a duplicated packet.
- Over VMess, a large upload arrives scrambled: when a mux frame is split across two VMess chunks and a further chunk arrives in the same read, bytes of the completed frame are lost.
katana 3.0.0 nodes are affected too, because they serve these cores from etemenanki-protocols 2.0.0. In the test suite, two Xray interop tests in app/tests/integration/e2e_xray_mux.rs fail at 054cf34 (2.0.0): xudp_attributes_replies_to_the_right_peer and vmess_mux_payload_spans_both_framings. Both run only when Xray can be built: build_xray in app/tests/support/mod.rs compiles it from the reference tree with go. When go is missing or the build fails, the tests print a SKIP: line and pass.
Cause. Two separate defects in Demux, the mux.cool server in protocols/src/mux/demux.rs, in 2.0.0:
- Repeated frames (VLESS, Trojan).
Demux::outholds the frames of the last call. The downlink calls (on_outbound,on_datagram,on_outbound_gone) clear it and refill it, andVlessCore::on_subandTrojanCore::on_substagedemux.out()without taking it. The uplink callDemux::feeddoes not clear it: it only appendsEndframes for declined sessions. The carrier then stages everythingtake_out()returns, so the frames already sent go out again. - Scrambled uploads (VMess).
VMessCorecallsfeed_wholeonce per opened chunk. A frame that straddles a chunk boundary is completed in the held buffer and forwarded from it withEffect::ForwardHeldorEffect::SendToHeld, which the runtime applies after the whole event. Eachfeed_wholecall starts withtrim_held, which drains the held buffer up topartial_at, so a second chunk in the same read removes bytes that a held forward from the first chunk still points at.
The 2.0.0 signatures (2.0.1 replaces feed_whole, see Status):
pub const MAX_SESSIONS: usize = 256;
impl<T> Demux<T> { pub fn out(&self) -> &[u8]; fn trim_held(&mut self);}
impl<T: Send + Sync + 'static> Demux<T> { pub fn feed<C>( &mut self, plain: &[u8], base: usize, fx: &mut Effects<'_, C>, ) -> io::Result<usize> where C: ProxyCoreDecode<Key = FlowKey, Target = Flow<T>>;
pub fn feed_whole<C>( &mut self, plain: &[u8], base: usize, fx: &mut Effects<'_, C>, ) -> io::Result<()> where C: ProxyCoreDecode<Key = FlowKey, Target = Flow<T>>;
pub fn take_out(&mut self) -> Vec<u8>; pub fn on_outbound(&mut self, key: SubKey, data: &[u8]); pub fn on_datagram(&mut self, key: SubKey, from: &Destination, data: &[u8]);}The repeated frame explains the XUDP test failure: after the reply from the first peer, the client sends toward the second peer, and that uplink event stages the first peer’s reply again, so the client can receive it before the second peer’s reply.
sequenceDiagram participant C as Xray client participant K as VlessCore participant D as Demux participant P as peer A sub-flow C->>K: New frame toward peer A K->>D: feed, then take_out (empty) P->>K: reply from peer A K->>D: on_datagram fills out K->>C: stage(demux.out()), out keeps the frame C->>K: Keep frame toward peer B K->>D: feed appends nothing K->>D: take_out returns the peer A frame K->>C: peer A reply staged a second time
Workaround. Turn mux off on Xray clients that connect to a VLESS, Trojan or VMess inbound of an affected build, or upgrade to etemenanki-protocols 2.0.1 (katana 3.0.1).
Status. Fixed in etemenanki-protocols 2.0.1. Both uplink calls, feed and Demux::feed_chunks, clear out at the start of each call, so out holds only the frames of the last call and a frame the carrier staged once is never staged again. feed_chunks replaces feed_whole: VMessCore passes it every chunk of one read in a single call, and it trims the held buffer once per read, so every held forward of that read still finds its bytes. 2.0.1 adds the regression tests a_downlink_frame_is_not_sent_again_by_the_next_uplink and vmess_keeps_every_frame_one_read_completes in protocols/tests/unit/mux/demux.rs, and extends vless_answers_mux_at_once_and_demultiplexes with a further uplink that must stage nothing. 2.0.1 fixes both e2e_xray_mux failures above. katana 3.0.1 builds on 2.0.1. The carrier design is described in mux.cool and XUDP.
A SOCKS UDP association hears whichever address sends first
Section titled “A SOCKS UDP association hears whichever address sends first”Symptom. With etemenanki-protocols 2.0.1 or earlier, UDP through a socks inbound goes silent when another host, or another socket on the client’s own host, sends a datagram to the association’s relay port before the client does. From then on the relay forwards that sender’s datagrams, sends every reply to it, and drops the client’s own datagrams until the association ends. The control connection stays open, so the client sees no error. katana has no SOCKS inbound and is not affected.
Cause. In 2.0.1 and earlier, SocksInbound::associate in protocols/src/socks/server.rs set the association’s client with get_or_insert from the first datagram the hub received, of any kind, and never compared it with the control connection’s peer. The handshake read the request’s DST.ADDR and DST.PORT and discarded them.
Workaround. On an older build, firewall UDP to the relay address (udp_bind, or else the IP the client connected to) so that only client addresses reach it. Each association binds its hub to port 0, so the kernel picks a new ephemeral port every time and the rule has to cover the address, not one port. If the inbound’s clients do not need UDP, set udp = false.
Status. Fixed in etemenanki-protocols 2.0.2. An etemenanki-app built from that release includes the fix, but the etemenanki-app package was not bumped, so etemenanki-app --version still prints 2.0.0. The public API is unchanged.
associate now builds an ExpectedSender before it binds the hub. It takes the control connection’s peer (source, None over a Unix socket), the DST.ADDR and DST.PORT of the request, which the crate-private handshake_with_udp_source in protocols/src/socks/handshake.rs returns next to the Handshake, and the hub IP.
const STATUS_NOT_ALLOWED: u8 = 0x02;
struct ExpectedSender { ip: IpAddr, // canonical port: Option<u16>, client: Option<SocketAddr>, // where replies go, once pinned}
impl ExpectedSender { fn new( peer: Option<IpAddr>, declared: Option<&Destination>, hub: IpAddr, ) -> Result<Self, &'static str>; fn admits(&self, from: SocketAddr) -> bool; fn pin(&mut self, from: SocketAddr); fn client(&self) -> Option<SocketAddr>;}
fn hears(hub: IpAddr, ip: IpAddr) -> bool;pub(crate) fn endpoint(addr: SocketAddr) -> (IpAddr, u16)ExpectedSender::new decides whom the association hears. A declared address counts only when it is an IP that is not unspecified; a domain or 0.0.0.0 names nothing.
| Control connection | The request names | The association hears |
|---|---|---|
TCP from peer P |
P with a non-zero port |
P on that port, from the start. |
TCP from peer P |
P with port 0, any other IP, an unspecified address or a domain |
P, on the port of the first datagram forwarded. The declared source is set aside, neither trusted nor refused. |
| Unix socket | an IP that is not unspecified, with a non-zero port | Exactly that address and port. |
| Unix socket | anything else | Nobody: refused with 0x02. |
The declared source is set aside rather than refused because clients name addresses their datagrams do not come from: sing-box names a loopback address of the other family whenever its first target is private, a client behind NAT names its LAN address, and PySocks names only a port. Naming another address never opens the association to that address.
Then hears(hub, ip) checks that the hub can receive from that IP: an IPv4 or IPv4-mapped hub hears only IPv4, the unspecified IPv6 hub (::) hears both families, and any other IPv6 hub hears only IPv6. If it cannot, or if a Unix-socket request names no exact source, associate replies 0x02 (“connection not allowed by ruleset”) and fails with PermissionDenied, before any hub is bound.
On the hub, each datagram meets admits before it is parsed. admits compares with endpoint, which puts the IP in canonical form (an IPv4-mapped IPv6 address is the IPv4 one) and leaves out IPv6 flow info and scope. It requires the expected IP, the named port if there is one, and, once the association is pinned, the pinned port. Anything else is dropped unread. pin runs only for a datagram that parses and carries a payload, so a neighbour on the client’s address cannot claim the association with junk, and the first pin holds. Replies go to client(), in the form the hub saw it.
What operators may notice after upgrading:
- A client on a Unix-socket
socksinbound must name the exact address and port its datagrams will come from. One that leaves either at zero, or names a domain, gets0x02. - A relay that could never hear the client, because it is bound in another address family, is refused with
0x02instead of going quiet. For example,udp_bind = "::1"with clients that connect over IPv4. - Datagrams from any IP other than the control connection’s are dropped. A client whose UDP leaves from a different address than its TCP connection is not heard.
- In
etemenanki-app, a refused association ends the connection with adebugline such assocks connection from Some(192.0.2.10) ended: socks: UDP associate from an address family the relay is not bound in, or, over a Unix socket,socks connection from None ended: socks: UDP associate over a unix socket must name its source address and port. - On the client side,
SocksUdpLinkinprotocols/src/socks/udp_link.rscompares a reply’s source with the relay throughendpointtoo, so a dual-stack socket hears an IPv4 relay whose replies arrive from its IPv4-mapped address. Its own request still names0.0.0.0:0, which the table above sets aside.
These tests pin the behaviour:
| Test | File | Pins |
|---|---|---|
udp_association_ignores_another_ip (Linux) |
protocols/tests/pipeline/socks.rs |
A host on 127.0.0.2 that sends first, and again later, is neither forwarded nor answered. |
udp_association_ignores_another_port_once_pinned |
protocols/tests/pipeline/socks.rs |
After the first datagram, another socket on the client’s IP is not heard. |
udp_association_holds_to_the_port_the_request_names |
protocols/tests/pipeline/socks.rs |
A request naming the client’s own address and port holds the association to it before any datagram. |
udp_association_sets_aside_a_source_it_cannot_hold_to |
protocols/tests/pipeline/socks.rs |
A request naming [::1]:0 from an IPv4 client is granted and works. |
udp_association_is_not_widened_by_the_request (Linux) |
protocols/tests/pipeline/socks.rs |
Naming another address does not let that address in. |
udp_association_over_a_unix_socket_needs_its_exact_source (Unix) |
protocols/tests/pipeline/socks.rs |
0.0.0.0:0 over a Unix socket gets 0x02 and a closed connection; an exact source is held to its port. |
udp_association_refuses_a_relay_that_cannot_hear_the_client |
protocols/tests/pipeline/socks.rs |
udp_bind = ::1 with an IPv4 client gets 0x02 and a closed connection. |
udp_link_ignores_datagrams_not_from_the_relay |
protocols/tests/pipeline/socks.rs |
SocksUdpLink ignores a well-formed reply from another socket. |
udp_link_on_a_dual_stack_socket_hears_an_ipv4_relay (Linux) |
protocols/tests/pipeline/socks.rs |
SocksUdpLink on a [::] socket hears an IPv4 relay. |
only_the_control_peer_is_heard_and_its_first_datagram_pins_the_port, an_ipv4_mapped_address_is_the_ipv4_one, a_request_naming_the_peer_pins_its_port_up_front, a_request_naming_any_other_source_is_set_aside, over_a_unix_socket_the_request_must_name_the_exact_source, a_relay_that_cannot_hear_the_client_is_refused |
protocols/tests/unit/socks/server.rs |
ExpectedSender and hears case by case, as in the table above. |
endpoint_sees_through_ipv4_mapping_and_ignores_flow_info |
protocols/tests/unit/socks/protocol.rs |
endpoint equates an IPv4-mapped address with the IPv4 one and ignores flow info. |
The driver and its relay are described in SOCKS.
katana 3.0.1
Section titled “katana 3.0.1”Several katana entries are about hot reload. For orientation, this is the path a saved config takes through apply_reload in src/runtime.rs and on to each node:
flowchart TB
load["config::load ok, after 500 ms debounce"] --> ob{"new.outbounds != cfg.outbounds?"}
ob -- yes --> pool["build_outbounds: new pool and Resolver"]
pool -- error --> reject["log error, nothing applied"]
pool -- ok --> nodes{"build_node for each added node, PanelClient::new for each changed kept node"}
ob -- no --> nodes
nodes -- error --> reject
nodes -- ok --> lvl["log level swapped if changed"]
lvl --> fan["StaticUpdate::Outbounds to every node, if the pool was rebuilt"]
fan --> rm["cancel and await nodes whose identity is gone"]
rm --> ids{"each new node: identity already running?"}
ids -- no --> spawn["spawn_built from the pre-built manager"]
ids -- yes --> eq{"NodeConfig changed?"}
eq -- yes --> upd["StaticUpdate::Config"]
eq -- no --> skip["nothing for this node"]
The node’s identity and the channel it listens on:
type NodeId = (String, String, u32, String, String);
fn identity(cfg: &NodeConfig) -> NodeId
struct NodeHandle { id: NodeId, cfg: NodeConfig, task: JoinHandle<()>, static_tx: mpsc::Sender<StaticUpdate>, shutdown: CancellationToken,}pub enum StaticUpdate { Config(Box<NodeConfig>), Outbounds(Arc<HashMap<CompactString, Arc<Outbound>>>),}identity returns (panel_type lowercased, api.host, api.node_id, api.key, panel node type). The panel node type is the node_type query parameter a newV2board node asks with (vless for a V2ray-family node, node_type V2ray, Vmess or Vless, with enable_vless, otherwise the lowercased node_type), because UniProxy finds a node by its id and that type; it is empty for sspanel, which finds a node by its id alone. Log lines show the identity as <panel>@<host>#<node_id>, followed by /<node type> when that is not empty, and never show the key.
Before anything is applied, every node the reload adds must build (its panel client, and its router against the new pool if there is one), and every kept node whose config changed must build its panel client. If one does not, katana logs reload: node <display_id>: <error>; keeping current config at ERROR and applies nothing, not even the outbound pool or the log level. a_reload_with_a_node_that_does_not_build_changes_nothing in tests/unit/runtime.rs pins this. A node that did not build at startup was never started, so every later reload builds it as a new node: while that entry still does not build, every reload is refused, until you fix or remove it. The node-side handling is in Node manager and the process side in Process runtime and reload.
A node whose bootstrap fails stays down
Section titled “A node whose bootstrap fails stays down”Symptom. With katana 3.0.0, katana keeps running and serves its other nodes, but one node never listens. At startup it logged one of these lines at ERROR:
| Condition | Log line |
|---|---|
The panel answers node_info with an error |
node <id>: node_info failed: <error> |
| The panel answers with nothing (HTTP 304) | node <id>: panel returned no node info |
The node description has port 0 (newV2board instead fails node_info with newV2board: server port must be > 0) |
node <id>: panel returned port 0 |
| Building or binding the listener fails | node <id>: initial start failed: <error> |
Later edits to that node log reload: reconfigured node … and change nothing. A panel that is unreachable for a few seconds at boot, DNS that is not ready yet, a port still held by another process, or node_type = "hy2" on Xboard is enough.
Cause. In katana 3.0.0, NodeManager::run in src/manager/node.rs makes one bootstrap attempt and returns on any of the failures above, which ends the node’s task. The runtime keeps the NodeHandle but never watches its JoinHandle, and a later static_tx.send(…) fails without a log line because the receiver is gone. run in src/runtime.rs exits with no nodes could be started only when no node task could be spawned at all, which is decided before any bootstrap runs.
impl NodeManager { pub async fn run( self: Arc<Self>, shutdown: CancellationToken, mut static_rx: mpsc::Receiver<StaticUpdate>, )}Workaround. With katana 3.0.0, fix the cause, then restart katana. To restart only the affected node, change a field of its 3.0.0 identity tuple, for example [node.api].timeout, and save: the reload removes the dead node and spawns a fresh one, which bootstraps again. In katana 3.0.1 timeout is no longer part of the identity, and no workaround is needed.
Status. Fixed in katana 3.0.1. NodeManager::bootstrap retries a failed attempt until the node is up. The wait starts at 1 s and doubles after each failure up to 60 s, and never exceeds controller.update_periodic. Each failure logs node <id>: <reason>; retrying in <n>s at ERROR, where <reason> is node_info failed: …, panel returned no node info, panel returned port 0, user_list failed: …, panel returned no user list or initial start failed: …. A failed user list now fails the attempt too, instead of starting the node with no users. Every attempt forgets the panel ETags, so the panel answers in full rather than with HTTP 304. A config edit saved during the wait cuts it short, so the next attempt runs at once, and cancellation (a reload that removes the node, or shutdown) still stops the retries. a_node_comes_up_once_the_panel_answers, a_node_whose_port_is_taken_comes_up_once_it_is_free and a_node_that_never_bootstraps_still_stops in tests/unit/e2e.rs pin the retry.
Some [node.api] edits are not applied
Section titled “Some [node.api] edits are not applied”Symptom. With katana 3.0.0, editing api.node_type, api.enable_vless, api.vless_flow, api.speed_limit, api.rule_list_path or api.disable_custom_config logs reload: reconfigured node <panel>@<host>#<id>, and the node keeps behaving as before. api.enable_vless is half applied: the listener is rebuilt (dropping the node’s connections), but the panel client still asks for and parses the old node type. Turning it off keeps serving VLESS, because build_protocol in src/inbound.rs serves VLESS when either the config flag or the panel client’s NodeInfo.enable_vless is set.
Two exceptions and one related gap:
- For a Hysteria 2 node described locally (
[node.hysteria].portset),NodeManager::node_infobuilds the node description from the live config, includingapi.speed_limit, at the next poll. The panel client still applies the old value as each user’s override, and the effective rate is the lower of the two, so lowering the limit takes effect and raising it does not. rule_list_pathis cached, but the file it names is read again at every rule refresh, so edits to the file’s contents are picked up.- Outside
[node.api], an edit tocontroller.disable_sniffingor to[node.hysteria]is stored and takes effect only at the node’s next listener rebuild, becauseapply_staticdoes not compare those fields.
Cause. In katana 3.0.0, each node’s PanelClient is built once, in spawn_node, and NodeManager holds it immutably. The panel clients copy these fields out of NodeConfig at construction. identity does not include them, so an edit to one becomes a StaticUpdate::Config, and NodeManager::apply_static looks only at controller.listen_ip, controller.cert, api.enable_vless and route. The new values are stored and never reach the client.
| Field | Cached by newv2board::Client |
Cached by sspanel::Client |
|---|---|---|
node_type |
yes, as node_type and the node_type query parameter |
yes |
enable_vless |
yes: query parameter, settings key and NodeInfo.enable_vless |
yes |
vless_flow |
not read | yes |
speed_limit |
yes, as speed_limit_mbps |
yes |
rule_list_path |
yes | yes |
disable_custom_config |
not read | yes |
Workaround. With katana 3.0.0, restart katana, or change [node.api].timeout in the same save so the node is respawned with a new client. Either way the node’s connections drop.
Status. Fixed in katana 3.0.1:
- On newV2board, an edit that changes the panel node type in the identity (a
node_typechange other than letter case, orenable_vlesson a V2ray or VMess node; a VLESS node is asked for asvlesseither way) respawns the node, because it now names another panel node. - Any other
[node.api]edit outside the identity reaches the running node asStaticUpdate::Config, andapply_staticbuilds a newPanelClientfrom the edited config. A newV2board client takes over the routes its predecessor read (inherit_routes), so audit rules derived from them hold until it reads the node config itself. The node then runs one panel poll at once with the new client, and the reconcile ladder drops only what the new answer requires: a client-only edit such asrule_list_pathortimeoutkeeps the node’s connections. - An edit to
controller.listen_ip,controller.cert,controller.disable_sniffing,api.enable_vless,[node.hysteria]or the route forces a listener rebuild in that poll. - An edit is taken whole or not at all. One whose panel client does not build is refused by the reload, as described above. One whose router does not build is refused by the node, which logs
node <id>: config edit refused, keeping the running one: <error>and keeps running as it was.
a_newv2board_type_edit_respawns_the_node, an_sspanel_api_edit_takes_effect_in_place and a_client_edit_takes_effect_without_dropping_connections in tests/unit/runtime.rs, and a_rebuilt_client_keeps_the_routes_it_has_not_read in tests/unit/api/newv2board.rs, pin this.
A [dns]-only edit is not applied
Section titled “A [dns]-only edit is not applied”Symptom. After an edit that changes only [dns], katana logs no reload: line, and the outbounds keep resolving through the old resolver.
Cause. build_outbounds builds the one Resolver that the built-in direct handler and every [[outbound]] share. apply_reload calls it only when new.outbounds != cfg.outbounds. It still stores the new config with *cfg = new, so the new [dns] takes effect the next time the pool is built.
Workaround. Restart katana, or save the [dns] change together with an [[outbound]] edit (which rebuilds every node’s listener, see the next entry).
Status. Open.
Any [[outbound]] edit rebuilds every node’s listener
Section titled “Any [[outbound]] edit rebuilds every node’s listener”Symptom. Adding, removing or changing any [[outbound]], even one no route uses, drops every connection on every node. katana logs reload: outbound pool rebuilt, and each node logs node <id>: router rebuilt and, if it has users, node <id>: listening on <ip>:<port>.
Cause. A rebuilt pool is sent to every node as StaticUpdate::Outbounds. apply_static stores it, recompiles the router with rebuild_router, and calls apply_route_change, which does a full rebuild: tear_down of the TransportManager, then bring_up. katana drops every connection on a route change on purpose, and tearing down the listener is the only way to end all of them, including those still in the handshake. The router holds the pool’s Arc<Outbound> handles, so a new pool always means a new router.
A [node.route] edit takes a different path to the same rebuild: apply_static compiles the new router first and refuses the edit if it does not compile (node <id>: config edit refused, keeping the running one: <error>, nothing rebuilt), then forces the rebuild through poll_cycle(true).
If the new pool lacks a tag a running node’s route names, build_router fails with route references unknown outbound tag: <tag>. The node logs node <id>: route rebuild failed, keeping current: <error>, keeps its old router, and still rebuilds its listener. A node that the same save adds is compiled against the new pool before anything is applied, so a missing tag in its route refuses the whole reload.
Workaround. Group outbound edits and make them when a reconnect is acceptable. Run katana --test -c <file> before saving: test_config compiles each node’s router against the new pool and catches a route that names a missing tag.
Status. Open. route_change_drops_connections in tests/unit/e2e.rs pins the drop for a route edit sent as StaticUpdate::Config; no test drives the Outbounds path.
Renewed certificates and replaced geodata are not picked up
Section titled “Renewed certificates and replaced geodata are not picked up”Symptom. After a certificate renewal that keeps the same file paths, the node keeps serving the old certificate. A replaced geoip or geosite file is not used.
Cause. build_transport in src/inbound.rs reads cert_file and key_file when bring_up builds a listener, and build_router in src/router.rs reads the geodata files when it compiles a router. A reload compares NodeConfig values, and a file replaced at the same path leaves them equal, so nothing is rebuilt.
Workaround. Restart katana after each renewal, for example from your ACME client’s deploy hook. Pointing cert_file and key_file at new paths also works, because a changed controller.cert rebuilds the listener. For geodata, save a change to [node.route] or [[outbound]], or restart.
Status. Open.
VLESS transport settings under networkSettings are ignored
Section titled “VLESS transport settings under networkSettings are ignored”Symptom. With current Xboard, a VLESS node on WebSocket is served at path / with no Host check, whatever path the panel shows clients, and a VLESS node on gRPC is served with an empty service name. Clients that use the panel’s path or service name cannot connect. VLESS over TCP, with or without TLS, and every VMess node are not affected.
Cause. The UniProxy response struct in src/api/newv2board.rs declares the settings twice, and parse_v2ray reads network_settings when enable_vless is set. Xboard’s ServerService::buildNodeConfig sends the settings of every node type under networkSettings, so a VLESS node finds none. build_transport then falls back to / for an empty WebSocket path and uses the empty service name as it is.
struct ServerConfig { #[serde(default, rename = "networkSettings")] network_settings: Option<NetworkSettings>, #[serde(default, rename = "network_settings")] vless_network_settings: Option<NetworkSettings>, // …}Workaround. On such a panel, serve VLESS over TCP and TLS, or set the WebSocket path to / in the panel. Panels that send VLESS settings under network_settings work as intended.
Status. Open. parse_v2ray_ws_tls in tests/unit/api/newv2board.rs covers the VMess key only; no test covers the VLESS key.
node_type = "hy2" is refused by Xboard
Section titled “node_type = "hy2" is refused by Xboard”Symptom. katana --test accepts node_type = "hy2", but at runtime every UniProxy request the node makes fails:
- When the panel describes the node (
[node.hysteria].portis 0), the node never comes up and retries its bootstrap forever, loggingnode <id>: node_info failed: GET /api/v1/server/UniProxy/config; retrying in <n>s. - When
[node.hysteria].portis set, katana does not ask the panel for the node, but the user list still fails, and a failed user list fails the bootstrap too. The node binds nothing and keeps retrying, loggingnode <id>: user_list failed: GET /api/v1/server/UniProxy/user; retrying in <n>s.
reload: lines show such a node as <panel>@<host>#<node_id>/hy2, because the panel node type is part of its identity.
Cause. NodeType::parse in src/api/mod.rs accepts hysteria2, hysteria and hy2. newv2board::Client::new sends the configured node_type, lowercased, as the node_type query parameter of every UniProxy request (only a V2ray-family node with enable_vless is sent as vless). Xboard validates that parameter against its own type list after mapping the aliases v2ray → vmess and hysteria2 → hysteria, and rejects hy2 with Invalid node type specified. --test never contacts the panel, so it cannot notice.
Workaround. Write node_type = "hysteria2" or "hysteria".
Status. Open. node_type_param_vless in tests/unit/api/newv2board.rs pins how the parameter is derived for the other node types.
Audit rules see only the requested address
Section titled “Audit rules see only the requested address”Symptom. A domain audit rule does not match a connection that the client opens to an IP address, even when its TLS server name or HTTP Host names that domain and a routing rule for the same domain matches it.
Cause. KatanaConnector::connect in src/connector.rs passes flow.sniffed to route_target for routing, but Dispatcher::forbidden matches the rules against dest_string(&flow.destination): the requested domain or IP literal, without the port. UDP packets are audited the same way, each against its own destination.
impl RuleManager { pub fn detect(&self, tag: &str, dest: &str, uid: Option<i64>) -> bool}Workaround. Add IP patterns to the audit rules, or block the domain with a routing rule to the block outbound. For TCP, routing sees the sniffed name while sniffing is on; UDP packets are routed by their destination only. See Destination audit.
Status. Open. a_forbidden_destination_is_refused_and_recorded in tests/unit/connector.rs pins that a rule is matched against the requested address without its port, and that the hit is recorded against the user.
Traffic counters live in memory, and a timed-out report can be billed twice
Section titled “Traffic counters live in memory, and a timed-out report can be billed twice”Symptom. Two effects:
- If katana is killed or crashes, the bytes it has not reported yet are lost. A clean shutdown sends one final report per node; if that report fails, its bytes are lost too.
- If the panel records a report but its answer arrives after the HTTP timeout, katana counts the report as failed and sends the same bytes again in the next cycle, and the panel bills them twice.
Cause. NodeTraffic in src/traffic.rs keeps every counter in memory and nothing is persisted. NodeManager::report_traffic commits a report only on success: it subtracts exactly the reported bytes with UserCounter::commit_reported and discards the residual rows. On any error, a timeout included, it calls restore_residuals and leaves the live counters untouched, so the next cycle reports those bytes again. katana sends no request identifier with a report to the UniProxy push endpoint or to SSPanel’s /mod_mu/users/traffic, so neither side can tell a resend from a new report, and katana cannot tell a lost report from a slow one. The timeout is [node.api].timeout seconds, and 0 means 5 (ApiConfig::timeout_secs).
impl UserCounter { pub fn commit_reported(&self, up: u64, down: u64)}
impl NodeTraffic { pub fn snapshot(&self) -> Vec<TrafficSnapshot> pub fn restore_residuals(&self, rows: Vec<(i64, u64, u64)>)}Workaround. Stop katana with SIGTERM or SIGINT rather than SIGKILL, so the final report runs. Raise [node.api].timeout above the panel’s slowest normal response time. Changing it rebuilds the node’s panel client in place; connections stay unless the node info read with the new client changes the node’s protocol or transport. Watch for repeated node <id>: report traffic: <error> warnings.
Status. Open. katana resends rather than drops by design: a dropped report would under-bill. restored_residuals_are_retried and commit_reported_preserves_concurrent in tests/unit/traffic.rs pin the resend and the exact-subtraction commit. Details are in Traffic accounting and, for operators, Traffic reporting.
No online-user, alive-IP or node-status reporting
Section titled “No online-user, alive-IP or node-status reporting”Symptom. The panel shows no online IPs per user, cannot enforce a device limit, and gets no node status from katana. The [node.api].device_limit key is accepted and never read.
Cause. PanelClient in src/api/mod.rs has five operations: node_info, user_list, report_user_traffic, node_rule and report_illegal. The newV2board client calls only the UniProxy config, user and push endpoints; alive, alivelist and status are never called. The SSPanel client never calls /mod_mu/users/aliveip and sends no node status.
Workaround. None in katana.
Status. Open: not implemented.