Contributing
Source files: 75 · checked against Etemenanki 596916d · katana v3.0.1
Etemenanki/Cargo.tomlEtemenanki/Cargo.lockEtemenanki/.cargo/config.tomlEtemenanki/.github/workflows/build.ymlEtemenanki/README.mdEtemenanki/concepts/Cargo.tomlEtemenanki/environment/Cargo.tomlEtemenanki/protocols/Cargo.tomlEtemenanki/app/Cargo.tomlEtemenanki/concepts/src/buffer.rsEtemenanki/concepts/src/runtime.rsEtemenanki/environment/src/dial/udp.rsEtemenanki/environment/src/routing.rsEtemenanki/protocols/src/socks/udp_link.rsEtemenanki/protocols/src/socks/server.rsEtemenanki/protocols/src/socks/protocol.rsEtemenanki/protocols/src/socks/handshake.rsEtemenanki/protocols/src/helpers/address_family.rsEtemenanki/protocols/src/mux/demux.rsEtemenanki/protocols/src/wireguard/device.rsEtemenanki/protocols/src/hysteria/server/config.rsEtemenanki/protocols/src/hysteria/server/inbound.rsEtemenanki/protocols/src/hysteria/server/datagrams.rsEtemenanki/app/src/serve.rsEtemenanki/app/src/transport.rsEtemenanki/app/src/config.rsEtemenanki/app/src/instance.rsEtemenanki/concepts/tests/runtime.rsEtemenanki/concepts/tests/client.rsEtemenanki/protocols/tests/pipeline/socks.rsEtemenanki/protocols/tests/unit/socks/server.rsEtemenanki/protocols/tests/unit/socks/protocol.rsEtemenanki/protocols/tests/pipeline/wireguard.rsEtemenanki/protocols/tests/unit/wireguard/device.rsEtemenanki/environment/tests/integration/udp.rsEtemenanki/environment/tests/unit/routing.rsEtemenanki/protocols/tests/unit/helpers/address_family.rsEtemenanki/protocols/tests/unit/vless/protocol.rsEtemenanki/protocols/tests/unit/vmess/protocol.rsEtemenanki/protocols/tests/unit/trojan/protocol.rsEtemenanki/protocols/tests/unit/mux/frame.rsEtemenanki/protocols/tests/unit/transports/grpc_framing.rsEtemenanki/protocols/tests/unit/dns/message.rsEtemenanki/protocols/tests/unit/hysteria/protocol.rsEtemenanki/protocols/tests/unit/hysteria/auth.rsEtemenanki/protocols/tests/unit/hysteria/server/datagrams.rsEtemenanki/app/tests/unit/config.rsEtemenanki/app/tests/unit/transport.rsEtemenanki/app/tests/integration/e2e_hysteria_inbound.rsEtemenanki/app/tests/support/mod.rskatana/Cargo.tomlkatana/Cargo.lockkatana/.cargo/config.tomlkatana/.github/workflows/build.ymlkatana/.github/workflows/release.ymlkatana/src/config.rskatana/src/serve.rskatana/src/manager/proxy.rskatana/src/manager/node.rskatana/src/runtime.rskatana/src/api/mod.rskatana/src/api/newv2board.rskatana/src/api/sspanel.rskatana/src/manager/mod.rskatana/src/traffic.rskatana/src/meter.rskatana/tests/unit/traffic.rskatana/tests/unit/meter.rskatana/tests/unit/serve.rskatana/tests/unit/outbound.rskatana/tests/unit/inbound.rskatana/tests/unit/e2e.rskatana/tests/unit/runtime.rskatana/tests/unit/api/newv2board.rskatana/tests/support/mod.rs
This page is the working agreement for anyone who changes the kernel or katana. It says which repository a change belongs in, which commands must pass before it lands, what every reviewer checks regardless of the change, how versions are chosen, in which order crates are released, which git operations are off limits, and how these documentation pages are kept in step with the code.
The rules are stable; the numbers are not. Versions, branches, tags and dependency resolutions change with every release, so every step below starts by reading them from the repository rather than from this page. Where this page and a repository disagree, the repository is right and this page is due for an update (see Keeping these docs current).
Where a change belongs
Section titled “Where a change belongs”The code lives in separate git repositories. None of them is a parent of the others, and there is no Cargo workspace spanning them: run every git, Cargo, test and release command inside the repository you are changing.
| You are changing | Repository | Path | Package |
|---|---|---|---|
| The sans-I/O core contract, the per-connection runtimes, links, connectors, shared net types | Etemenanki | concepts/ |
etemenanki-concepts |
| Host dialers (TCP, UDP, QUIC), the shared socket policy, the route model | Etemenanki | environment/ |
etemenanki-environment |
| A protocol, a transport, mux, sniffing, DNS | Etemenanki | protocols/ |
etemenanki-protocols |
| TOML configuration, inbound and outbound construction, router composition, generations and hot reload of the standalone proxy | Etemenanki | app/ |
etemenanki-app |
| Panel clients, node lifecycle, admission, traffic accounting, speed limits, audit, katana’s own configuration and reload | katana | src/ |
katana |
Three boundaries catch people out:
- Routing semantics belong in
environment/src/routing.rs. The route model was vendored unchanged from the standalone routing crate (harranu) asetemenanki_environment::routing, and both the app and katana route with that copy. A change made only in harranu reaches neither. When harranu itself needs a change, it is released on its own, never as part of a kernel release; weigh whether the same change also belongs inenvironment/. - Reference trees are not implementations.
Xray-core/andhysteria/in Etemenanki are upstream sources (git submodules) kept for protocol reference and built by the interop tests of both repositories;Xboard/,V2bX/andXrayR/in katana are kept for panel reference. A bug that shows up against them is fixed in the Rust code, never by editing the reference tree. - katana sees only published kernels. katana depends on
etemenanki-concepts,etemenanki-environmentandetemenanki-protocolsby version from a private Cargo registry. A kernel fix reaches katana only after the affected crates are published and katana’s lockfile is moved to them. Workspace and crates has the details, including which features katana enables (vendored-opensslandhysteriaonetemenanki-protocols, notun).
Files ignored by .gitignore (scratch copies, build output) are never authoritative. Review and change only tracked sources.
Before you start
Section titled “Before you start”-
Read the live state of the repository. Every later decision depends on it.
Terminal window git status --short --branchgit branch --show-currentgit remote -vgit log --oneline --decorate --max-count=12git tag --sort=-v:refname | head -20 -
Read the manifests.
Cargo.toml,Cargo.lockand.cargo/config.tomlsay which versions exist, which registry a crate publishes to and how the registry authenticates. For the package graph, usecargo metadata --no-deps --format-version 1. The version a dependency actually resolves to is inCargo.lockorcargo metadata, not in the requirement inCargo.toml. -
Read the call chain, not just the function. For a change to a public item in
concepts,environmentorprotocols, search katana for every use of it: a signature change there is a breaking release (see Versioning). -
Walk the failure paths. For configuration, protocol, network or lifecycle code, ask what happens on malformed input, on a peer that stops reading, on cancellation, and when a dependency (DNS, a port, the panel) is not there. The review checklist lists what has gone wrong before.
Fix the cause, not the test. Keep the change as small as it can be while still complete, leave unrelated code alone, and do not add a dependency the change does not need.
Validation gates
Section titled “Validation gates”A change lands only when its repository’s gates pass. The release gates are run in addition, immediately before publishing or tagging.
cargo fmt --all -- --checkcargo test --workspacecargo clippy --workspace --all-targets --all-features -- -D warnings# before a release, additionally:cargo check --workspace --lockedcargo fmt --all -- --checkcargo test --lockedcargo clippy --all-targets --all-features -- -D warnings# before a release, additionally:cargo metadata --locked --no-deps --format-version 1cargo check --lockedcargo fmt --checkcargo testcargo clippy --all-targets --all-features -- -D warnings| Gate | What it protects |
|---|---|
cargo fmt … --check |
One formatting, so diffs show only real changes |
cargo test |
Behaviour, including the tests named in the checklist below |
cargo clippy … -D warnings |
Every warning is an error. In etemenanki-environment and etemenanki-protocols this is also what enforces the crate-wide denies on unwrap_used, expect_used, indexing_slicing and arithmetic_side_effects, which keep untrusted bytes from panicking the process; cargo build and cargo test do not check them |
--all-features |
Lints the code behind features no workspace member turns on: quic on etemenanki-environment and vendored-openssl on etemenanki-protocols. (hysteria and tun are compiled anyway, because etemenanki-app enables them.) |
--locked |
Fails instead of silently changing Cargo.lock, so what is tested is what is released |
Neither repository’s CI runs these gates. Etemenanki’s build.yml only builds the etemenanki-app release binary, and katana’s build.yml only builds the katana release binary with --locked, so a green CI run says nothing about tests or clippy. Run the gates locally.
A green cargo test can still have skipped things. The Xray and Hysteria interop tests build the upstream binaries with go from the reference trees. When go is missing or the build fails, the support code prints a SKIP: line and each test that needs the binary returns early and passes, some after printing skipping: hysteria binary unavailable (app/tests/support/mod.rs, and katana’s tests/support/mod.rs, which looks for the trees in the sibling ../Etemenanki checkout). Changes to a protocol, a transport, TLS, SOCKS or WireGuard should be checked with those interop tests actually running. Testing lists what each suite needs.
Before asking for review, check the diff itself:
git diff --checkgit diff --statgit diffLook for unrelated files, leftover debug output, temporary or generated files, a lockfile that moved more than the change explains, and anything that looks like a credential, token or private key.
Review checklist
Section titled “Review checklist”These invariants come from the architecture and from past reviews. They are not a list of open bugs. Each one describes a property the code must have and must not lose. Where the verified revision does not fully meet an item yet, the item says so, and a change in that area should move towards the rule, not away from it. Apply every item that touches the code under review, and add a test for any mechanism a change introduces or moves.
Network code
Section titled “Network code”UDP datagrams come from the association’s peer
Section titled “UDP datagrams come from the association’s peer”A UDP relay must only accept datagrams from the peer it is serving, and a client must only accept replies from the relay it associated with. A packet from any other address is dropped, never relayed or delivered.
-
Server: the SOCKS server’s UDP association loop in
protocols/src/socks/server.rshears one client (RFC 1928 §7), described by anExpectedSenderthat is built before the relay socket is bound. Over TCP the client is the control connection’s peer:admitsrequires every datagram’s IP to be that peer’s, and a datagram from any other IP is dropped unread. The first datagram that parses and carries a payload pins the port (pin; the first pin holds), so a neighbour on the same address cannot claim the association by sending junk first. AUDP ASSOCIATEwhose DST.ADDR names the peer’s own IP with a non-zero port pins that port up front. A request that names anything else (another address, an unspecified address, a domain) is set aside rather than trusted or refused, because it is not where the datagrams come from: a client behind NAT names its LAN address, and sing-box names a loopback address whenever its first target is private. Replies go only to the pinned client, in the form the relay socket saw it. Over a Unix socket there is no peer, so the request must name its exact IP and a non-zero port. When there is no client the relay could hold to, the server answers0x02(STATUS_NOT_ALLOWED, “connection not allowed by ruleset”) and closes the control connection. That happens for a Unix-socket request that names a domain or leaves the address or the port at zero, and for a relay (udp_bind) in an address family the client’s IP is not in:hearslets an IPv4 or IPv4-mapped relay hear only IPv4, the unspecified IPv6 relay::hear both families, and any other IPv6 relay hear only IPv6. An association never outlives its control connection: EOF or an error on the TCP stream ends it, and so doesRELAY_IDLE_TIMEOUTwithout traffic. Datagrams with a non-zeroFRAGbyte are discarded byparse_udp_packetinprotocols/src/socks/protocol.rs.struct ExpectedSender {ip: IpAddr, // canonical; every datagram must come from itport: Option<u16>, // when the request named it alongside `ip`client: Option<SocketAddr>, // the first sender forwarded for; replies go here}impl ExpectedSender {fn new(peer: Option<IpAddr>, declared: Option<&Destination>, hub: IpAddr) -> Result<Self, &'static str>;fn admits(&self, from: SocketAddr) -> bool;fn pin(&mut self, from: SocketAddr);fn client(&self) -> Option<SocketAddr>;}fn hears(hub: IpAddr, ip: IpAddr) -> bool; -
Comparison: both sides compare senders with
endpointinprotocols/src/socks/protocol.rs: the IP in canonical form, so an IPv4-mapped IPv6 address from a dual-stack socket is the IPv4 one, and the port. IPv6 flow info and scope are left out.pub(crate) fn endpoint(addr: SocketAddr) -> (IpAddr, u16); -
Client:
SocksUdpLinkinprotocols/src/socks/udp_link.rsskips any datagram whose source is not the relay address the server named, tested asendpoint(from) != endpoint(self.relay), so a dual-stack client socket still hears an IPv4 relay. Its ownUDP ASSOCIATEnames no source (all zeros, as RFC 1928 has a client do when it does not know), so a server that checks sources holds it to the control connection’s address and the port of its first datagram:impl<S> DatagramLink for SocksUdpLink<S>whereS: AsyncRead + AsyncWrite + Unpin,{type Addr = Destination;fn poll_recv_from(&mut self,cx: &mut Context<'_>,buf: &mut ReadBuf<'_>,) -> Poll<io::Result<Destination>>;}
A change to either side needs a negative test that sends from a third address and asserts nothing is delivered. The existing ones are udp_association_ignores_another_ip, udp_association_ignores_another_port_once_pinned, udp_association_is_not_widened_by_the_request and udp_link_ignores_datagrams_not_from_the_relay in protocols/tests/pipeline/socks.rs (the two that send from 127.0.0.2 run on Linux only). The same file pins the refusals with udp_association_over_a_unix_socket_needs_its_exact_source and udp_association_refuses_a_relay_that_cannot_hear_the_client, and the positive round trip is new_server_vs_new_client_udp. The unit tests in protocols/tests/unit/socks/server.rs pin ExpectedSender directly, the hears rules included through ExpectedSender::new, and endpoint_sees_through_ipv4_mapping_and_ignores_flow_info (protocols/tests/unit/socks/protocol.rs) pins the comparison. SOCKS describes the association in full.
Queues are bounded and backpressure reaches the producer
Section titled “Queues are bounded and backpressure reaches the producer”Every per-connection buffer or queue has a fixed upper bound. Moving items from a bounded channel into an unbounded collection to keep a producer running defeats the bound: when the consumer stalls, the stall must travel back to whoever is producing.
-
Mechanism:
ProxyServerRuntimeallocates exactly three buffers ofBUF_SIZEbytes per connection, once, and never grows them: the transport read buffer (ReadBuffer<const N: usize>), the transport staging buffer (WriteBuffer<const N: usize>), both inconcepts/src/buffer.rs, and an outbound scratch array. The contract inconcepts/src/runtime.rsis that a forward that cannot complete holds every effect behind it, and while a forward of transport bytes waits the transport is not read; an outbound is not read while the staging buffer lacksSTAGING_RESERVEplus one byte (plusMAX_DATAGRAMfor a datagram outbound). A client that stops reading therefore stops the downlink at the outbound socket. -
WireGuard driver: the driver task in
protocols/src/wireguard/device.rsmoves every connection’s bytes between bounded channels and smoltcp sockets, so besides the runtime’s buffers it keeps per-connection state of its own. Each connection’s way up is anUplink: the application’s channel, bounded atCHANNEL_CAP = 256items, and a slot for at most one item the socket has not accepted yet. The driver reads the channel only while that slot is empty (poll_refillstaysPendingwhile an item is held), so a remote or tunnel that stops draining one connection fills that connection’s socket (a TCP socket buffersTCP_BUFFER = 64 * 1024bytes each way), then its channel, and then blocks the application’s sender, while the other connections keep moving. It hands one socket at mostCHANNEL_CAPitems per pass, so no producer keeps the driver to itself. On the way down, a socket is read only while its channel has room. A UDP datagram that finds its association’s send ring full is held in the same slot until the next poll empties the ring; one that can never be sent (larger than the whole ring, or to an unaddressable endpoint) is dropped rather than retried.struct Uplink<T> {rx: Option<mpsc::Receiver<T>>,held: Option<T>,}impl<T> Uplink<T> {fn take(&mut self) -> Option<T>;fn hold(&mut self, item: T);fn poll_refill(&mut self, cx: &mut Context<'_>) -> Poll<()>;} -
Tests:
stalled_outbound_holds_uplink_but_not_other_downlink,frame_larger_than_the_buffer_is_an_erroranddatagram_outbound_is_never_truncated_by_staging_backpressure(concepts/tests/runtime.rs);backpressure_from_the_wire_reaches_the_writer(concepts/tests/client.rs);a_stalled_tcp_flow_blocks_its_writer(protocols/tests/pipeline/wireguard.rs), which checks that a flow whose remote reads nothing blocks its writer once the buffers above are full and does not hold up another flow on the same tunnel;a_held_item_keeps_the_channel_unread(protocols/tests/unit/wireguard/device.rs).
Session tables that grow with peer input carry an explicit cap too, for example MAX_SESSIONS = 256 per connection for Hysteria 2 UDP sessions (protocols/src/hysteria/server/datagrams.rs) and for mux sub-flows (protocols/src/mux/demux.rs).
A connection limit covers the whole relay
Section titled “A connection limit covers the whole relay”A semaphore that claims to bound active connections must hold its permit until the relay has finished. A permit that covers only the handshake, the decode or the transport stage bounds that stage, not active connections. At every spawn, check that the permit moves into the spawned task.
| Where | Constant | Held by |
|---|---|---|
App stream inbounds (app/src/serve.rs) |
MAX_LIVE_CONNECTIONS_PER_INBOUND = 65_536 |
An Arc<OwnedSemaphorePermit> taken in run_stream_inbound, cloned into every stream a transport yields, and bound as _session for the whole of serve_connection |
katana stream nodes (src/serve.rs, src/manager/proxy.rs) |
MAX_LIVE_CONNECTIONS_PER_NODE = 65_536 |
The same pattern: ProxyManager::accept_stream moves the session permit into the spawned task next to serve_stream |
Hysteria 2 circuits (protocols/src/hysteria/server/inbound.rs, datagrams.rs) |
circuit_permits: in the app, the inbound’s max_circuits, default DEFAULT_MAX_CIRCUITS = 65_536; in katana, MAX_LIVE_CONNECTIONS_PER_NODE |
A permit per proxy stream, bound as _permit inside the task that awaits the runtime, and one per UDP session, held by the session entry |
The stage limits next to them (MAX_HANDSHAKES_PER_INBOUND = 2048 in the app, MAX_TRANSPORT_STAGES_PER_NODE = 2048 and MAX_PREAUTH_STREAMS_PER_NODE = 512 in katana) are deliberately separate and cover only the handshake stage; they are not connection limits. the_circuit_limit_refuses_new_sessions (protocols/tests/unit/hysteria/server/datagrams.rs) pins the circuit cap for UDP sessions.
Dual-stack sends pick a usable address
Section titled “Dual-stack sends pick a usable address”When a name resolves to several addresses, the caller must pick one its local sockets can actually reach, and keep the others as fallbacks, rather than taking the first answer and dropping the packet when that family is unavailable.
-
Mechanism:
select_candidate_ipsinprotocols/src/helpers/address_family.rsfilters the resolver’s answer by the outbound’sAddressFamilyStrategyand byFamilySupport(which families the outbound can source traffic from:FamilySupport::both()for a dialer that leaves source selection to the kernel,FamilySupport::from_addrsfor a userspace netstack with its own interface addresses), and orders it with a stable sort so the other family stays as a fallback.DualStackUdpinenvironment/src/dial/udp.rsbinds one socket per requested family, leaves out a family whose bind fails, and sends each datagram from the socket of the peer’s family.pub fn select_candidate_ips(resolved: Vec<IpAddr>,strategy: AddressFamilyStrategy,support: FamilySupport,) -> Vec<IpAddr>;impl DualStackUdp {pub fn socket_for(&self, peer: &SocketAddr) -> io::Result<&UdpSocket>;} -
Tests:
auto_skips_families_without_a_local_address,prefer_ipv4_keeps_ipv6_as_fallbackandipv4_only_filters_to_ipv4(protocols/tests/unit/helpers/address_family.rs);v6_socket_is_v6_only_and_dual_stack_picks_by_family(environment/tests/integration/udp.rs).
Fixed fields are validated strictly
Section titled “Fixed fields are validated strictly”A parser checks every fixed field (version, reserved bytes, lengths, command and type codes) and refuses a value it does not know. Never assume the peer is a correct implementation, and never let a declared length allocate or read beyond a bound before it is checked. In environment and protocols the clippy denies above make unchecked indexing and unchecked arithmetic a lint error, so the code has to handle an out-of-range index or an overflowing length instead of panicking on it.
| Parser | Tests that pin the strictness |
|---|---|
| VLESS | rejects_bad_version, rejects_nonempty_addons, response_header_rejects_version (protocols/tests/unit/vless/protocol.rs) |
| VMess | rejects_unknown_version, rejects_unknown_command, rejects_unknown_security, rejects_invalid_option_relationship (protocols/tests/unit/vmess/protocol.rs) |
| Trojan | rejects_missing_crlf, rejects_oversize_udp_packet (protocols/tests/unit/trojan/protocol.rs) |
| mux.cool | rejects_unknown_status_and_network, rejects_truncated_metadata (protocols/tests/unit/mux/frame.rs) |
| gRPC framing | decoder_rejects_unexpected_field, decoder_rejects_an_oversized_length_before_buffering_the_body (protocols/tests/unit/transports/grpc_framing.rs) |
| DNS | truncated_and_malformed_input_never_panics, rejects_an_over_long_name_and_label (protocols/tests/unit/dns/message.rs) |
| Hysteria 2 | varint_above_62_bits_is_refused_not_panicked, tcp_request_rejects_a_foreign_frame_type (protocols/tests/unit/hysteria/protocol.rs) |
Configuration and reload
Section titled “Configuration and reload”Configuration fails closed
Section titled “Configuration fails closed”A typo must stop the program, not fall back to a default that changes what the configuration means. Stable configuration structs reject unknown fields, and a value that is not recognised is an error, not “anything else”.
-
Mechanism: the configuration structs carry
#[serde(deny_unknown_fields)](33 of them inapp/src/config.rs, 12 in katana’ssrc/config.rs).tls_layerinapp/src/transport.rsvalidatessecurityagainst the network instead of comparing it with"tls", so an unknown value is refused rather than becoming plaintext:pub fn tls_layer(network: &str, security: Option<&str>, ctx: &str) -> io::Result<bool>; -
Tests:
a_mistyped_key_is_rejected_rather_than_ignoredandan_inbound_defaults_to_loopback(app/tests/unit/config.rs);an_unknown_security_is_rejected_on_every_networkandtcp_with_tls_is_rejected_and_names_the_fix(app/tests/unit/transport.rs); in katana,reject_unknown_sni_is_refused_rather_than_ignored,acme_cert_mode_rejectedandhysteria_refuses_an_unknown_obfs_rather_than_disabling_it(tests/unit/inbound.rs) andan_invalid_address_family_is_rejected_for_every_protocol(tests/unit/outbound.rs).
The binaries show it directly. A misspelled key under [inbound.stream] (log timestamp removed):
$ etemenanki-app --test -c typo.tomlERROR etemenanki_app: configuration invalid: TOML parse error at line 7, column 1 |7 | securty = "tls" | ^^^^^^^unknown field `securty`, expected one of `network`, `security`, `tls`, `ws`, `grpc`The process exits with status 1 and binds nothing.
Pay particular attention to keys that decide exposure or security: listen address, port, protocol, network, security, TLS settings, outbounds and routes.
Reload is transactional
Section titled “Reload is transactional”A reload must not tear down the running configuration and then discover that the new one cannot start. Parse, validate and build first; switch only when that succeeded; keep the old state when it did not. A failed reload must also be retryable once the outside cause (a busy port, a missing file) is gone, not only when the file’s bytes change.
flowchart LR read["read file"] --> parse["parse_bytes"] parse -->|error| keep["keep old generation"] parse --> build["build: router, inbounds, outbounds"] build -->|error| keep build --> cancel["cancel old generation"] cancel --> spawn["spawn_generation(built, false)"]
-
App mechanism:
Instance::reloadinapp/src/instance.rsrunsconfig::parse_bytesandbuildbefore touching the running generation, and logsreload: parse failed, keeping current configorreload: build failed, keeping current configon failure. Only then does it cancel the old generation’s token, wait for its accept loops to release their listeners, and callspawn_generation(built, false), where a single inbound’s bind failure is logged and the other inbounds still start. The swap drops every in-flight connection of the old generation.pub fn build(cfg: &Config) -> io::Result<Built>;async fn spawn_generation(built: Built, strict: bool) -> io::Result<Generation>;impl Instance {pub async fn reload(&self);} -
Where the app falls short at the verified revision: the new listeners are bound only after the old generation has released its ports, so an inbound whose bind fails on reload stays down; the old generation is already gone by then.
Instance::reloadalso returns early when the file’s bytes equal the last bytes it tried (a parse or build failure records them too), so an unchanged file is not tried again after the outside cause is fixed. Edit the file or restart the process to retry. A change to this path should close these gaps rather than widen them. -
katana mechanism: a file that does not load is refused first, in the watcher loop, with
config reload failed, keeping current.apply_reloadinsrc/runtime.rsthen builds everything the reload needs before touching anything: the new outbound pool when[[outbound]]changed, every node the reload adds (its panel client and its router; nothing is bound or spawned yet), and the panel client of every running node whose entry changed. If any of these fails, it applies nothing and logsreload: bad outbounds, keeping current configorreload: node <display_id>: <error>; keeping current config. Only then are node changes applied, node by node: a removed or re-identified node is stopped, a new one is spawned from what was already built, and a changed one is sent the edit. The running node takes that edit whole or not at all: a route that does not compile or a panel client that does not build is refused withnode <id>: config edit refused, keeping the running one: <error>. A node the reload adds or respawns that fails to come up, for example on a port that is still taken, keeps retrying as described under Node tasks are supervised, so it recovers once the outside cause is gone; a running node whose listener rebuild fails tries again at every poll. -
Tests:
a_reload_rebinds_the_udp_port(app/tests/integration/e2e_hysteria_inbound.rs) checks that a Hysteria 2 inbound gets its UDP port back across a reload while the old generation holds a live connection;a_reload_with_a_node_that_does_not_build_changes_nothing(katanatests/unit/runtime.rs) checks that an entry with a mistypednode_typeleaves the running node, its live connection and the route edit made alongside untouched, and the other tests in that file pin which edits respawn a node and which reach it in place;route_change_drops_connectionsandrepeated_user_refreshes_never_disturb_a_live_connection(katanatests/unit/e2e.rs) pin what a katana reload and a panel refresh keep. Generations and reload and Runtime and reload describe both paths in full.
katana
Section titled “katana”Node tasks are supervised
Section titled “Node tasks are supervised”A transient failure of the first panel request, of DNS, of a bind or of bring-up must not end a node’s task for good with nothing to restart it. A long-lived node task needs retry with backoff under a supervisor, and must still stop promptly on cancellation.
-
Bootstrap:
NodeManager::bootstrapinsrc/manager/node.rscallstry_bootstrapuntil the node is up or itsCancellationTokenis cancelled. An attempt fails whennode_infofails, returns nothing or reports port 0, whenuser_listfails or returns nothing (a node up without its users would serve nobody), or whenbring_upfails, for example on a port that is still taken. Each attempt first clears the panel client’s ETags, so a panel that answers a repeated request with 304 Not Modified still gives it a full answer. After a failure the node logsnode <id>: <reason>; retrying in <n>sand waitsBOOTSTRAP_RETRY_MIN(1 s), doubling with each failure up toBOOTSTRAP_RETRY_MAX(60 s), or up to the poll period (controller.update_periodic) when that is shorter. While it waits it applies static updates as they arrive, which keeps the runtime’s bounded channel to it from filling, and a config edit ends the wait at once, since the edit may be the fix. Cancellation is honoured both during an attempt and during the wait.const BOOTSTRAP_RETRY_MIN: Duration = Duration::from_secs(1);const BOOTSTRAP_RETRY_MAX: Duration = Duration::from_secs(60);impl NodeManager {pub async fn run(self: Arc<Self>, shutdown: CancellationToken, static_rx: mpsc::Receiver<StaticUpdate>);async fn bootstrap(&self, shutdown: &CancellationToken, static_rx: &mut mpsc::Receiver<StaticUpdate>) -> bool;async fn try_bootstrap(&self) -> Result<(), String>;} -
Steady state:
NodeManager::poll_cyclefalls back to the last applied node info when a panel request fails, returns nothing or reports port 0, and to the last applied user list when that request fails or returns nothing; it logs a failed request or a port of 0 and tries again on the next tick. The loop exits only when itsCancellationTokenis cancelled. Whether the node stopped in the steady state or while still bootstrapping,runthen stops the listener withtear_downand flushes the final traffic and audit reports. -
Tests:
a_node_comes_up_once_the_panel_answers,a_node_whose_port_is_taken_comes_up_once_it_is_free(against a panel that answers repeated requests with 304) anda_node_that_never_bootstraps_still_stops(tests/unit/e2e.rs). Node manager describes the whole loop.
Reload rebuilds whatever caches configuration
Section titled “Reload rebuilds whatever caches configuration”If an object caches configuration fields when it is built, a reload that changes those fields must rebuild or update that object. Updating the manager’s copy of the config while the client keeps the old one is the failure this rule exists for.
-
Identity: a node’s identity in
src/runtime.rsis the tuple of fields that decide which panel node it serves. A reload that changes the identity removes the node and spawns a new one, with a newPanelClientand a fresh traffic registry, so one panel node’s counters are never billed to another; a reload that keeps it sendsStaticUpdate::Configto the running node.type NodeId = (String, String, u32, String, String);fn identity(cfg: &NodeConfig) -> NodeId;pub fn panel_node_type(cfg: &NodeConfig) -> String;The tuple is
(panel_type, api.host, api.node_id, api.key, panel node type), withpanel_typelowercased. The panel node type comes frompanel_node_typeinsrc/api/mod.rs: newV2board finds a node by its id and thenode_typeit is asked for, so there it is"vless"for a V2ray-family node (node_typeV2ray,VmessorVless) withenable_vless, else the lowercasedapi.node_type; SSPanel finds a node by id alone, so there it is empty.display_idshows the node type after the id when it is not empty. -
Cached fields:
Client::newinsrc/api/newv2board.rsandsrc/api/sspanel.rscopies every field it needs frompanel_typeand[node.api]into the client, including the timeout,api.node_type,api.enable_vless,api.speed_limitandapi.rule_list_path.NodeManager::apply_staticinsrc/manager/node.rstherefore builds a newPanelClientwheneverpanel_typeor anything in[node.api]changed, before it stores the edit. The new client takes over the routes the old one read from a newV2board node config (PanelClient::inherit_routes), so audit rules derived from them keep applying until it reads the config itself. It holds no ETags, so the node then runs one full panel poll with it and takes the narrowest teardown the answer needs. An edit such asapi.speed_limitorapi.rule_list_pathkeeps every connection unless the panel’s new answer changes the protocol or transport; an edit to a local setting the listener is built from (controller.listen_ip,controller.cert,controller.disable_sniffing,api.enable_vless,[node.hysteria]or[node.route]) forces a new listener generation. -
What to check: a new field a client caches is covered only if it is inside
[node.api]orpanel_type; otherwise it must join that comparison inapply_static, or be pushed to the client on reload. A field that decides which panel node is served belongs in the identity. -
Tests:
an_sspanel_node_is_its_panel_node_id,a_newv2board_node_is_also_the_type_it_asks_for,a_node_is_logged_by_its_panel_node_not_its_key,a_newv2board_type_edit_respawns_the_node,an_sspanel_api_edit_takes_effect_in_placeanda_client_edit_takes_effect_without_dropping_connections(tests/unit/runtime.rs);a_rebuilt_client_keeps_the_routes_it_has_not_read(tests/unit/api/newv2board.rs).
Rate limiting is byte-accurate
Section titled “Rate limiting is byte-accurate”A token bucket charges the actual number of bytes and produces a wait or a debt in proportion. A chunk larger than the bucket’s capacity must not pass after a single refill period, and the achieved rate must not depend on the relay’s buffer or chunk size.
-
Mechanism:
TokenBucketinsrc/traffic.rsrefills atratebytes per second and banks at most one second of it (arateof 0 is unlimited).chargealways takes the full amount, letting the balance go negative, and returns the instant the debt is repaid;Gateinsrc/meter.rscharges each transfer after it moves and holds both directions of every flow of that user back while the shared bucket is in debt.impl TokenBucket {pub fn new(rate: u64) -> Self;pub fn charge(&self, n: usize) -> Option<Instant>;pub fn ready_at(&self) -> Option<Instant>;}pub fn determine_rate(node_bps: u64, user_bps: u64) -> u64; -
Tests:
token_bucket_rate_limits,a_chunk_larger_than_the_burst_is_still_limited,debt_holds_back_the_next_charge_tooandan_idle_bucket_banks_one_second_and_no_more(tests/unit/traffic.rs);the_limit_holds_however_the_writes_are_sizedandthe_limit_is_shared_by_both_directions(tests/unit/meter.rs). All buttoken_bucket_rate_limits, which uses the real clock and a tolerance, run on a paused tokio clock, so their timings are exact.
A limit on active authenticated connections, per node or per user, would need admission control that covers the whole relay in the same way as the connection limits above; the pre-auth limit covers only the handshake stage. See Admission.
No credentials in logs
Section titled “No credentials in logs”Never log a credential, UUID, token, key, panel secret, or a URL whose query carries one. When a credential is malformed, log a safe identifier and the kind of error, never the value.
-
Mechanism:
error_for_statusinsrc/api/mod.rsreplaces reqwest’s own status check and strips the URL, because the panel key travels in the query string.display_idinsrc/runtime.rsrenders a node as panel type, host, node id and, on newV2board, the panel node type, neverapi.key.NodeManager::try_bootstraprenders a failed panel request by its outermost error message only, because the cause beneath carries the request URL and its query carries the key. Outbound builders refuse a malformed UUID without quoting it.pub fn error_for_status(resp: reqwest::Response) -> Result<reqwest::Response>;fn display_id(id: &NodeId) -> String; -
Tests:
a_malformed_uuid_is_refused_without_echoing_it(katanatests/unit/outbound.rs);a_node_is_logged_by_its_panel_node_not_its_key(katanatests/unit/runtime.rs);an_unencodable_password_is_refused_without_quoting_it(Etemenankiprotocols/tests/unit/hysteria/auth.rs).
A change to HTTP error handling in a panel client must keep URLs and tokens out of every error it can return.
Traffic accounting is exact
Section titled “Traffic accounting is exact”Five properties must survive every change to the traffic code:
| Property | Mechanism (src/traffic.rs, src/manager/node.rs) |
Test (tests/unit/traffic.rs) |
|---|---|---|
| A failed report loses nothing | Live counters are only reduced after success; residual rows go back through restore_residuals on failure |
restored_residuals_are_retried |
| A retry does not bill twice | commit_reported subtracts exactly what was reported; a residual or fully drained row leaves the table when it is snapshotted |
rate_change_drains_old_counter_and_reports_once |
| Residuals are recoverable | A departed user’s bytes become residual rows keyed by uid | dropped_user_bytes_become_residuals, a_departed_users_late_bytes_are_still_reported |
| Removal and re-adding are well defined | A removed user leaves the table; a credential rebound to another uid gets a fresh counter, and the old uid’s bytes are reported under the old uid | set_users_drops_absent, rebound_credential_reports_the_old_uid_separately |
| Concurrent increments are kept | commit_reported uses fetch_sub instead of resetting to zero |
commit_reported_preserves_concurrent |
sequenceDiagram
participant N as NodeManager
participant T as NodeTraffic
participant P as Panel
N->>T: snapshot()
T-->>N: live rows, draining rows, residual rows
N->>P: report_user_traffic(non-zero rows merged by uid)
alt report succeeded
N->>T: commit_reported(up, down) per row with a counter
else report failed
N->>T: restore_residuals(rows without a counter)
end
impl UserCounter { pub fn commit_reported(&self, up: u64, down: u64);}impl NodeTraffic { pub fn snapshot(&self) -> Vec<TrafficSnapshot>; pub fn restore_residuals(&self, rows: Vec<(i64, u64, u64)>);}End to end, two tests in tests/unit/e2e.rs run a node against a fake panel: vmess_traffic_is_metered_and_reported checks the totals the panel receives, and a_retired_user_stops_while_the_rest_keep_their_connections checks that removing one user stops that user’s connections and leaves the others running.
Route model
Section titled “Route model”The route model in environment/src/routing.rs keeps first-match semantics: the first rule that matches wins, and a target no rule matches goes to the default outbound. Changes to matchers or geodata loading are reviewed for case normalisation, regex validation, CIDR prefix validation, malformed protobuf input and rule order. A port range whose lower bound exceeds its upper bound is refused by parse_port_match rather than compiled into a rule that never matches:
pub fn parse_port_match(s: &str) -> io::Result<RouteMatch>;port_parser_rejects_an_inverted_range and port_ranges_are_inclusive (environment/tests/unit/routing.rs) pin it. The same review applies to harranu, which carries the same model.
Versioning
Section titled “Versioning”Every published crate follows SemVer as Cargo interprets it. Cargo reads a bare requirement such as version = "2.0.0" as a caret requirement, and what “compatible” means depends on whether the version is below 1.0:
| Version series | Requirement ^X.Y.Z accepts |
Breaking change | Compatible addition | Compatible fix |
|---|---|---|---|---|
0.y.z with y >= 1 (before 1.0) |
>=0.y.z, <0.(y+1).0 |
Bump y |
Bump z |
Bump z |
x.y.z with x >= 1 |
>=x.y.z, <(x+1).0.0 |
Bump x |
Bump y |
Bump z |
etemenanki-concepts, etemenanki-protocols and etemenanki-app each moved through the 0.y series at their own pace until they were released together as 1.0.0. All four kernel crates moved to 2.0.0 together (etemenanki-environment was first published at that number), a major release because the task-and-channel pipeline was removed rather than kept alongside the new one. At the verified revision etemenanki-protocols is at 2.0.2 (after the 2.0.1 patch), and the other three are still at 2.0.0. katana is at 3.0.1: it took a major version when it moved to kernel 1.0.0 (katana 2.0.0) and again for kernel 2.0.0 (katana 3.0.0), and 3.0.1 is a patch release that fixes node bootstrap and reload and moves its lockfile to etemenanki-protocols 2.0.1 while its Cargo.toml keeps the 2.0.0 requirement.
What counts as breaking:
- Library crates: removing or renaming a public item, changing a public signature, a trait’s required items or a public type’s fields, or changing a feature’s meaning. Adding a public item or an optional feature is a compatible addition.
- Behaviour: a change in what existing configuration means, what goes on the wire, or what a caller can observe is judged like an API change, even when every signature stays the same. For
etemenanki-appand katana, removing a configuration key or changing its meaning is breaking.
etemenanki-protocols 2.0.1 is one recorded exception. It replaced Demux::feed_whole with feed_chunks; etemenanki_protocols::mux and Demux are public, so by the rule above that removal is breaking. The release shipped as a patch because nothing outside the crate called the method, and the release commit’s body says so. That judgement does not change the rule: a patch that removes a public item needs the same check (search katana and every other dependent for uses) and the same written justification, and when any dependent might use the item, the release is breaking.
etemenanki-protocols 2.0.2 also shipped as a patch, and its public API is unchanged: the source a UDP ASSOCIATE names reaches the SOCKS inbound through the crate-private handshake_with_udp_source in protocols/src/socks/handshake.rs, while the public handshake keeps its signature. The release commit’s body records the behaviour that changed with it: a client over a Unix socket whose request names no exact source is now refused, and so is an association whose relay could never hear its client.
If you cannot tell whether a change is compatible, say so in review and settle it before anything is published. Bump only the crates that changed; a crate whose code and dependency requirements did not change keeps its version.
Releasing
Section titled “Releasing”A release is a deliberate act with outside effects (a registry upload, pushed commits, a tag and the binaries built from it), so it happens only when a release has actually been decided, never as a side effect of landing a change.
Crates are published in dependency order, and katana moves only after the kernel crates it needs are visible in the registry:
flowchart LR c["etemenanki-concepts"] --> e["etemenanki-environment"] e --> p["etemenanki-protocols"] p --> a["etemenanki-app"] p --> k["katana: update lockfile, bump, tag vX.Y.Z"] e --> k c --> k
Skip every crate that was not bumped. When a lower crate changes, check each crate above it:
| Changed | Check the manifests of | Also check |
|---|---|---|
etemenanki-concepts |
environment/Cargo.toml, protocols/Cargo.toml, app/Cargo.toml |
Source changes needed in each |
etemenanki-environment |
protocols/Cargo.toml, app/Cargo.toml |
Source changes needed in each |
etemenanki-protocols |
app/Cargo.toml |
Source changes needed in the app |
The inter-crate dependencies are written as path dependencies with a version and a registry, so the published manifest points at the registry version. If a crate needs something only a newer lower crate provides, raise its requirement to that version. Either way, confirm in Cargo.lock that the workspace resolves to the version you are about to publish.
Kernel release
Section titled “Kernel release”-
Preflight. Read the live state of Etemenanki (and katana, if it will follow) as in Before you start, and read
.cargo/config.tomlfor the registry and its credential provider. The registry token is never in the repository; it comes from the environment or Cargo’s credential store. If the branch must be synced first, use a fast-forward only pull (see Git rules) and stop if it cannot fast-forward. -
Decide the scope. List what actually changed with
git status --short,git diff --name-onlyandgit diff --cached --name-only, and map paths to packages:concepts/**,environment/**,protocols/**,app/**. Root manifest, feature and lockfile changes are judged by which crates they affect. -
Choose the versions by the rules in Versioning, and edit only the bumped crates’
version, plus theversionrequirement other workspace crates declare on them where it must rise. The workspace members’ own entries inCargo.lockfollow on the next Cargo command run without--locked(the gatecargo test --workspaceis one);cargo check --workspace --lockedfails until they have. Check thatCargo.lockchanged only in those crates’ entries. -
Run every Etemenanki gate, including
cargo check --workspace --locked. A failure stops the release. -
Commit. The message names what is released:
release etemenanki-protocols X.Y.Zfor one crate,release etemenanki-protocols X.Y.Z, etemenanki-app X.Y.Zfor several, orrelease etemenanki X.Y.Zwhen all four move together. The body says why the bump has the size it has. Checkgit statusandgit show --statafterwards. -
Publish in order, one bumped crate at a time, and confirm each exact version is visible before the next:
Terminal window cargo publish -p <package> --registry <registry>cargo info --registry <registry> <package>@<version>Each crate’s manifest has
publish = [...]naming the one registry it may go to. The index can lag for a moment, so retryingcargo infois reasonable. -
Push the branch with
git push origin <branch>. Etemenanki has no release tags; its release commits carry the versions. Do not start a tag scheme.
katana release
Section titled “katana release”-
Update the kernel dependencies that the release needs, once they are visible in the registry, one package at a time:
Terminal window cargo update -p <package> --precise <version>Raise the requirement in
Cargo.tomltoo if katana needs the new version’s API or behaviour. Then checkCargo.lock: each updated crate resolves to the version just published, itssourceline is the private registry’s index and it carries achecksum, and nothing else moved. -
Bump katana’s own version in
Cargo.toml, read from the file rather than remembered. A compatible fix is a patch; follow the history for anything larger (katana has taken a major version each time it moved to a new kernel major). Update thekatanaentry inCargo.locktoo, with any Cargo command run without--lockedsuch ascargo check: the--lockedgates refuse a lockfile that no longer matchesCargo.toml. Check that nothing else in the lockfile moved. -
Run every katana gate, including
cargo metadata --locked --no-deps --format-version 1andcargo check --locked. -
Commit and tag. The history uses a commit titled
release vX.Y.Zand an annotated tagvX.Y.Zwith the same message:Terminal window git commit -m "release vX.Y.Z"git tag -a vX.Y.Z -m "release vX.Y.Z"git rev-parse HEADgit rev-parse 'vX.Y.Z^{commit}'The tags are annotated, so compare
HEADwith the tag’s peeled commitvX.Y.Z^{commit}; a plaingit rev-parse vX.Y.Zprints the tag object instead. -
Push the branch, then the tag:
Terminal window git push origin <branch>git push origin vX.Y.ZThe tag push starts
release.yml, which buildsx86_64-unknown-linux-gnuandx86_64-unknown-linux-muslbinaries with--locked, fails if either links OpenSSL dynamically or if the musl binary is not fully static, and publisheskatana-vX.Y.Z-linux-x86_64-gnu,katana-vX.Y.Z-linux-x86_64-musland a.sha256for each as the release’s assets.
If the tag already exists and points at another commit, do not move or force-update it.
After a release
Section titled “After a release”In every repository you changed, confirm that local and remote agree:
git status --short --branchgit rev-parse HEADgit ls-remote origin refs/heads/<branch>git ls-remote origin 'refs/tags/vX.Y.Z*' # katana: the peeled ^{} line is the release commitcargo info --registry <registry> <package>@<version> # every crate publishedFinally, check that katana’s Cargo.lock resolves every etemenanki-* crate to the version this release published, and bring these pages up to the new release as described in Keeping these docs current.
What stops a release
Section titled “What stops a release”Stop, and report the actual error, when any of these happens. Do not work around it.
| Situation | Not acceptable as a workaround |
|---|---|
| The branch cannot fast-forward, or has diverged from the remote | Merging or rebasing and carrying on |
| Registry authentication fails, or the registry configuration is missing | Publishing somewhere else |
| A publish reports the version exists | Overwriting it, or bumping until it goes through |
| A dependency does not resolve, or the lockfile moves far more than expected | Deleting Cargo.lock or running a blanket cargo update |
| Any gate fails | Skipping the test or lowering the clippy level |
| The new version does not appear in the registry | Moving katana to it anyway |
| A tag exists and points at another commit | Moving the tag |
| SemVer compatibility or the tag convention is unclear | Guessing |
Git rules
Section titled “Git rules”Uncommitted work belongs to whoever made it and outranks automation.
- Never run
git reset --hard,git checkout -- <file>or other destructive restores, and do not rungit cleancasually. - Never delete, overwrite or hide someone else’s uncommitted changes to make a command or a test pass.
- Never force-push a branch, and never rewrite history that has been pushed.
- Check
git statusbefore syncing with the remote. When local changes exist, sync withgit pull --ff-only --autostash. If the pull cannot fast-forward, stop: merging or rebasing is a decision for a person, not a release step. - Tags are never moved once pushed.
Keeping these docs current
Section titled “Keeping these docs current”Every developer page lists, in its frontmatter, the repository files it describes and the revision it was last checked against:
sources: Etemenanki: [concepts/src/core.rs, concepts/src/runtime.rs] katana: [src/traffic.rs]verified: Etemenanki: 596916d katana: v3.0.1In the docs repository, pnpm check:drift (scripts/check-drift.mjs) reads that frontmatter from every page and reports a problem when:
- a repository listed under
sourceshas noverifiedentry, or the revision does not exist in the sibling checkout; - a listed file does not exist at the verified revision;
- any listed file has commits between the verified revision and the released revision recorded for that repository in the docs repository’s
docs-versions.json(it prints them). The comparison deliberately stops at the release rather than the checkout’sHEAD, which may hold unreleased work; a repository missing from that file is compared up toHEAD; - a translated page’s
sourcesorverifieddiffer from its English page.
It ends with check-drift: N tracked page(s), M problem(s) and exits with status 1 when there is any problem. After each release, move that repository’s ref in docs-versions.json to the release commit or tag and run the check. For every page it reports, re-read the code, update the prose, move verified to the new revision and update the Chinese page’s frontmatter in the same commit. A change to configuration keys also means updating the field tables under src/data/reference/ and the examples under examples/, which pnpm check:examples runs through the real binaries with --test; pnpm build then fails on any broken link or anchor.
The site is public and the code repositories are private. Pages name code by repository-relative path and symbol, never by line number or link, and they never contain real secrets, registry names or URLs, or anything about a particular deployment.