Skip to content

Contributing

Source files: 75 · checked against Etemenanki 596916d · katana v3.0.1
  • Etemenanki/Cargo.toml
  • Etemenanki/Cargo.lock
  • Etemenanki/.cargo/config.toml
  • Etemenanki/.github/workflows/build.yml
  • Etemenanki/README.md
  • Etemenanki/concepts/Cargo.toml
  • Etemenanki/environment/Cargo.toml
  • Etemenanki/protocols/Cargo.toml
  • Etemenanki/app/Cargo.toml
  • Etemenanki/concepts/src/buffer.rs
  • Etemenanki/concepts/src/runtime.rs
  • Etemenanki/environment/src/dial/udp.rs
  • Etemenanki/environment/src/routing.rs
  • Etemenanki/protocols/src/socks/udp_link.rs
  • Etemenanki/protocols/src/socks/server.rs
  • Etemenanki/protocols/src/socks/protocol.rs
  • Etemenanki/protocols/src/socks/handshake.rs
  • Etemenanki/protocols/src/helpers/address_family.rs
  • Etemenanki/protocols/src/mux/demux.rs
  • Etemenanki/protocols/src/wireguard/device.rs
  • Etemenanki/protocols/src/hysteria/server/config.rs
  • Etemenanki/protocols/src/hysteria/server/inbound.rs
  • Etemenanki/protocols/src/hysteria/server/datagrams.rs
  • Etemenanki/app/src/serve.rs
  • Etemenanki/app/src/transport.rs
  • Etemenanki/app/src/config.rs
  • Etemenanki/app/src/instance.rs
  • Etemenanki/concepts/tests/runtime.rs
  • Etemenanki/concepts/tests/client.rs
  • Etemenanki/protocols/tests/pipeline/socks.rs
  • Etemenanki/protocols/tests/unit/socks/server.rs
  • Etemenanki/protocols/tests/unit/socks/protocol.rs
  • Etemenanki/protocols/tests/pipeline/wireguard.rs
  • Etemenanki/protocols/tests/unit/wireguard/device.rs
  • Etemenanki/environment/tests/integration/udp.rs
  • Etemenanki/environment/tests/unit/routing.rs
  • Etemenanki/protocols/tests/unit/helpers/address_family.rs
  • Etemenanki/protocols/tests/unit/vless/protocol.rs
  • Etemenanki/protocols/tests/unit/vmess/protocol.rs
  • Etemenanki/protocols/tests/unit/trojan/protocol.rs
  • Etemenanki/protocols/tests/unit/mux/frame.rs
  • Etemenanki/protocols/tests/unit/transports/grpc_framing.rs
  • Etemenanki/protocols/tests/unit/dns/message.rs
  • Etemenanki/protocols/tests/unit/hysteria/protocol.rs
  • Etemenanki/protocols/tests/unit/hysteria/auth.rs
  • Etemenanki/protocols/tests/unit/hysteria/server/datagrams.rs
  • Etemenanki/app/tests/unit/config.rs
  • Etemenanki/app/tests/unit/transport.rs
  • Etemenanki/app/tests/integration/e2e_hysteria_inbound.rs
  • Etemenanki/app/tests/support/mod.rs
  • katana/Cargo.toml
  • katana/Cargo.lock
  • katana/.cargo/config.toml
  • katana/.github/workflows/build.yml
  • katana/.github/workflows/release.yml
  • katana/src/config.rs
  • katana/src/serve.rs
  • katana/src/manager/proxy.rs
  • katana/src/manager/node.rs
  • katana/src/runtime.rs
  • katana/src/api/mod.rs
  • katana/src/api/newv2board.rs
  • katana/src/api/sspanel.rs
  • katana/src/manager/mod.rs
  • katana/src/traffic.rs
  • katana/src/meter.rs
  • katana/tests/unit/traffic.rs
  • katana/tests/unit/meter.rs
  • katana/tests/unit/serve.rs
  • katana/tests/unit/outbound.rs
  • katana/tests/unit/inbound.rs
  • katana/tests/unit/e2e.rs
  • katana/tests/unit/runtime.rs
  • katana/tests/unit/api/newv2board.rs
  • katana/tests/support/mod.rs

This page is the working agreement for anyone who changes the kernel or katana. It says which repository a change belongs in, which commands must pass before it lands, what every reviewer checks regardless of the change, how versions are chosen, in which order crates are released, which git operations are off limits, and how these documentation pages are kept in step with the code.

The rules are stable; the numbers are not. Versions, branches, tags and dependency resolutions change with every release, so every step below starts by reading them from the repository rather than from this page. Where this page and a repository disagree, the repository is right and this page is due for an update (see Keeping these docs current).

The code lives in separate git repositories. None of them is a parent of the others, and there is no Cargo workspace spanning them: run every git, Cargo, test and release command inside the repository you are changing.

You are changing Repository Path Package
The sans-I/O core contract, the per-connection runtimes, links, connectors, shared net types Etemenanki concepts/ etemenanki-concepts
Host dialers (TCP, UDP, QUIC), the shared socket policy, the route model Etemenanki environment/ etemenanki-environment
A protocol, a transport, mux, sniffing, DNS Etemenanki protocols/ etemenanki-protocols
TOML configuration, inbound and outbound construction, router composition, generations and hot reload of the standalone proxy Etemenanki app/ etemenanki-app
Panel clients, node lifecycle, admission, traffic accounting, speed limits, audit, katana’s own configuration and reload katana src/ katana

Three boundaries catch people out:

  • Routing semantics belong in environment/src/routing.rs. The route model was vendored unchanged from the standalone routing crate (harranu) as etemenanki_environment::routing, and both the app and katana route with that copy. A change made only in harranu reaches neither. When harranu itself needs a change, it is released on its own, never as part of a kernel release; weigh whether the same change also belongs in environment/.
  • Reference trees are not implementations. Xray-core/ and hysteria/ in Etemenanki are upstream sources (git submodules) kept for protocol reference and built by the interop tests of both repositories; Xboard/, V2bX/ and XrayR/ in katana are kept for panel reference. A bug that shows up against them is fixed in the Rust code, never by editing the reference tree.
  • katana sees only published kernels. katana depends on etemenanki-concepts, etemenanki-environment and etemenanki-protocols by version from a private Cargo registry. A kernel fix reaches katana only after the affected crates are published and katana’s lockfile is moved to them. Workspace and crates has the details, including which features katana enables (vendored-openssl and hysteria on etemenanki-protocols, no tun).

Files ignored by .gitignore (scratch copies, build output) are never authoritative. Review and change only tracked sources.

  1. Read the live state of the repository. Every later decision depends on it.

    Terminal window
    git status --short --branch
    git branch --show-current
    git remote -v
    git log --oneline --decorate --max-count=12
    git tag --sort=-v:refname | head -20
  2. Read the manifests. Cargo.toml, Cargo.lock and .cargo/config.toml say which versions exist, which registry a crate publishes to and how the registry authenticates. For the package graph, use cargo metadata --no-deps --format-version 1. The version a dependency actually resolves to is in Cargo.lock or cargo metadata, not in the requirement in Cargo.toml.

  3. Read the call chain, not just the function. For a change to a public item in concepts, environment or protocols, search katana for every use of it: a signature change there is a breaking release (see Versioning).

  4. Walk the failure paths. For configuration, protocol, network or lifecycle code, ask what happens on malformed input, on a peer that stops reading, on cancellation, and when a dependency (DNS, a port, the panel) is not there. The review checklist lists what has gone wrong before.

Fix the cause, not the test. Keep the change as small as it can be while still complete, leave unrelated code alone, and do not add a dependency the change does not need.

A change lands only when its repository’s gates pass. The release gates are run in addition, immediately before publishing or tagging.

Terminal window
cargo fmt --all -- --check
cargo test --workspace
cargo clippy --workspace --all-targets --all-features -- -D warnings
# before a release, additionally:
cargo check --workspace --locked
Gate What it protects
cargo fmt … --check One formatting, so diffs show only real changes
cargo test Behaviour, including the tests named in the checklist below
cargo clippy … -D warnings Every warning is an error. In etemenanki-environment and etemenanki-protocols this is also what enforces the crate-wide denies on unwrap_used, expect_used, indexing_slicing and arithmetic_side_effects, which keep untrusted bytes from panicking the process; cargo build and cargo test do not check them
--all-features Lints the code behind features no workspace member turns on: quic on etemenanki-environment and vendored-openssl on etemenanki-protocols. (hysteria and tun are compiled anyway, because etemenanki-app enables them.)
--locked Fails instead of silently changing Cargo.lock, so what is tested is what is released

Neither repository’s CI runs these gates. Etemenanki’s build.yml only builds the etemenanki-app release binary, and katana’s build.yml only builds the katana release binary with --locked, so a green CI run says nothing about tests or clippy. Run the gates locally.

A green cargo test can still have skipped things. The Xray and Hysteria interop tests build the upstream binaries with go from the reference trees. When go is missing or the build fails, the support code prints a SKIP: line and each test that needs the binary returns early and passes, some after printing skipping: hysteria binary unavailable (app/tests/support/mod.rs, and katana’s tests/support/mod.rs, which looks for the trees in the sibling ../Etemenanki checkout). Changes to a protocol, a transport, TLS, SOCKS or WireGuard should be checked with those interop tests actually running. Testing lists what each suite needs.

Before asking for review, check the diff itself:

Terminal window
git diff --check
git diff --stat
git diff

Look for unrelated files, leftover debug output, temporary or generated files, a lockfile that moved more than the change explains, and anything that looks like a credential, token or private key.

These invariants come from the architecture and from past reviews. They are not a list of open bugs. Each one describes a property the code must have and must not lose. Where the verified revision does not fully meet an item yet, the item says so, and a change in that area should move towards the rule, not away from it. Apply every item that touches the code under review, and add a test for any mechanism a change introduces or moves.

UDP datagrams come from the association’s peer

Section titled “UDP datagrams come from the association’s peer”

A UDP relay must only accept datagrams from the peer it is serving, and a client must only accept replies from the relay it associated with. A packet from any other address is dropped, never relayed or delivered.

  • Server: the SOCKS server’s UDP association loop in protocols/src/socks/server.rs hears one client (RFC 1928 §7), described by an ExpectedSender that is built before the relay socket is bound. Over TCP the client is the control connection’s peer: admits requires every datagram’s IP to be that peer’s, and a datagram from any other IP is dropped unread. The first datagram that parses and carries a payload pins the port (pin; the first pin holds), so a neighbour on the same address cannot claim the association by sending junk first. A UDP ASSOCIATE whose DST.ADDR names the peer’s own IP with a non-zero port pins that port up front. A request that names anything else (another address, an unspecified address, a domain) is set aside rather than trusted or refused, because it is not where the datagrams come from: a client behind NAT names its LAN address, and sing-box names a loopback address whenever its first target is private. Replies go only to the pinned client, in the form the relay socket saw it. Over a Unix socket there is no peer, so the request must name its exact IP and a non-zero port. When there is no client the relay could hold to, the server answers 0x02 (STATUS_NOT_ALLOWED, “connection not allowed by ruleset”) and closes the control connection. That happens for a Unix-socket request that names a domain or leaves the address or the port at zero, and for a relay (udp_bind) in an address family the client’s IP is not in: hears lets an IPv4 or IPv4-mapped relay hear only IPv4, the unspecified IPv6 relay :: hear both families, and any other IPv6 relay hear only IPv6. An association never outlives its control connection: EOF or an error on the TCP stream ends it, and so does RELAY_IDLE_TIMEOUT without traffic. Datagrams with a non-zero FRAG byte are discarded by parse_udp_packet in protocols/src/socks/protocol.rs.

    struct ExpectedSender {
    ip: IpAddr, // canonical; every datagram must come from it
    port: Option<u16>, // when the request named it alongside `ip`
    client: Option<SocketAddr>, // the first sender forwarded for; replies go here
    }
    impl ExpectedSender {
    fn new(peer: Option<IpAddr>, declared: Option<&Destination>, hub: IpAddr) -> Result<Self, &'static str>;
    fn admits(&self, from: SocketAddr) -> bool;
    fn pin(&mut self, from: SocketAddr);
    fn client(&self) -> Option<SocketAddr>;
    }
    fn hears(hub: IpAddr, ip: IpAddr) -> bool;
  • Comparison: both sides compare senders with endpoint in protocols/src/socks/protocol.rs: the IP in canonical form, so an IPv4-mapped IPv6 address from a dual-stack socket is the IPv4 one, and the port. IPv6 flow info and scope are left out.

    pub(crate) fn endpoint(addr: SocketAddr) -> (IpAddr, u16);
  • Client: SocksUdpLink in protocols/src/socks/udp_link.rs skips any datagram whose source is not the relay address the server named, tested as endpoint(from) != endpoint(self.relay), so a dual-stack client socket still hears an IPv4 relay. Its own UDP ASSOCIATE names no source (all zeros, as RFC 1928 has a client do when it does not know), so a server that checks sources holds it to the control connection’s address and the port of its first datagram:

    impl<S> DatagramLink for SocksUdpLink<S>
    where
    S: AsyncRead + AsyncWrite + Unpin,
    {
    type Addr = Destination;
    fn poll_recv_from(
    &mut self,
    cx: &mut Context<'_>,
    buf: &mut ReadBuf<'_>,
    ) -> Poll<io::Result<Destination>>;
    }

A change to either side needs a negative test that sends from a third address and asserts nothing is delivered. The existing ones are udp_association_ignores_another_ip, udp_association_ignores_another_port_once_pinned, udp_association_is_not_widened_by_the_request and udp_link_ignores_datagrams_not_from_the_relay in protocols/tests/pipeline/socks.rs (the two that send from 127.0.0.2 run on Linux only). The same file pins the refusals with udp_association_over_a_unix_socket_needs_its_exact_source and udp_association_refuses_a_relay_that_cannot_hear_the_client, and the positive round trip is new_server_vs_new_client_udp. The unit tests in protocols/tests/unit/socks/server.rs pin ExpectedSender directly, the hears rules included through ExpectedSender::new, and endpoint_sees_through_ipv4_mapping_and_ignores_flow_info (protocols/tests/unit/socks/protocol.rs) pins the comparison. SOCKS describes the association in full.

Queues are bounded and backpressure reaches the producer

Section titled “Queues are bounded and backpressure reaches the producer”

Every per-connection buffer or queue has a fixed upper bound. Moving items from a bounded channel into an unbounded collection to keep a producer running defeats the bound: when the consumer stalls, the stall must travel back to whoever is producing.

  • Mechanism: ProxyServerRuntime allocates exactly three buffers of BUF_SIZE bytes per connection, once, and never grows them: the transport read buffer (ReadBuffer<const N: usize>), the transport staging buffer (WriteBuffer<const N: usize>), both in concepts/src/buffer.rs, and an outbound scratch array. The contract in concepts/src/runtime.rs is that a forward that cannot complete holds every effect behind it, and while a forward of transport bytes waits the transport is not read; an outbound is not read while the staging buffer lacks STAGING_RESERVE plus one byte (plus MAX_DATAGRAM for a datagram outbound). A client that stops reading therefore stops the downlink at the outbound socket.

  • WireGuard driver: the driver task in protocols/src/wireguard/device.rs moves every connection’s bytes between bounded channels and smoltcp sockets, so besides the runtime’s buffers it keeps per-connection state of its own. Each connection’s way up is an Uplink: the application’s channel, bounded at CHANNEL_CAP = 256 items, and a slot for at most one item the socket has not accepted yet. The driver reads the channel only while that slot is empty (poll_refill stays Pending while an item is held), so a remote or tunnel that stops draining one connection fills that connection’s socket (a TCP socket buffers TCP_BUFFER = 64 * 1024 bytes each way), then its channel, and then blocks the application’s sender, while the other connections keep moving. It hands one socket at most CHANNEL_CAP items per pass, so no producer keeps the driver to itself. On the way down, a socket is read only while its channel has room. A UDP datagram that finds its association’s send ring full is held in the same slot until the next poll empties the ring; one that can never be sent (larger than the whole ring, or to an unaddressable endpoint) is dropped rather than retried.

    struct Uplink<T> {
    rx: Option<mpsc::Receiver<T>>,
    held: Option<T>,
    }
    impl<T> Uplink<T> {
    fn take(&mut self) -> Option<T>;
    fn hold(&mut self, item: T);
    fn poll_refill(&mut self, cx: &mut Context<'_>) -> Poll<()>;
    }
  • Tests: stalled_outbound_holds_uplink_but_not_other_downlink, frame_larger_than_the_buffer_is_an_error and datagram_outbound_is_never_truncated_by_staging_backpressure (concepts/tests/runtime.rs); backpressure_from_the_wire_reaches_the_writer (concepts/tests/client.rs); a_stalled_tcp_flow_blocks_its_writer (protocols/tests/pipeline/wireguard.rs), which checks that a flow whose remote reads nothing blocks its writer once the buffers above are full and does not hold up another flow on the same tunnel; a_held_item_keeps_the_channel_unread (protocols/tests/unit/wireguard/device.rs).

Session tables that grow with peer input carry an explicit cap too, for example MAX_SESSIONS = 256 per connection for Hysteria 2 UDP sessions (protocols/src/hysteria/server/datagrams.rs) and for mux sub-flows (protocols/src/mux/demux.rs).

A semaphore that claims to bound active connections must hold its permit until the relay has finished. A permit that covers only the handshake, the decode or the transport stage bounds that stage, not active connections. At every spawn, check that the permit moves into the spawned task.

Where Constant Held by
App stream inbounds (app/src/serve.rs) MAX_LIVE_CONNECTIONS_PER_INBOUND = 65_536 An Arc<OwnedSemaphorePermit> taken in run_stream_inbound, cloned into every stream a transport yields, and bound as _session for the whole of serve_connection
katana stream nodes (src/serve.rs, src/manager/proxy.rs) MAX_LIVE_CONNECTIONS_PER_NODE = 65_536 The same pattern: ProxyManager::accept_stream moves the session permit into the spawned task next to serve_stream
Hysteria 2 circuits (protocols/src/hysteria/server/inbound.rs, datagrams.rs) circuit_permits: in the app, the inbound’s max_circuits, default DEFAULT_MAX_CIRCUITS = 65_536; in katana, MAX_LIVE_CONNECTIONS_PER_NODE A permit per proxy stream, bound as _permit inside the task that awaits the runtime, and one per UDP session, held by the session entry

The stage limits next to them (MAX_HANDSHAKES_PER_INBOUND = 2048 in the app, MAX_TRANSPORT_STAGES_PER_NODE = 2048 and MAX_PREAUTH_STREAMS_PER_NODE = 512 in katana) are deliberately separate and cover only the handshake stage; they are not connection limits. the_circuit_limit_refuses_new_sessions (protocols/tests/unit/hysteria/server/datagrams.rs) pins the circuit cap for UDP sessions.

When a name resolves to several addresses, the caller must pick one its local sockets can actually reach, and keep the others as fallbacks, rather than taking the first answer and dropping the packet when that family is unavailable.

  • Mechanism: select_candidate_ips in protocols/src/helpers/address_family.rs filters the resolver’s answer by the outbound’s AddressFamilyStrategy and by FamilySupport (which families the outbound can source traffic from: FamilySupport::both() for a dialer that leaves source selection to the kernel, FamilySupport::from_addrs for a userspace netstack with its own interface addresses), and orders it with a stable sort so the other family stays as a fallback. DualStackUdp in environment/src/dial/udp.rs binds one socket per requested family, leaves out a family whose bind fails, and sends each datagram from the socket of the peer’s family.

    pub fn select_candidate_ips(
    resolved: Vec<IpAddr>,
    strategy: AddressFamilyStrategy,
    support: FamilySupport,
    ) -> Vec<IpAddr>;
    impl DualStackUdp {
    pub fn socket_for(&self, peer: &SocketAddr) -> io::Result<&UdpSocket>;
    }
  • Tests: auto_skips_families_without_a_local_address, prefer_ipv4_keeps_ipv6_as_fallback and ipv4_only_filters_to_ipv4 (protocols/tests/unit/helpers/address_family.rs); v6_socket_is_v6_only_and_dual_stack_picks_by_family (environment/tests/integration/udp.rs).

A parser checks every fixed field (version, reserved bytes, lengths, command and type codes) and refuses a value it does not know. Never assume the peer is a correct implementation, and never let a declared length allocate or read beyond a bound before it is checked. In environment and protocols the clippy denies above make unchecked indexing and unchecked arithmetic a lint error, so the code has to handle an out-of-range index or an overflowing length instead of panicking on it.

Parser Tests that pin the strictness
VLESS rejects_bad_version, rejects_nonempty_addons, response_header_rejects_version (protocols/tests/unit/vless/protocol.rs)
VMess rejects_unknown_version, rejects_unknown_command, rejects_unknown_security, rejects_invalid_option_relationship (protocols/tests/unit/vmess/protocol.rs)
Trojan rejects_missing_crlf, rejects_oversize_udp_packet (protocols/tests/unit/trojan/protocol.rs)
mux.cool rejects_unknown_status_and_network, rejects_truncated_metadata (protocols/tests/unit/mux/frame.rs)
gRPC framing decoder_rejects_unexpected_field, decoder_rejects_an_oversized_length_before_buffering_the_body (protocols/tests/unit/transports/grpc_framing.rs)
DNS truncated_and_malformed_input_never_panics, rejects_an_over_long_name_and_label (protocols/tests/unit/dns/message.rs)
Hysteria 2 varint_above_62_bits_is_refused_not_panicked, tcp_request_rejects_a_foreign_frame_type (protocols/tests/unit/hysteria/protocol.rs)

A typo must stop the program, not fall back to a default that changes what the configuration means. Stable configuration structs reject unknown fields, and a value that is not recognised is an error, not “anything else”.

  • Mechanism: the configuration structs carry #[serde(deny_unknown_fields)] (33 of them in app/src/config.rs, 12 in katana’s src/config.rs). tls_layer in app/src/transport.rs validates security against the network instead of comparing it with "tls", so an unknown value is refused rather than becoming plaintext:

    pub fn tls_layer(network: &str, security: Option<&str>, ctx: &str) -> io::Result<bool>;
  • Tests: a_mistyped_key_is_rejected_rather_than_ignored and an_inbound_defaults_to_loopback (app/tests/unit/config.rs); an_unknown_security_is_rejected_on_every_network and tcp_with_tls_is_rejected_and_names_the_fix (app/tests/unit/transport.rs); in katana, reject_unknown_sni_is_refused_rather_than_ignored, acme_cert_mode_rejected and hysteria_refuses_an_unknown_obfs_rather_than_disabling_it (tests/unit/inbound.rs) and an_invalid_address_family_is_rejected_for_every_protocol (tests/unit/outbound.rs).

The binaries show it directly. A misspelled key under [inbound.stream] (log timestamp removed):

$ etemenanki-app --test -c typo.toml
ERROR etemenanki_app: configuration invalid: TOML parse error at line 7, column 1
|
7 | securty = "tls"
| ^^^^^^^
unknown field `securty`, expected one of `network`, `security`, `tls`, `ws`, `grpc`

The process exits with status 1 and binds nothing.

Pay particular attention to keys that decide exposure or security: listen address, port, protocol, network, security, TLS settings, outbounds and routes.

A reload must not tear down the running configuration and then discover that the new one cannot start. Parse, validate and build first; switch only when that succeeded; keep the old state when it did not. A failed reload must also be retryable once the outside cause (a busy port, a missing file) is gone, not only when the file’s bytes change.

flowchart LR
  read["read file"] --> parse["parse_bytes"]
  parse -->|error| keep["keep old generation"]
  parse --> build["build: router, inbounds, outbounds"]
  build -->|error| keep
  build --> cancel["cancel old generation"]
  cancel --> spawn["spawn_generation(built, false)"]
  • App mechanism: Instance::reload in app/src/instance.rs runs config::parse_bytes and build before touching the running generation, and logs reload: parse failed, keeping current config or reload: build failed, keeping current config on failure. Only then does it cancel the old generation’s token, wait for its accept loops to release their listeners, and call spawn_generation(built, false), where a single inbound’s bind failure is logged and the other inbounds still start. The swap drops every in-flight connection of the old generation.

    pub fn build(cfg: &Config) -> io::Result<Built>;
    async fn spawn_generation(built: Built, strict: bool) -> io::Result<Generation>;
    impl Instance {
    pub async fn reload(&self);
    }
  • Where the app falls short at the verified revision: the new listeners are bound only after the old generation has released its ports, so an inbound whose bind fails on reload stays down; the old generation is already gone by then. Instance::reload also returns early when the file’s bytes equal the last bytes it tried (a parse or build failure records them too), so an unchanged file is not tried again after the outside cause is fixed. Edit the file or restart the process to retry. A change to this path should close these gaps rather than widen them.

  • katana mechanism: a file that does not load is refused first, in the watcher loop, with config reload failed, keeping current. apply_reload in src/runtime.rs then builds everything the reload needs before touching anything: the new outbound pool when [[outbound]] changed, every node the reload adds (its panel client and its router; nothing is bound or spawned yet), and the panel client of every running node whose entry changed. If any of these fails, it applies nothing and logs reload: bad outbounds, keeping current config or reload: node <display_id>: <error>; keeping current config. Only then are node changes applied, node by node: a removed or re-identified node is stopped, a new one is spawned from what was already built, and a changed one is sent the edit. The running node takes that edit whole or not at all: a route that does not compile or a panel client that does not build is refused with node <id>: config edit refused, keeping the running one: <error>. A node the reload adds or respawns that fails to come up, for example on a port that is still taken, keeps retrying as described under Node tasks are supervised, so it recovers once the outside cause is gone; a running node whose listener rebuild fails tries again at every poll.

  • Tests: a_reload_rebinds_the_udp_port (app/tests/integration/e2e_hysteria_inbound.rs) checks that a Hysteria 2 inbound gets its UDP port back across a reload while the old generation holds a live connection; a_reload_with_a_node_that_does_not_build_changes_nothing (katana tests/unit/runtime.rs) checks that an entry with a mistyped node_type leaves the running node, its live connection and the route edit made alongside untouched, and the other tests in that file pin which edits respawn a node and which reach it in place; route_change_drops_connections and repeated_user_refreshes_never_disturb_a_live_connection (katana tests/unit/e2e.rs) pin what a katana reload and a panel refresh keep. Generations and reload and Runtime and reload describe both paths in full.

A transient failure of the first panel request, of DNS, of a bind or of bring-up must not end a node’s task for good with nothing to restart it. A long-lived node task needs retry with backoff under a supervisor, and must still stop promptly on cancellation.

  • Bootstrap: NodeManager::bootstrap in src/manager/node.rs calls try_bootstrap until the node is up or its CancellationToken is cancelled. An attempt fails when node_info fails, returns nothing or reports port 0, when user_list fails or returns nothing (a node up without its users would serve nobody), or when bring_up fails, for example on a port that is still taken. Each attempt first clears the panel client’s ETags, so a panel that answers a repeated request with 304 Not Modified still gives it a full answer. After a failure the node logs node <id>: <reason>; retrying in <n>s and waits BOOTSTRAP_RETRY_MIN (1 s), doubling with each failure up to BOOTSTRAP_RETRY_MAX (60 s), or up to the poll period (controller.update_periodic) when that is shorter. While it waits it applies static updates as they arrive, which keeps the runtime’s bounded channel to it from filling, and a config edit ends the wait at once, since the edit may be the fix. Cancellation is honoured both during an attempt and during the wait.

    const BOOTSTRAP_RETRY_MIN: Duration = Duration::from_secs(1);
    const BOOTSTRAP_RETRY_MAX: Duration = Duration::from_secs(60);
    impl NodeManager {
    pub async fn run(self: Arc<Self>, shutdown: CancellationToken, static_rx: mpsc::Receiver<StaticUpdate>);
    async fn bootstrap(&self, shutdown: &CancellationToken, static_rx: &mut mpsc::Receiver<StaticUpdate>) -> bool;
    async fn try_bootstrap(&self) -> Result<(), String>;
    }
  • Steady state: NodeManager::poll_cycle falls back to the last applied node info when a panel request fails, returns nothing or reports port 0, and to the last applied user list when that request fails or returns nothing; it logs a failed request or a port of 0 and tries again on the next tick. The loop exits only when its CancellationToken is cancelled. Whether the node stopped in the steady state or while still bootstrapping, run then stops the listener with tear_down and flushes the final traffic and audit reports.

  • Tests: a_node_comes_up_once_the_panel_answers, a_node_whose_port_is_taken_comes_up_once_it_is_free (against a panel that answers repeated requests with 304) and a_node_that_never_bootstraps_still_stops (tests/unit/e2e.rs). Node manager describes the whole loop.

Reload rebuilds whatever caches configuration

Section titled “Reload rebuilds whatever caches configuration”

If an object caches configuration fields when it is built, a reload that changes those fields must rebuild or update that object. Updating the manager’s copy of the config while the client keeps the old one is the failure this rule exists for.

  • Identity: a node’s identity in src/runtime.rs is the tuple of fields that decide which panel node it serves. A reload that changes the identity removes the node and spawns a new one, with a new PanelClient and a fresh traffic registry, so one panel node’s counters are never billed to another; a reload that keeps it sends StaticUpdate::Config to the running node.

    type NodeId = (String, String, u32, String, String);
    fn identity(cfg: &NodeConfig) -> NodeId;
    pub fn panel_node_type(cfg: &NodeConfig) -> String;

    The tuple is (panel_type, api.host, api.node_id, api.key, panel node type), with panel_type lowercased. The panel node type comes from panel_node_type in src/api/mod.rs: newV2board finds a node by its id and the node_type it is asked for, so there it is "vless" for a V2ray-family node (node_type V2ray, Vmess or Vless) with enable_vless, else the lowercased api.node_type; SSPanel finds a node by id alone, so there it is empty. display_id shows the node type after the id when it is not empty.

  • Cached fields: Client::new in src/api/newv2board.rs and src/api/sspanel.rs copies every field it needs from panel_type and [node.api] into the client, including the timeout, api.node_type, api.enable_vless, api.speed_limit and api.rule_list_path. NodeManager::apply_static in src/manager/node.rs therefore builds a new PanelClient whenever panel_type or anything in [node.api] changed, before it stores the edit. The new client takes over the routes the old one read from a newV2board node config (PanelClient::inherit_routes), so audit rules derived from them keep applying until it reads the config itself. It holds no ETags, so the node then runs one full panel poll with it and takes the narrowest teardown the answer needs. An edit such as api.speed_limit or api.rule_list_path keeps every connection unless the panel’s new answer changes the protocol or transport; an edit to a local setting the listener is built from (controller.listen_ip, controller.cert, controller.disable_sniffing, api.enable_vless, [node.hysteria] or [node.route]) forces a new listener generation.

  • What to check: a new field a client caches is covered only if it is inside [node.api] or panel_type; otherwise it must join that comparison in apply_static, or be pushed to the client on reload. A field that decides which panel node is served belongs in the identity.

  • Tests: an_sspanel_node_is_its_panel_node_id, a_newv2board_node_is_also_the_type_it_asks_for, a_node_is_logged_by_its_panel_node_not_its_key, a_newv2board_type_edit_respawns_the_node, an_sspanel_api_edit_takes_effect_in_place and a_client_edit_takes_effect_without_dropping_connections (tests/unit/runtime.rs); a_rebuilt_client_keeps_the_routes_it_has_not_read (tests/unit/api/newv2board.rs).

A token bucket charges the actual number of bytes and produces a wait or a debt in proportion. A chunk larger than the bucket’s capacity must not pass after a single refill period, and the achieved rate must not depend on the relay’s buffer or chunk size.

  • Mechanism: TokenBucket in src/traffic.rs refills at rate bytes per second and banks at most one second of it (a rate of 0 is unlimited). charge always takes the full amount, letting the balance go negative, and returns the instant the debt is repaid; Gate in src/meter.rs charges each transfer after it moves and holds both directions of every flow of that user back while the shared bucket is in debt.

    impl TokenBucket {
    pub fn new(rate: u64) -> Self;
    pub fn charge(&self, n: usize) -> Option<Instant>;
    pub fn ready_at(&self) -> Option<Instant>;
    }
    pub fn determine_rate(node_bps: u64, user_bps: u64) -> u64;
  • Tests: token_bucket_rate_limits, a_chunk_larger_than_the_burst_is_still_limited, debt_holds_back_the_next_charge_too and an_idle_bucket_banks_one_second_and_no_more (tests/unit/traffic.rs); the_limit_holds_however_the_writes_are_sized and the_limit_is_shared_by_both_directions (tests/unit/meter.rs). All but token_bucket_rate_limits, which uses the real clock and a tolerance, run on a paused tokio clock, so their timings are exact.

A limit on active authenticated connections, per node or per user, would need admission control that covers the whole relay in the same way as the connection limits above; the pre-auth limit covers only the handshake stage. See Admission.

Never log a credential, UUID, token, key, panel secret, or a URL whose query carries one. When a credential is malformed, log a safe identifier and the kind of error, never the value.

  • Mechanism: error_for_status in src/api/mod.rs replaces reqwest’s own status check and strips the URL, because the panel key travels in the query string. display_id in src/runtime.rs renders a node as panel type, host, node id and, on newV2board, the panel node type, never api.key. NodeManager::try_bootstrap renders a failed panel request by its outermost error message only, because the cause beneath carries the request URL and its query carries the key. Outbound builders refuse a malformed UUID without quoting it.

    pub fn error_for_status(resp: reqwest::Response) -> Result<reqwest::Response>;
    fn display_id(id: &NodeId) -> String;
  • Tests: a_malformed_uuid_is_refused_without_echoing_it (katana tests/unit/outbound.rs); a_node_is_logged_by_its_panel_node_not_its_key (katana tests/unit/runtime.rs); an_unencodable_password_is_refused_without_quoting_it (Etemenanki protocols/tests/unit/hysteria/auth.rs).

A change to HTTP error handling in a panel client must keep URLs and tokens out of every error it can return.

Five properties must survive every change to the traffic code:

Property Mechanism (src/traffic.rs, src/manager/node.rs) Test (tests/unit/traffic.rs)
A failed report loses nothing Live counters are only reduced after success; residual rows go back through restore_residuals on failure restored_residuals_are_retried
A retry does not bill twice commit_reported subtracts exactly what was reported; a residual or fully drained row leaves the table when it is snapshotted rate_change_drains_old_counter_and_reports_once
Residuals are recoverable A departed user’s bytes become residual rows keyed by uid dropped_user_bytes_become_residuals, a_departed_users_late_bytes_are_still_reported
Removal and re-adding are well defined A removed user leaves the table; a credential rebound to another uid gets a fresh counter, and the old uid’s bytes are reported under the old uid set_users_drops_absent, rebound_credential_reports_the_old_uid_separately
Concurrent increments are kept commit_reported uses fetch_sub instead of resetting to zero commit_reported_preserves_concurrent
sequenceDiagram
  participant N as NodeManager
  participant T as NodeTraffic
  participant P as Panel
  N->>T: snapshot()
  T-->>N: live rows, draining rows, residual rows
  N->>P: report_user_traffic(non-zero rows merged by uid)
  alt report succeeded
    N->>T: commit_reported(up, down) per row with a counter
  else report failed
    N->>T: restore_residuals(rows without a counter)
  end
impl UserCounter {
pub fn commit_reported(&self, up: u64, down: u64);
}
impl NodeTraffic {
pub fn snapshot(&self) -> Vec<TrafficSnapshot>;
pub fn restore_residuals(&self, rows: Vec<(i64, u64, u64)>);
}

End to end, two tests in tests/unit/e2e.rs run a node against a fake panel: vmess_traffic_is_metered_and_reported checks the totals the panel receives, and a_retired_user_stops_while_the_rest_keep_their_connections checks that removing one user stops that user’s connections and leaves the others running.

The route model in environment/src/routing.rs keeps first-match semantics: the first rule that matches wins, and a target no rule matches goes to the default outbound. Changes to matchers or geodata loading are reviewed for case normalisation, regex validation, CIDR prefix validation, malformed protobuf input and rule order. A port range whose lower bound exceeds its upper bound is refused by parse_port_match rather than compiled into a rule that never matches:

pub fn parse_port_match(s: &str) -> io::Result<RouteMatch>;

port_parser_rejects_an_inverted_range and port_ranges_are_inclusive (environment/tests/unit/routing.rs) pin it. The same review applies to harranu, which carries the same model.

Every published crate follows SemVer as Cargo interprets it. Cargo reads a bare requirement such as version = "2.0.0" as a caret requirement, and what “compatible” means depends on whether the version is below 1.0:

Version series Requirement ^X.Y.Z accepts Breaking change Compatible addition Compatible fix
0.y.z with y >= 1 (before 1.0) >=0.y.z, <0.(y+1).0 Bump y Bump z Bump z
x.y.z with x >= 1 >=x.y.z, <(x+1).0.0 Bump x Bump y Bump z

etemenanki-concepts, etemenanki-protocols and etemenanki-app each moved through the 0.y series at their own pace until they were released together as 1.0.0. All four kernel crates moved to 2.0.0 together (etemenanki-environment was first published at that number), a major release because the task-and-channel pipeline was removed rather than kept alongside the new one. At the verified revision etemenanki-protocols is at 2.0.2 (after the 2.0.1 patch), and the other three are still at 2.0.0. katana is at 3.0.1: it took a major version when it moved to kernel 1.0.0 (katana 2.0.0) and again for kernel 2.0.0 (katana 3.0.0), and 3.0.1 is a patch release that fixes node bootstrap and reload and moves its lockfile to etemenanki-protocols 2.0.1 while its Cargo.toml keeps the 2.0.0 requirement.

What counts as breaking:

  • Library crates: removing or renaming a public item, changing a public signature, a trait’s required items or a public type’s fields, or changing a feature’s meaning. Adding a public item or an optional feature is a compatible addition.
  • Behaviour: a change in what existing configuration means, what goes on the wire, or what a caller can observe is judged like an API change, even when every signature stays the same. For etemenanki-app and katana, removing a configuration key or changing its meaning is breaking.

etemenanki-protocols 2.0.1 is one recorded exception. It replaced Demux::feed_whole with feed_chunks; etemenanki_protocols::mux and Demux are public, so by the rule above that removal is breaking. The release shipped as a patch because nothing outside the crate called the method, and the release commit’s body says so. That judgement does not change the rule: a patch that removes a public item needs the same check (search katana and every other dependent for uses) and the same written justification, and when any dependent might use the item, the release is breaking.

etemenanki-protocols 2.0.2 also shipped as a patch, and its public API is unchanged: the source a UDP ASSOCIATE names reaches the SOCKS inbound through the crate-private handshake_with_udp_source in protocols/src/socks/handshake.rs, while the public handshake keeps its signature. The release commit’s body records the behaviour that changed with it: a client over a Unix socket whose request names no exact source is now refused, and so is an association whose relay could never hear its client.

If you cannot tell whether a change is compatible, say so in review and settle it before anything is published. Bump only the crates that changed; a crate whose code and dependency requirements did not change keeps its version.

A release is a deliberate act with outside effects (a registry upload, pushed commits, a tag and the binaries built from it), so it happens only when a release has actually been decided, never as a side effect of landing a change.

Crates are published in dependency order, and katana moves only after the kernel crates it needs are visible in the registry:

flowchart LR
  c["etemenanki-concepts"] --> e["etemenanki-environment"]
  e --> p["etemenanki-protocols"]
  p --> a["etemenanki-app"]
  p --> k["katana: update lockfile, bump, tag vX.Y.Z"]
  e --> k
  c --> k

Skip every crate that was not bumped. When a lower crate changes, check each crate above it:

Changed Check the manifests of Also check
etemenanki-concepts environment/Cargo.toml, protocols/Cargo.toml, app/Cargo.toml Source changes needed in each
etemenanki-environment protocols/Cargo.toml, app/Cargo.toml Source changes needed in each
etemenanki-protocols app/Cargo.toml Source changes needed in the app

The inter-crate dependencies are written as path dependencies with a version and a registry, so the published manifest points at the registry version. If a crate needs something only a newer lower crate provides, raise its requirement to that version. Either way, confirm in Cargo.lock that the workspace resolves to the version you are about to publish.

  1. Preflight. Read the live state of Etemenanki (and katana, if it will follow) as in Before you start, and read .cargo/config.toml for the registry and its credential provider. The registry token is never in the repository; it comes from the environment or Cargo’s credential store. If the branch must be synced first, use a fast-forward only pull (see Git rules) and stop if it cannot fast-forward.

  2. Decide the scope. List what actually changed with git status --short, git diff --name-only and git diff --cached --name-only, and map paths to packages: concepts/**, environment/**, protocols/**, app/**. Root manifest, feature and lockfile changes are judged by which crates they affect.

  3. Choose the versions by the rules in Versioning, and edit only the bumped crates’ version, plus the version requirement other workspace crates declare on them where it must rise. The workspace members’ own entries in Cargo.lock follow on the next Cargo command run without --locked (the gate cargo test --workspace is one); cargo check --workspace --locked fails until they have. Check that Cargo.lock changed only in those crates’ entries.

  4. Run every Etemenanki gate, including cargo check --workspace --locked. A failure stops the release.

  5. Commit. The message names what is released: release etemenanki-protocols X.Y.Z for one crate, release etemenanki-protocols X.Y.Z, etemenanki-app X.Y.Z for several, or release etemenanki X.Y.Z when all four move together. The body says why the bump has the size it has. Check git status and git show --stat afterwards.

  6. Publish in order, one bumped crate at a time, and confirm each exact version is visible before the next:

    Terminal window
    cargo publish -p <package> --registry <registry>
    cargo info --registry <registry> <package>@<version>

    Each crate’s manifest has publish = [...] naming the one registry it may go to. The index can lag for a moment, so retrying cargo info is reasonable.

  7. Push the branch with git push origin <branch>. Etemenanki has no release tags; its release commits carry the versions. Do not start a tag scheme.

  1. Update the kernel dependencies that the release needs, once they are visible in the registry, one package at a time:

    Terminal window
    cargo update -p <package> --precise <version>

    Raise the requirement in Cargo.toml too if katana needs the new version’s API or behaviour. Then check Cargo.lock: each updated crate resolves to the version just published, its source line is the private registry’s index and it carries a checksum, and nothing else moved.

  2. Bump katana’s own version in Cargo.toml, read from the file rather than remembered. A compatible fix is a patch; follow the history for anything larger (katana has taken a major version each time it moved to a new kernel major). Update the katana entry in Cargo.lock too, with any Cargo command run without --locked such as cargo check: the --locked gates refuse a lockfile that no longer matches Cargo.toml. Check that nothing else in the lockfile moved.

  3. Run every katana gate, including cargo metadata --locked --no-deps --format-version 1 and cargo check --locked.

  4. Commit and tag. The history uses a commit titled release vX.Y.Z and an annotated tag vX.Y.Z with the same message:

    Terminal window
    git commit -m "release vX.Y.Z"
    git tag -a vX.Y.Z -m "release vX.Y.Z"
    git rev-parse HEAD
    git rev-parse 'vX.Y.Z^{commit}'

    The tags are annotated, so compare HEAD with the tag’s peeled commit vX.Y.Z^{commit}; a plain git rev-parse vX.Y.Z prints the tag object instead.

  5. Push the branch, then the tag:

    Terminal window
    git push origin <branch>
    git push origin vX.Y.Z

    The tag push starts release.yml, which builds x86_64-unknown-linux-gnu and x86_64-unknown-linux-musl binaries with --locked, fails if either links OpenSSL dynamically or if the musl binary is not fully static, and publishes katana-vX.Y.Z-linux-x86_64-gnu, katana-vX.Y.Z-linux-x86_64-musl and a .sha256 for each as the release’s assets.

If the tag already exists and points at another commit, do not move or force-update it.

In every repository you changed, confirm that local and remote agree:

Terminal window
git status --short --branch
git rev-parse HEAD
git ls-remote origin refs/heads/<branch>
git ls-remote origin 'refs/tags/vX.Y.Z*' # katana: the peeled ^{} line is the release commit
cargo info --registry <registry> <package>@<version> # every crate published

Finally, check that katana’s Cargo.lock resolves every etemenanki-* crate to the version this release published, and bring these pages up to the new release as described in Keeping these docs current.

Stop, and report the actual error, when any of these happens. Do not work around it.

Situation Not acceptable as a workaround
The branch cannot fast-forward, or has diverged from the remote Merging or rebasing and carrying on
Registry authentication fails, or the registry configuration is missing Publishing somewhere else
A publish reports the version exists Overwriting it, or bumping until it goes through
A dependency does not resolve, or the lockfile moves far more than expected Deleting Cargo.lock or running a blanket cargo update
Any gate fails Skipping the test or lowering the clippy level
The new version does not appear in the registry Moving katana to it anyway
A tag exists and points at another commit Moving the tag
SemVer compatibility or the tag convention is unclear Guessing

Uncommitted work belongs to whoever made it and outranks automation.

  • Never run git reset --hard, git checkout -- <file> or other destructive restores, and do not run git clean casually.
  • Never delete, overwrite or hide someone else’s uncommitted changes to make a command or a test pass.
  • Never force-push a branch, and never rewrite history that has been pushed.
  • Check git status before syncing with the remote. When local changes exist, sync with git pull --ff-only --autostash. If the pull cannot fast-forward, stop: merging or rebasing is a decision for a person, not a release step.
  • Tags are never moved once pushed.

Every developer page lists, in its frontmatter, the repository files it describes and the revision it was last checked against:

sources:
Etemenanki: [concepts/src/core.rs, concepts/src/runtime.rs]
katana: [src/traffic.rs]
verified:
Etemenanki: 596916d
katana: v3.0.1

In the docs repository, pnpm check:drift (scripts/check-drift.mjs) reads that frontmatter from every page and reports a problem when:

  • a repository listed under sources has no verified entry, or the revision does not exist in the sibling checkout;
  • a listed file does not exist at the verified revision;
  • any listed file has commits between the verified revision and the released revision recorded for that repository in the docs repository’s docs-versions.json (it prints them). The comparison deliberately stops at the release rather than the checkout’s HEAD, which may hold unreleased work; a repository missing from that file is compared up to HEAD;
  • a translated page’s sources or verified differ from its English page.

It ends with check-drift: N tracked page(s), M problem(s) and exits with status 1 when there is any problem. After each release, move that repository’s ref in docs-versions.json to the release commit or tag and run the check. For every page it reports, re-read the code, update the prose, move verified to the new revision and update the Chinese page’s frontmatter in the same commit. A change to configuration keys also means updating the field tables under src/data/reference/ and the examples under examples/, which pnpm check:examples runs through the real binaries with --test; pnpm build then fails on any broken link or anchor.

The site is public and the code repositories are private. Pages name code by repository-relative path and symbol, never by line number or link, and they never contain real secrets, registry names or URLs, or anything about a particular deployment.