Skip to content

Failover between upstreams

This recipe builds a client-side etemenanki-app that sends your traffic through one of three upstream proxy servers. It uses the first server while that server accepts connections, moves new connections to the second when the first stops answering, and moves back once the first recovers. A variant at the end spreads connections over all healthy servers instead.

Read it if you run more than one upstream server and want the client to route around an outage without you editing the config. It also explains the health probe in detail, because the probe decides everything a balancer does, and it checks less than you might expect.

flowchart LR
  App["Application"] --> In["socks-in 127.0.0.1:1080"]
  In --> R{"Route"}
  R -- "private addresses" --> Direct["direct"]
  R -- "everything else (default)" --> B["balancer upstream"]
  B -- "1st choice" --> P["primary: Trojan"]
  B -- "2nd choice" --> S["secondary: VLESS"]
  B -- "3rd choice" --> T["tertiary: VMess"]
  Prober["Health prober"] -. "TCP connect" .-> P
  Prober -. "TCP connect" .-> S
  Prober -. "TCP connect" .-> T
  • Three outbounds, primary, secondary and tertiary, each reach a different server. The protocols differ on purpose to show that members do not have to match. Yours can all use the same protocol.
  • A [[balancer]] named upstream groups them. With strategy = "failover" the order of its outbounds list is the priority.
  • [route].default = "upstream" sends every flow that no rule catches to the balancer, so the balancer is the fallback for all traffic.
  • A background prober opens a TCP connection to each member’s server and port on a timer. The balancer only picks members whose last probe succeeded.

The balancer’s probe only checks that a server accepts TCP connections. It does not check your password, UUID, TLS name or WebSocket path (see What the probe does not check). Make sure each upstream works on its own first, so that a healthy probe really means a working path.

  1. Write a config with only the first upstream outbound and [route].default pointing at it. Copy its [[outbound]] block from the recipe for its protocol, such as VLESS over WebSocket and TLS or VMess over gRPC and TLS.

  2. Check the config and start it:

    Terminal window
    etemenanki-app --test -c config.toml
    etemenanki-app -c config.toml
  3. From another terminal, send a request through it:

    Terminal window
    curl --socks5-hostname 127.0.0.1:1080 https://example.com/ -o /dev/null -w '%{http_code}\n'

    A status code such as 200 means the whole path works: TCP, TLS, the proxy protocol and your credentials.

  4. Stop the proxy with Ctrl-C, point [route].default at the next outbound, and repeat until every upstream has passed on its own.

config.toml
# The "Failover between upstreams" recipe: a local SOCKS proxy that sends
# everything through one of three upstream servers. "primary" is used while it
# accepts TCP connections, then "secondary", then "tertiary". Private addresses
# go out directly.
# Replace the server names, the password and the UUIDs with your servers' values.
[log]
level = "info"
[[inbound]]
tag = "socks-in"
protocol = "socks"
listen = "127.0.0.1"
port = 1080
[[outbound]]
tag = "direct"
protocol = "freedom"
# Member 1: Trojan over TLS.
[[outbound]]
tag = "primary"
protocol = "trojan"
server = "proxy1.example.com"
port = 443
[outbound.stream]
network = "tls"
[outbound.stream.tls]
server_name = "proxy1.example.com"
[outbound.settings]
password = "replace-with-a-long-random-password"
# Member 2: VLESS over WebSocket and TLS.
[[outbound]]
tag = "secondary"
protocol = "vless"
server = "proxy2.example.com"
port = 443
[outbound.stream]
network = "ws"
security = "tls"
[outbound.stream.ws]
path = "/vless"
[outbound.stream.tls]
server_name = "proxy2.example.com"
[outbound.settings]
id = "11111111-2222-3333-4444-555555555555"
# Member 3: VMess over gRPC and TLS.
[[outbound]]
tag = "tertiary"
protocol = "vmess"
server = "proxy3.example.com"
port = 443
[outbound.stream]
network = "grpc"
security = "tls"
[outbound.stream.grpc]
service_name = "tunnel"
[outbound.stream.tls]
server_name = "proxy3.example.com"
[outbound.settings]
id = "11111111-2222-3333-4444-666666666666"
# All three members carry UDP, so the balancer can take UDP flows too.
# List order is the priority under "failover".
[[balancer]]
tag = "upstream"
outbounds = ["primary", "secondary", "tertiary"]
strategy = "failover" # the default, written out for clarity
probe_interval = 15 # seconds between probes of each member (default 30)
probe_timeout = 5 # seconds a probe may take (default 5)
[route]
# Set explicitly: without it the default would be "direct", the first outbound.
default = "upstream"
[[route.rule]]
outbound = "direct"
cidr = ["10.0.0.0/8", "172.16.0.0/12", "192.168.0.0/16", "127.0.0.0/8", "fc00::/7"]

What each part does:

  • direct comes first among the outbounds only because it is a common layout. It does not matter for the balancer, but it does matter for the default route: without [route].default, flows that match no rule go to the first [[outbound]], which here is direct. A balancer never becomes the default on its own, so this config sets default = "upstream" explicitly.
  • The three members are ordinary outbounds. Nothing in an [[outbound]] block changes when it becomes a balancer member, and you can still route to a member’s own tag directly in a rule.
  • All three members carry UDP. Trojan, VLESS and VMess outbounds carry UDP, so the balancer can take UDP flows as well as TCP. An http or shadowsocks member would silently drop every UDP flow the balancer hands it; see UDP to an outbound without datagrams if you need one.
  • [[balancer]] lists the members in priority order. probe_interval = 15 halves the default of 30 seconds, so a dead server is noticed sooner. Choosing the probe timings explains the trade-off.
  • The direct rule keeps private and loopback destinations away from the upstreams. Rules are checked in order before the default applies, so this rule wins for those addresses. A cidr rule only sees destinations that arrive as an IP address: a name such as nas.lan that resolves to 192.168.1.10 still goes to the balancer (see Domains are not resolved for routing).

A [[balancer]] accepts exactly these keys. Any other key stops the config from loading, for example unknown field `probe_intervall`, expected one of `tag`, `outbounds`, `strategy`, `probe_interval`, `probe_timeout` .

KeyTypeRequiredDefaultDescription
tagstringyes—Name of the balancer. [route].default and [[route.rule]].outbound refer to it exactly as they refer to an outbound. It must not repeat any [[outbound]] tag or the tag of another balancer; both fail with balancer tag <tag> collides with an outbound tag.
outboundsarray of stringsyes—The member outbounds, by exact tag. At least one (a balancer needs at least one outbound). Only [[outbound]] tags are accepted: a missing tag or the tag of another balancer fails with balancer <tag> references unknown outbound tag: <member>. Every member needs a server and port that a TCP connect can probe. hysteria2 members are always refused, and freedom, blackhole and wireguard members are refused because they normally have no server and port; both fail with balancer <tag>: outbound <member> has no upstream a TCP health probe can reach, so it cannot be balanced. Under failover the list order is the priority. A tag listed twice is accepted and counts twice.
strategystring (enum)no"failover"How a healthy member is chosen for each new flow: failover takes the first healthy member in outbounds order, round_robin takes each healthy member in turn. Matched exactly and case-sensitively; anything else, such as "roundrobin" or "Failover", fails with unknown balancer strategy "<value>" (expected "failover" or "round_robin").
probe_intervalu64no30Seconds to wait after one health probe of a member finishes before the next one starts. Each member is probed on its own schedule, and the first probe runs as soon as the configuration starts. 0 is accepted and makes the probes run back to back with no pause, opening connections to the upstream continuously.
probe_timeoutu64no5Seconds one probe may take, name resolution included, before the member counts as down. 0 is accepted but leaves a probe only until the next timer tick, about a millisecond, so a member is marked down unless its connect completes almost at once. Use at least 1.

probe_interval and probe_timeout are whole seconds written as TOML integers. A string such as "30s" or a negative number is a parse error.

  1. Check the config. --test builds everything, balancer included, but does not start the prober, so it does not tell you whether the servers are reachable:

    Terminal window
    etemenanki-app --test -c config.toml
    Configuration OK.
  2. Start it:

    Terminal window
    etemenanki-app -c config.toml

    The prober starts with the configuration and probes every member at once. A member that answers its first probe logs nothing, because every member already counts as healthy before its first probe. A member that fails it logs a line at info level:

    balancer member tertiary is now down
  3. Make primary unreachable. Stop the proxy service on that server, or block it from the client for a moment. On Linux, for example:

    Terminal window
    sudo iptables -I OUTPUT -p tcp -d proxy1.example.com --dport 443 -j REJECT

    iptables only covers IPv4. If proxy1.example.com also has an IPv6 address, add the same rule with ip6tables, or the probe connects over IPv6 and primary stays healthy.

    Within one probe_interval the log shows:

    balancer member primary is now down

    From then on, new connections go to secondary. Run the curl command from Before you start again to confirm. If your upstreams have different exit addresses, an IP echo service shows which one carried the request. etemenanki-app does not log which member a flow went to.

  4. Undo the block:

    Terminal window
    sudo iptables -D OUTPUT -p tcp -d proxy1.example.com --dport 443 -j REJECT

    Remove the ip6tables rule too if you added one.

    Within one probe_interval the log shows balancer member primary is now up, and new connections go back to primary.

Connections that were already open when a member went down are not moved. A TCP connection stays on the member it started on until it closes, and a UDP association keeps the member its sub-link to the balancer opened with. Only new connections follow the balancer’s current choice.

If your upstreams are equivalent and you would rather use all of them than keep some on standby, change the strategy:

config.toml
[[balancer]]
tag = "upstream"
outbounds = ["primary", "secondary", "tertiary"]
strategy = "round_robin"
probe_interval = 15
probe_timeout = 5

Each new connection now goes to the next healthy member in turn. A member that goes down drops out of the rotation, and it rejoins when it answers again.

failover round_robin
Which member a new connection gets The first healthy member in outbounds order The next healthy member in rotation
What the list order means Priority Only the rotation order
Load on the servers All on one server, the others idle Spread across every healthy server
When a member recovers New connections go back to it at once It rejoins the rotation
Every member down The first member in outbounds The first member in outbounds
Good for A preferred server with standbys Several equivalent servers

A few things to know about round_robin:

  • It rotates per connection, not per destination. Two connections to the same website can leave through different servers. Sites that tie a login session to your IP address may notice, so use failover if your upstreams have different exit addresses and that matters.
  • The rotation skips members that are down. One shared counter walks over the members that are healthy at that moment, so when the healthy set changes the rotation continues over the new set.
  • Weights are possible only by repetition. Listing a tag twice, as in outbounds = ["primary", "primary", "secondary"], gives it two slots in the rotation. Each slot is a separate member with its own probe, so that server is also probed twice as often. There is no weight setting.

You can also keep both: a failover balancer as [route].default and a separate round_robin balancer that a [[route.rule]] points at. Each balancer probes its own members, so an outbound that belongs to both is probed twice. Balancers cannot contain other balancers.

Each balancer runs one background task per member. The task repeats a single cycle for as long as the configuration runs:

  1. Resolve the member’s server with the resolver configured in [dns] (see DNS), the same resolver and cache the outbound itself uses. An IP address needs no lookup.
  2. Try a plain TCP connection to each resolved address in turn, and close it as soon as one connects. Nothing is sent over it.
  3. If a connection succeeded within probe_timeout seconds, the member is healthy. A lookup failure, a refused connection or the timeout makes it down. The timeout covers the lookup and all addresses together.
  4. Log a line if the state changed, then wait probe_interval seconds and start again.
stateDiagram-v2
  [*] --> Healthy: config starts, before the first probe
  Healthy --> Down: one probe fails or times out
  Down --> Healthy: one probe connects
  Healthy --> Healthy: probe connects
  Down --> Down: probe fails

Members start out healthy. If they started out down, the balancer would have nothing to send traffic to until the first round of probes finished. There is no threshold in either direction: one failed probe marks a member down and one successful probe marks it up again.

The prober stops when the configuration stops, and it runs whether or not anything is routed to the balancer.

The probe answers one question: does something accept TCP connections at this member’s server and port? That is the failure a balancer exists to route around, an unreachable or stopped server. It does not use the member’s [outbound.stream] settings or its credentials, because checking those would mean running a second proxy client inside the prober.

Some finer points:

  • The probe ignores address_family. It tries every address the resolver returns, IPv4 and IPv6 alike, in the order the resolver returns them. A member restricted to ipv4_only can probe healthy over IPv6 while its traffic can only use IPv4.
  • One address that swallows packets can take the whole timeout. The addresses are tried one after another within a single probe_timeout. If the first address never answers, the probe times out before it reaches the second, and the member counts as down even though the second address works. Point server at a name or address whose first answer is reachable from the client.
  • Servers see the probes. Each probe is a TCP connection that closes without sending a byte. Proxy servers often log these as failed or empty handshakes. With three members and probe_interval = 15, each server sees about four such connections a minute from each client.

When no member is healthy, the balancer sends new connections to the first member in outbounds, primary in this recipe, instead of dropping them. Probing can be wrong for a minute, for example when the client’s own network blips, and a connection sent to a server that may have recovered has some chance to work, while a dropped one has none.

In practice, when all your upstreams are really down:

  • every new connection that reaches the balancer fails, and client applications see connection errors;
  • the log shows a balancer member … is now down line for each member; the failed connections themselves are logged only at debug level, one … connection from … ended: … line each;
  • traffic that a rule sends elsewhere, such as the private ranges sent to direct here, keeps working;
  • as soon as any member answers a probe, new connections go to it.

The balancer never falls back to direct on its own. If you want traffic to leave directly when every upstream is gone, that is a routing decision you have to write yourself, and it usually defeats the purpose of the proxy.

Two numbers decide how the balancer behaves, and they pull in opposite directions.

How long until a dead server is noticed. A probe runs probe_interval seconds after the previous one finished. If a server dies right after a successful probe, the next probe starts up to probe_interval seconds later. A server that refuses connections is marked down almost at once. A server whose packets vanish, the usual case when a host or its network is down, keeps the probe waiting for the full probe_timeout. So the worst case is about probe_interval + probe_timeout seconds, during which new connections to that server fail. Recovery is noticed by the first probe after the server comes back: within about probe_interval seconds when the failing probes were refused, and within about probe_interval + probe_timeout seconds when they timed out.

How many probes the servers see. Each member gets roughly one TCP connection every probe_interval seconds from every client that runs this config.

Profile probe_interval probe_timeout Worst case to mark a silent server down Probes per member per minute
Default 30 5 about 35 s about 2
This recipe 15 5 about 20 s about 4
Quick 5 3 about 8 s about 12

Some guidance:

  • Keep probe_timeout at 3 seconds or more on real networks. A probe is one TCP handshake, and when the first SYN packet is lost, Linux sends it again after 1 second and again 2 seconds later. With probe_timeout = 1, a single lost packet marks a healthy server down, and because there is no threshold, failover traffic then bounces between members. Every probe of a member named by a domain also includes the lookup, which is slow whenever it misses the resolver’s cache.
  • Do not set probe_timeout higher than it needs to be. It only matters for silent failures, and it adds directly to the time those take to notice.
  • Pick probe_interval by how long you can tolerate failing connections. Below about 5 seconds the gain is small and the probe traffic grows quickly, especially with many clients.
  • Never use 0 for either. Both are accepted. probe_interval = 0 makes each member’s probes run back to back with almost no pause, a constant stream of connections to your servers. probe_timeout = 0 gives a probe only until the next timer tick, about a millisecond, so any member that is not on the local network is marked down on every probe. When that happens to all members, the balancer treats them all as down and traffic always goes to the first member: there is no failover at all.

A member must have a TCP endpoint for the probe to connect to. The config refuses any member that does not, rather than treating it as always healthy, which would hide the outage the balancer exists to catch.

Member protocol Accepted Why
socks, http, trojan, vless, vmess, shadowsocks yes Each dials a TCP server and port, which the probe can connect to
hysteria2 no It has a server and port, but a Hysteria 2 server listens on UDP only. A TCP probe would fail every time, mark the member down on its first probe and never mark it up again
wireguard no It has no server; its peer is settings.endpoint, which is UDP
freedom, blackhole no There is no upstream server to probe

The protocol aliases count too: hysteria and hy2 are refused like hysteria2, direct like freedom and block like blackhole.

A refused member stops the config from loading. --test reports it as:

configuration invalid: balancer upstream: outbound direct has no upstream a TCP health probe can reach, so it cannot be balanced

A normal start exits with the same message after failed to start:, and a reload that introduces it is refused and keeps the running config (see Reloading).

If you want a Hysteria 2 or WireGuard server as a fallback, a balancer cannot do it. Route to that outbound with a rule or make it [route].default on its own; switching to or from it is then a config change. Adding direct as a last-resort member is refused for the same reason, and see When every upstream is down for why you probably do not want it.

etemenanki-app reloads when you save the config file. The balancer is rebuilt with everything else: the old probers stop, every member starts out healthy again, and the new probers probe at once. A reload also closes every open connection. The reload log line lists changes to inbounds, outbounds, the route and the log section only, so a save that changes only a [[balancer]] block logs config reload: no changes although the new settings do take effect. See Hot reload.

Each of these fails etemenanki-app --test, where the message follows configuration invalid:, and a normal start, where it follows failed to start:. On a reload it follows reload: build failed, keeping current config: and the previous config keeps running. A mistyped key or a wrong value type in [[balancer]] is a TOML parse error instead, such as unknown field `probe_intervall` or invalid type: string "30s", expected u64, reported after TOML parse error at line …. On a reload a parse error follows reload: parse failed, keeping current config:, and the previous config keeps running as well.

Message Cause Fix
balancer <tag>: outbound <member> has no upstream a TCP health probe can reach, so it cannot be balanced A hysteria2, wireguard, freedom or blackhole member Remove it from outbounds and route to it directly
balancer <tag> references unknown outbound tag: <member> A member tag is misspelt, not defined, or is another balancer List only [[outbound]] tags
balancer tag <tag> collides with an outbound tag The balancer’s tag repeats an outbound or another balancer Give it a unique name
a balancer needs at least one outbound outbounds = [] List at least one member
unknown balancer strategy "<value>" (expected "failover" or "round_robin") A misspelt or differently cased strategy Write failover or round_robin exactly
route references unknown outbound tag: <tag> [route].default or a rule names a balancer that does not exist Fix the tag or define the balancer

And one mistake that produces no error: leaving out [route].default. The config loads, the balancer probes its members, and all traffic goes to the first [[outbound]] instead, which is direct in this recipe.