Why Onion Latency Swings Wildly: What Weeks of Monitoring Actually Show
Ask anyone who operates an onion status tracker which metric misbehaves most, and uptime rarely tops the list. Latency does. Across weeks of automated probes against the same hidden services, response times swing from crisp sub-second replies to multi-second stalls, sometimes within a few hours of each other.
That variance is not random noise, and it is not usually a sign that a site is failing. Most of it traces back to how Tor works underneath: circuits are built hop by hop on demand, relays queue traffic under load, and clients quietly rotate their entry guards. Understanding those three forces turns confusing charts into readable stories.Every circuit is built from scratch
Unlike the ordinary web, Tor has no persistent fast lane. Each circuit is constructed incrementally - guard, middle, then final hop or rendezvous point - with a cryptographic handshake at every extension. The Tor Project's own OnionPerf measurements show circuit build times forming a heavily long-tailed distribution: medians look healthy, but the upper quartile stretches dramatically (Tor Metrics, circuit build times). Onion services pay more than everyone else. A full client-to-service path can involve up to seven relay hops spliced at a rendezvous point, so setup cost roughly doubles compared to a standard three-hop exit circuit. Researchers studying hidden-service quality of service concluded years ago that connection establishment, not raw bandwidth, dominates what users perceive as slowness.Congestion leaves fingerprints everywhere
When relays receive cells faster than they can drain them, round-trip times inflate across every circuit sharing those nodes. In May 2024 the Tor Project publicly confirmed weeks of unusually high load degrading both onion and general traffic, and asked operators to investigate mitigations (Tor infrastructure status). Probes run during such windows tell an unmistakable story. The problem is structural. As the Tor design team put it in the proposal for RTT-based congestion control:"Lack of congestion control is the reason why Tor has an inherent speed limit ... Because onion services paths are more than twice the length of Exit paths (and thus more than twice the circuit latency), onion service throughput will always have less than half the throughput of Exit throughput." (Proposal 324)One congested guard relay serving a busy onion service can therefore drag dozens of unrelated sites into slow territory simultaneously. During sustained flood attacks, Tor even shipped a proof-of-work defense for onion services to make cheap request storms expensive (Tor Project blog).
Guard changes silently reset your baseline
A monitoring probe, like any well-behaved Tor client, pins itself to a small set of entry guards and sticks with them for weeks. When a guard rotates out - historically on a 30-to-60-day schedule, or sooner if it goes dark - every new circuit starts through unfamiliar territory. The COGS research framework, built on eight months of live network data, documented just how much natural churn plus scheduled rotation reshapes a client's day-to-day routing (Elahi et al., WPES 2012). For trackers, the effect appears as a step change rather than gradual drift. Latency baselines shift almost overnight after a rotation; a service that averaged 800 milliseconds may settle near two seconds with nothing changing on the server itself. That is the probe's new route talking, not the site's decline.What weeks of probes actually see
After enough observation, the same recurring shapes appear again and again:- Diurnal waves, as global relay load rises during peak hours in Europe and North America.
- Step changes in baseline latency following guard rotation or relay failure.
- Brief correlated spikes across many sites, typical of network-wide congestion events.
- Timeout clusters that recover on retry, pointing at circuit setup rather than the destination.