Introduction: the failures point at the wrong things
The first run of an existing suite inside a closed network produces a list of failures that look unrelated to each other and unrelated to the network. Tests time out. A browser crashes at random. A report renders as unstyled text. Certificate errors appear on hosts that have valid certificates.
Almost none of it announces itself as a connectivity problem, because software written for a connected world does not fail loudly when the world disappears. It waits, retries, falls back, and eventually times out somewhere far from the cause.
This is the catalogue. Sorted roughly by how long each one takes to diagnose the first time.
1Shared memory, disguised as flakiness
The most expensive entry on the list, because it looks exactly like a race condition in your own code.
Chromium uses /dev/shm for rendering. In a container the default is 64 MB, which is enough right up until a page with several heavy tabs, at which point the renderer dies. The test does not report "out of shared memory"; it reports a closed target, a timeout, or nothing at all. Then it passes on retry, because the next run happened to be lighter.
services:
tests:
shm_size: "2gb" # or: ipc: hostTeams lose weeks to this one. If a suite is unstable in containers and stable outside them, check this before reading a single line of test code.
2Fonts from a CDN
A report or a page that pulls a font from a public CDN behaves differently offline depending on the browser: some wait for the timeout, some fall back immediately. Either way, text renders in a fallback face at different metrics.
That matters in two places. Screenshot comparisons fail on pixel differences that have nothing to do with the change under test. And reports handed to a customer render in Times New Roman, which reads as broken software regardless of the content.
The check takes a second:
grep -rEo 'https?://[^"'"'"' )]+' reports/html/ | sort -uAnything that is not the application under test is a dependency you did not know you had.
3Certificate revocation checks
A browser asked to validate a certificate may try to check whether it has been revoked, over OCSP, against a responder on the public internet. With no route out, that request does not fail fast. It waits.
The symptom is a page that takes 10 to 30 seconds to start loading and then works fine, in a suite that is otherwise fast. It usually shows up on the first HTTPS request after the browser starts, which makes it look like a slow login.
This is worth understanding before reaching for flags: the fix belongs with the customer's PKI configuration, and disabling revocation checking wholesale is a conversation with their security team, not a switch you flip in a config file.
4Telemetry, from things you did not choose
Frameworks, test runners, package managers and editors all phone home by default. Some respect a single environment variable, some need a flag per invocation, some only stop when the network refuses the connection.
Offline, each one costs a timeout somewhere. The worst behave differently on a network that drops packets silently versus one that rejects the connection, which is why a suite can behave differently in two closed environments that both "have no internet".
Audit it once, on a clean machine, with the same capture used for any other outbound question:
sudo tcpdump -n -i any \
'(tcp[tcpflags] & tcp-syn != 0) and not dst net 10.0.0.0/8 and not dst net 172.16.0.0/12 and not dst net 192.168.0.0/16' \
-w /tmp/egress.pcapThen run install, a full suite and a report browse, and look at what shows up.
5Time
An isolated network that has no NTP source, or has one nobody validated, drifts. Two consequences arrive quietly.
Tests that assert on dates start failing in a band around midnight, in a way that depends on which machine ran them. And tokens with short lifetimes get rejected when the issuing service and the test machine disagree about the current time by more than the allowed skew, which surfaces as an authentication failure with a perfectly correct password.
Check the skew between the test machine and the application host before debugging any authentication problem in a closed environment. It takes ten seconds and rules out a whole category.
6External APIs in tests
Payment gateways, address lookup, SMS, mapping, an external identity provider. In the connected world these were a mild dependency; here they are a wall.
What makes this harder than it looks is that the failure is often partial. A test that calls a mocked payment gateway but loads a real map tile from the internet fails in a place that has nothing to do with payments, and the developer who wrote it never saw the tile because their machine had a route out.
7Package installs at the wrong moment
A pipeline that installs dependencies during deployment rather than during build works fine until the machine doing the deploying has no registry. Same for a container that runs apt-get in its entrypoint, or a test that installs a fixture library on first use.
The rule that prevents all of these: installation happens where there is a network, execution happens where there is not. Anything that installs at run time is a connectivity dependency wearing a different hat.
8Timeouts that look like flakiness
Every item above expresses itself, eventually, as a timeout. This is why the first week in a closed environment feels like the suite became unstable rather than like something specific broke.
A practical way to separate the categories: run the suite twice, once with the network namespace fully isolated and once with it merely unable to route outside. If the failure counts differ, you are looking at something waiting for a network rather than something using one, and the fix is a configuration change rather than a code change.
9The order to work through them
- Shared memory, if anything runs in containers. Cheapest fix, largest effect.
- Capture outbound traffic once and look at the list. It answers items 2, 4 and 6 at the same time.
- Check the clock against the application host.
- Grep the reports for external references.
- Move installs to build time, if anything installs at run time.
- Only then start reading test code.
Most teams do this list in reverse, which is why the first month is expensive.
10Conclusion
None of these are exotic. Each one is well understood in isolation, and each one is invisible in a connected environment, which is exactly why the first offline run produces a list of failures that seem to have nothing in common.
They do have something in common: software written for a connected world fails slowly rather than loudly when the world is gone. Once you know that, the diagnosis order changes, and a month of confusion turns into an afternoon of configuration.
The single most useful habit, and the one that answers several items at once: capture what your suite actually tries to reach, on a clean machine, before you assume you know.