Back vallusby ExpandIt The build has to work on a machine that cannot download anything
An airman moving palletised cargo on a forklift in a warehouse

Closed environments · build pipeline

The build has to work on a machine that cannot download anything

Vendoring Node.js dependencies for CI inside a closed network

npm · CI · vendoring

110322-F-YC711-028 (5550420623).jpg - US Air Force from USA, public domain (Wikimedia Commons)

Introduction: npm ci is a network call

Every pipeline eventually reduces to one assumption: that a machine can fetch what it needs at the moment it needs it. Inside a closed network that assumption fails, and it fails at the least convenient point, in the middle of a build, on someone else's infrastructure, with a log that says nothing more useful than a timeout against a registry address.

There are three ways out, they suit different situations, and mixing them badly is how teams end up with a pipeline that works only on the machine it was built on.


1Decide where the network boundary is

Before choosing a technique, be explicit about which stage still has a route out. Almost every workable arrangement puts the boundary in the same place:

connected                          | isolated
-----------------------------------|--------------------------------
resolve and fetch dependencies     |
build the artefact                 |
package everything into one thing  |  transfer
                                   |  install the artefact
                                   |  run

Everything that resolves a name or opens a socket to a registry happens on the left. The right-hand side only unpacks and executes. Most offline build failures are simply an activity that ended up on the wrong side of that line.


2The lockfile is not optional

In a connected environment a missing lock is a latent risk. Here it is a defect. Without it, two installs of the same package.json weeks apart produce different trees, and you have no way to see the difference on a machine you cannot inspect.

bash
npm ci          # installs exactly the lock; fails if it disagrees with package.json
npm install     # may change the lock, which is the last thing you want here

Two rules that follow, and both get broken constantly:

Commit the lockfile. Including in repositories where "it is only a test project".

One lockfile per installed tree. If a shared frontend is vendored into two projects, each has its own lock. A dependency bump or a CVE fix has to be applied to both, and the one nobody remembers is where the vulnerable version keeps shipping.


3Three techniques

Vendoring the tarballs

Fetch every package as a tarball on the connected side, carry the directory, install from it. Blunt, completely self-contained, and it needs nothing on the isolated side.

bash
# connected
npm ci
npm pack $(node -p "Object.keys(require('./package-lock.json').packages)
  .filter(Boolean).map(p => p.replace('node_modules/','')).join(' ')") 2>/dev/null

# isolated
npm ci --offline --cache ./npm-cache

Good when the isolated side has no infrastructure at all. Painful when several projects need the same thing, because each carries its own copy.

Packaging node_modules whole

Install on a machine matching the target, archive the resulting tree, unpack on the far side.

Simplest to explain and the most fragile in one specific way: native modules are compiled for a platform and a Node major version. Build the archive on the same operating system, architecture and Node line as the target, or it will fail at run time rather than at install time, which is much harder to diagnose.

An internal registry proxy

Nexus, Artifactory or Verdaccio inside the network, proxying the public registry through a controlled path or fed manually.

bash
export NPM_CONFIG_REGISTRY=https://nexus.internal/repository/npm-proxy/
npm ci

The right answer once more than one team needs this. It is infrastructure with an owner, which is the cost, and it buys you caching, audit and a single place to blocklist a package.


4Where the artefact should be built

The strongest arrangement removes dependency installation from the isolated side entirely: build the image whe
The strongest arrangement removes dependency installation from the isolated side entirely: build the image where there is a network.Aerial view of shipping containers, and big container cranes at Tacoma's container port -a.jpg - Brian Harris, public domain (Wikimedia Commons)

The strongest arrangement removes dependency installation from the isolated side entirely: build a container image on the connected side, with everything already inside, and transfer the image.

dockerfile
# manifests first: this layer changes rarely and rebuilds fast
COPY package.json package-lock.json ./
RUN npm ci --omit=dev

# source afterwards: this layer changes on every commit
COPY . .

Splitting those two steps is not cosmetic. It is the difference between a rebuild that takes tens of seconds and one that takes a few, and offline it also means the dependency layer can be built once and moved unchanged.

One detail that catches teams shipping into closed environments: the image must build where there is no git repository. Code often arrives as an archive rather than a clone, so a build that stamps its version from git rev-parse will simply fail. Provide a fallback:

dockerfile
ARG BUILD_VERSION=unknown
ENV APP_VERSION=${BUILD_VERSION}

5Verifying before you ship

Three checks, all cheap, all skipped more often than they should be.

Install with the network off. Not "with the registry unreachable", genuinely disconnected. A proxy that silently answers from cache will hide the problem until the cache is cold.

Check for run-time installs. Grep entrypoints, scripts and CI definitions for anything that installs during execution rather than during build:

bash
grep -rnE '(npm|npx|pip|apt-get) (install|ci)' --include='*.sh' --include='Dockerfile' \
  --include='*.yml' . | grep -v node_modules

Diff the tree. After an offline install, compare the resulting node_modules against one produced online for the same lock. They should be identical; when they are not, something resolved differently and you want to know now.


6What tends to go wrong

Optional dependencies quietly missing. Some packages skip optional native components when they cannot be fetched, and succeed. The failure appears later, at run time, as a missing binding.

Peer dependency resolution differing. Different npm versions resolve peers differently. Pin the Node and npm versions on both sides, in the image if you have one.

Post-install scripts reaching out. A surprising number of packages download something in postinstall. These fail offline, sometimes non-fatally, leaving a half-configured package.

A transitive dependency that is a git URL. Resolvable only with access to that host. Find these before the transfer:

bash
grep -o '"resolved": "git[^"]*"' package-lock.json | sort -u

7Conclusion

Vendoring is not a technique so much as a discipline about where the network boundary sits. Get that right and the specific tool barely matters; get it wrong and no tool helps.

  1. Resolve and build on the connected side, execute on the isolated one. Every offline build failure is an activity that crossed that line.
  2. The lockfile is a hard requirement, and there is one per installed tree. Shared code vendored twice means two locks to patch.
  3. Verify by genuinely disconnecting. A warm cache will tell you what you want to hear right up until the first cold machine.

Download the offline copyOne HTML file with everything inside - opens with no internet at all.