Back vallusby ExpandIt The review that decides whether your tool ever runs
A corridor and a closed door inside a decommissioned facility

Closed environments · security review

The review that decides whether your tool ever runs

Getting test tooling through a customer's security review in a closed environment

Security review · SBOM · supply chain

Dülmen, Kirchspiel, ehem. Sondermunitionslager Visbeck, Gebäude 30, Innenansicht -- 2023 -- 8939.jpg - Dietmar Rabich, CC BY-SA 4.0 (Wikimedia Commons)

Introduction: the stage nobody budgets for

Every closed-environment deployment has a moment where the engineering work stops mattering. The runner is packaged, the images are built, the offline install has been rehearsed. Then the customer's security team asks for a questionnaire to be filled in, and everything waits.

This is where deployments die. Not on architecture, not on price, not on a failed proof of concept. They die because the vendor treated the security review as paperwork to be handled at the end, discovered it takes eight weeks, and by then the budget quarter had closed.

The review is not an obstacle placed in your way by people who do not understand software. It is the customer's own control, run by people who will personally answer for what runs inside their perimeter. Treat it as an engineering deliverable with its own requirements, and it becomes predictable. Treat it as an interruption, and it will behave like one.

What follows is what the process actually looks like from the vendor's side, what gets asked, what the honest answers are, and which mistakes cost the most time. No customer is identifiable here: the classes of finding and the shape of the process are common across the sector.

This piece belongs to a series on test automation in closed environments, and goes deeper into the one stage the overview only mentions in passing.


1Who is on the other side

The other side is not one gatekeeper. It is architecture, application security, infrastructure and compliance,
The other side is not one gatekeeper. It is architecture, application security, infrastructure and compliance, each protecting something different.Combined Space Operations Center (CSpOC) staff monitors workstations (Screen images have been altered for operational security) (9169412).jpg - U.S. Space Force S4S by David Dozoretz, public domain (Wikimedia Commons)

"Security" is rarely one person, and the parts do not want the same things.

RoleWhat they are protectingWhat convinces them
Security architectThe perimeter and the data flow modelA diagram of every connection your tool opens, and evidence it opens no others
Application securityCode quality and the supply chainSBOM, dependency provenance, an answer to each SAST class
Infrastructure / platformThe hosts and the runtimeResource profile, privileges required, how it is patched
Compliance / riskThe audit trail and supplier registerContracts, data-processing terms, retention, subprocessors
The sponsor inside the companyTheir own credibilityAnswers they can forward without translating

That last row is the one vendors underestimate. Your reviewer is usually not the person who wants your tool. Someone else does: a QA lead, an engineering manager, whoever hit the problem you solve. They spend their own political capital pushing your tool through, and every unanswered question spends more of it.

Write for them, not for the reviewer. Every document you produce should be something your sponsor can forward without a covering explanation. If it needs one, you have written it for the wrong reader.


2The timeline, honestly

Ask a vendor how long security review takes and you get "a few weeks". Here is a shape that matches reality more often, for a tool that runs on customer infrastructure and touches test data:

PhaseTypical elapsed timeWhat determines it
Intake and questionnaire1 to 2 weeksHow fast you return it, and how few follow-ups it triggers
Document review1 to 3 weeksWhether the pack is complete on first submission
Technical assessment (SAST/SCA, sometimes pentest)2 to 6 weeksTheir queue, not your code
Findings and responses1 to 4 weeksRound trips; each one costs a week
Architecture and network sign-off1 to 2 weeksWhether you asked for firewall exceptions
Final approval and registration1 to 2 weeksChange-advisory cadence, often fortnightly

Two things drive that range more than anything you can control on the day:

Round trips. Each incomplete answer costs roughly a week, because your reply enters a queue rather than a conversation. Three round trips is a month. This is why the document pack matters far more than its contents suggest: it is not about impressing anyone, it is about not going around the loop again.

Their calendar. Change advisory boards meet on a fixed cadence. Miss one and you wait for the next, no matter how good your answers were. Ask early which day of the month it meets, and work backwards from it.

The practical consequence: start the review at the beginning of the engagement, in parallel with the technical work, not after the proof of concept succeeds. A proof of concept that ends with "great, now we start the security review" has not saved anyone time.


3The questionnaire, and what each question is really asking

Security questionnaires look like bureaucracy because they are written to cover every category of supplier, from cloud SaaS to a desktop utility. Most questions are not aimed at you. The ones that are, are asking something narrower than they appear.

"Where is data stored and processed?" They are not asking about your database schema. They are asking whether anything leaves their perimeter. For a tool that runs on their infrastructure, the honest answer is short and strong: results, screenshots, traces and history are written to their disk, in formats readable without your software, and nothing is transmitted anywhere. Say it in one sentence, then say what would be transmitted if a particular feature were enabled, and how it is disabled.

"Do you have access to customer data?" The expected answer from a SaaS vendor is a long qualification about support access and role separation. If your answer is genuinely "no, we have no access at all, there is no channel by which we could", say exactly that and explain why architecturally. This is the single strongest thing an on-premise tool can say, and vendors routinely bury it in hedging language that makes it sound like a SaaS answer.

"How are updates delivered?" They are asking whether your tool can change itself. Auto-update is a code-execution path from your infrastructure into theirs, and in a closed environment it is disqualifying. The answer they want is that updates are packages the customer transfers and installs deliberately, with checksums and signatures, on their schedule - the procedure for carrying a new version across is a document in its own right.

"What third parties are involved?" Subprocessors, telemetry endpoints, error reporting, analytics, fonts and scripts from a CDN, licence servers. Every one is an entry on their supplier register that someone has to justify at audit. A tool with zero of them removes an entire workstream from their side, which is worth more to them than most features.

"How do you authenticate users?" They are checking whether you bring your own identity silo. Local accounts are tolerated; integration with their directory or OIDC provider is what they want, because deprovisioning a leaver has to work everywhere at once.

"What privileges does it require?" Root, host network, privileged containers, access to the Docker socket. Each of these turns a routine approval into an escalation. Know your real minimum before they ask, and know what breaks without each one.


4The document pack

The document pack is not there to impress anyone. It exists to prevent round trips, and each round trip costs
The document pack is not there to impress anyone. It exists to prevent round trips, and each round trip costs about a week.Legal Contract & Signature - Warm Tones.jpg - Blogtrepreneur, CC BY 2.0 (Wikimedia Commons)

Assemble this once, keep it versioned alongside the product, and regenerate it per release. The goal is that the first submission is complete, because completeness is what prevents round trips.

Architecture and data-flow description. One or two pages, not a sales deck. Every component, every port, every connection, and the direction of each. Mark clearly which connections exist only when an optional feature is enabled. Reviewers read this first and form their model of your tool from it; if it is vague, everything afterwards is treated with suspicion.

Software Bill of Materials. Generated per release, in CycloneDX or SPDX. Which one depends on who is reading: security platforms and DevSecOps tooling generally expect CycloneDX, while procurement, legal and federal-adjacent reviewers usually expect SPDX. Producing both costs one line in your build and prevents a round trip.

bash
# per release, alongside the artefact itself
syft scan dir:. -o cyclonedx-json > sbom.cdx.json
syft scan dir:. -o spdx-json      > sbom.spdx.json
sha256sum sbom.*.json dist/*.tar.gz  > SHA256SUMS
gpg --detach-sign --armor SHA256SUMS

Vulnerability report against that SBOM. Do not wait for them to run the scan and send you the results. Run it yourself, and submit it with the findings already triaged. A report you produced showing twelve findings with reasoned dispositions lands very differently from twelve findings they discovered.

Signatures and checksums, with a verification procedure. Not just the values: the exact commands the customer runs to verify them, tested on a machine that has never seen your build system.

Network behaviour statement. The list of outbound connections, and the method by which the customer can verify it themselves. This one has its own section below, because it is the single most valuable document in the pack.

Licence and dependency notices. Every transitive dependency, its licence, and confirmation that none of them are copyleft in a way that affects the customer. Legal review runs in parallel and is a common silent blocker.

Patch and support policy. How fast you respond to a critical CVE in a dependency, how you notify customers who cannot receive automatic updates, and how long a given release is supported. Closed-environment customers upgrade slowly, so "supported for six months" means something very different to them than to a SaaS customer.


5SAST findings and how to answer them

At some point your code, or your container image, goes through a scanner. Sometimes the customer runs it; sometimes they accept yours. Either way, findings come back, and how you answer them determines whether the review takes one round or three.

The single most important thing to understand: you are not being asked to reach zero findings. You are being asked to demonstrate that you understand your own code. A vendor who disposes of every finding as "false positive" is less credible than one who fixes four, accepts two with reasoning, and admits one is a genuine weakness with a fix scheduled.

Classes that come back most often, and what an answer that ends the conversation looks like:

Dependency CVEs in transitive packages. The bulk of every report. Triage each into one of three states: fixed in the release under review, not reachable because the affected code path is never executed, or accepted with a date. "Not reachable" must be argued, not asserted: name the function, say why it is never called. Reviewers see the unargued version constantly and discount it.

Command execution. A test runner spawns processes by design, so this class is guaranteed. The answer is not "it is intentional", it is the boundary: what can reach those arguments, who can supply them, what is validated, and what privileges the spawned process has. If the honest answer is that a user with permission to define a test run can execute code on the runner, say so plainly and describe the role model that contains it. Reviewers can accept a documented boundary; they cannot accept a surprise.

Path traversal in artefact handling. Anything that unpacks archives, serves reports or accepts uploaded traces will attract this. Show the normalisation, show the containment check, and mention the test that covers it.

Secrets in code. Test fixtures and example configuration files trigger this constantly. Fix it rather than explaining it, even for obvious dummies. It is cheap to fix, and a report where the only remaining findings are ones you argued is much stronger than one where you argued about a fake password.

Missing security headers on the served interface. Genuinely worth fixing before the scan rather than after: content security policy, frame options, transport security. It is an afternoon of work and it removes a whole page from the report.

TLS configuration. In a closed environment the certificate is usually issued by the customer's own CA. Make sure your tool accepts a customer-supplied chain properly, and that whatever "accept self-signed" switch you offer is off by default, scoped, and logged when enabled.

Container image findings. Base image CVEs, running as root, no health check. Use a minimal base, pin it by digest, run as a non-root user, and rebuild for each release so the base is current. A large share of image findings are simply staleness.

The format that works for the response is a table, one row per finding, with four columns: identifier, your disposition, the reasoning in one or two sentences, and the release in which it is addressed. Send it as a document, not as comments inside their tracker, because your sponsor needs something forwardable.


6Proving there is no call-home

This is the question everything else in a closed environment reduces to, and it is the one where vendors write paragraphs instead of giving evidence.

An assertion is worth little. "Our product does not send telemetry" is what every vendor says, including the ones whose crash reporter fires on the first exception. What changes the conversation is handing the customer a procedure to verify it themselves, on their hardware, without trusting you.

bash
# On the host, before the tool is started. Capture new TCP connections (the
# first SYN, not the replies) and all UDP, excluding the internal ranges.
sudo tcpdump -n -i any \
  '((tcp[tcpflags] & (tcp-syn|tcp-ack)) == tcp-syn or udp) and not dst net 10.0.0.0/8 and not dst net 172.16.0.0/12 and not dst net 192.168.0.0/16' \
  -w /tmp/egress.pcap

# Run a full workload: install, licence check, a scheduled run, report browsing.
# Then list every destination that was contacted. The field after ">" is the
# destination whatever else tcpdump prints on the line (newer versions add the
# interface and direction when capturing on "any").
tcpdump -nr /tmp/egress.pcap | awk '{for (i = 1; i < NF; i++) if ($i == ">") print $(i + 1)}' | cut -d. -f1-4 | sort -u

Note the three RFC 1918 ranges rather than just 10.0.0.0/8, and UDP alongside TCP. A filter that excludes only one of the ranges, or watches only TCP while QUIC and NTP go out over UDP, is the kind of detail a reviewer notices, and noticing it undermines everything else you handed them.

The internal ranges hide one thing worth seeing: queries to the customer's own DNS resolver. A call-home attempt in a closed network usually ends there, as a lookup that never resolves, so capture port 53 in a second run and read which names were asked for.

Run this yourself first, on a clean machine, and include the output in the pack. Then invite them to repeat it. Two things follow. The reviewer gets evidence rather than a claim, and you find out what your own software actually does, which is not always what you believe. Update checks, font fetches, crash reporters, licence validation and package managers invoked during first run are all common surprises.

Two adjacent points worth stating explicitly in the same document:

The licence must be a file, not a request. Any licensing scheme that periodically contacts a server is a call-home by another name, and it means the customer's installation stops working at an unpredictable moment with no way to fix it locally. Verification has to be offline and deterministic.

Name what would connect if enabled. If an optional integration reaches an external system, list it, say it is off by default, and say how to confirm it is off. Reviewers trust a vendor who volunteers the exceptions far more than one whose statement is absolute until they find something.


7The network and identity questions

Two specific areas generate more rework than the rest of the review combined.

Inbound ports. Enumerate exactly what listens, on which interface, and what authenticates each one. A tool that opens one port for its interface and one per project should say so, rather than describing a range. If any port serves artefacts without authentication, expect that to become a finding, and have the answer ready before it does.

Identity. Local accounts get a tool through evaluation and block it at production approval. What the customer wants is their directory or OIDC provider, group mapping onto roles, and deprovisioning that takes effect immediately. If you support it, say which flows and which claims. If you do not yet, give a date rather than a hedge, because they are planning around it either way.

Egress to the application under test. Worth stating separately, because it is the one connection your tool legitimately needs. Making the distinction yourself, rather than letting them discover it, is the difference between an approved data flow and a finding.


8What not to do

Five mistakes, in rough order of how much time each costs.

Do not ask for a firewall exception. It is the most expensive sentence in the engagement. An exception needs a request, a justification, an approval and a review cycle, and it turns your tool from something that fits their environment into something that changes it. If your product genuinely requires outbound access to function, that is a product problem to fix, not a policy problem to negotiate.

Do not argue with findings. Disagree with them, in writing, with reasoning, and accept the ones you cannot defend. Arguing signals that the next finding will be argued too, which is exactly what makes a reviewer look harder.

Do not say "it is only a test tool". It runs code on their infrastructure, holds their test data, reaches their applications and often holds credentials for them. The category that sounds harmless to you is the one that historically gets least scrutiny and therefore most attention now.

Do not send a sales deck when asked for architecture. Reviewers read a deck as an attempt to avoid the question, and it costs you credibility for the rest of the process.

Do not promise a feature to unblock approval. It goes into their register as a commitment with a date attached, and it will be checked at the next audit whether or not anyone remembers the conversation.


9A checklist

Assemble this before the questionnaire arrives, not after:

Documents

Evidence

Answers ready in advance

Process


10Conclusion

The security review is not a gate placed in front of the real work. In a closed environment it is the real work, because it is the point at which someone decides whether your software is allowed to exist inside a perimeter they are responsible for.

Three things determine the outcome more than the quality of your code:

  1. Completeness on first submission. Round trips are measured in weeks. The pack exists to prevent them, not to impress anyone.
  2. Evidence instead of assertion. Anyone can claim there is no telemetry. Handing over the procedure to verify it, and the output of having run it, is a different category of answer.
  3. Volunteering the exceptions. The vendor who names the one connection that exists, and the one finding they cannot defend, is trusted on everything else. The vendor whose statement is absolute is trusted until the first contradiction, and then not at all.

A product built for these environments makes the review shorter by having fewer things to explain: no telemetry to justify, no subprocessors to register, no auto-update to disable, no external dependency to whitelist. That is not a marketing property. It is the difference between an approval that takes six weeks and one that takes six months.