Introduction: the category everyone assumes away
Test reporting has quietly standardised on a shape: results go to a hosted service, the team looks at trends in a browser, someone gets a Slack message. It is a good shape, and it is unavailable to anyone whose test machines cannot open a connection outwards.
What follows is an honest survey of the options that remain, what each one actually gives you, and where each one stops. It includes what our own product does not do, because a comparison that only flatters the author is worth nothing to someone making a purchasing decision.
1What a report has to replace
In a connected environment the report is a convenience: if it is unclear, you log into the machine and look. In a closed one, that fallback does not exist, so the report has to carry everything the absent engineer would have gone looking for.
That is the criterion to judge every option against, and it is not "does it look nice":
| Requirement | Why it matters |
|---|---|
| Opens with no network | A report that fetches anything renders wrong or not at all |
| Carries execution context | Environment, versions, configuration, resources |
| Carries evidence per failure | Screenshot, video, trace, console and network logs |
| Machine-readable alongside | For gates, trends and anything automated |
| Readable without its producer | An archived report must open in five years |
| History across runs | One result cannot distinguish a regression from a known failure |
2The options, and where each one ends
The framework's own HTML report
Gives you: per-run results, evidence attached to failures, and with Playwright a trace viewer that replays the run step by step. Self-contained enough to hand over as a folder.
Ends at: one run. There is no history, no trend, no comparison with yesterday. Also worth checking that your version genuinely embeds everything; some report templates pull a font or a script from a CDN, which is invisible until it renders unstyled somewhere else.
Fits: small suites, a handful of projects, teams whose main question is "what failed in this run".
Allure
Gives you: the same per-run detail plus history and trends when you keep previous results, categorisation of failures, and a presentation that non-engineers read without help.
Ends at: it needs a Java runtime to generate, which is a package on the approved-software list in some environments and a conversation in others. History depends on you preserving the results directory between runs, which is a small operational discipline that gets forgotten.
Fits: most teams that need trends and can accept a JRE on the machine.
Allure TestOps, ReportPortal, Testkube
Gives you: the full dashboard experience: analytics across runs, integrations, and in Allure TestOps' case test-case management as well. All three can be self-hosted, which is what makes them candidates at all.
Ends at: they are services. Self-hosting one means a database, a deployment, backups, upgrades and an owner, inside an environment where every one of those is a change request. The evaluation question is not whether the software is good; it is whether the customer's platform team will accept another stateful service.
Fits: organisations with a platform team and enough suites to justify the operational weight.
Grafana over your own metrics
Gives you: trends next to the rest of the infrastructure dashboards, in a tool the operations team already runs and already trusts.
Ends at: it shows aggregates, not evidence. Nobody diagnoses a failing test from a Grafana panel. Treat it as a complement to a report, never a replacement.
3The acceptance test for any of them
Whatever you pick, verify the offline claim yourself rather than believing the documentation:
# 1. generate a report as you normally would
# 2. list every external reference it contains
grep -rEo 'https?://[^"'"'"' )]+' reports/html/ | sort -uThe output should contain nothing but the application under test. Fonts, analytics, icon sets and script CDNs all show up here, and every one of them is something that will render differently, or hang, on a machine with no route out.
Then do the stronger version: copy the report to a machine that genuinely has no network and open it. Cached assets on your own laptop will hide the problem otherwise.
4Two formats, two audiences
Whichever human-facing report you choose, produce a machine-readable one alongside it. They answer different questions and neither substitutes for the other.
| Machine format | Human format | |
|---|---|---|
| Example | JUnit XML, JSON | HTML, Allure |
| Consumer | CI gates, trend scripts, other tools | a person diagnosing a failure |
| Role | aggregation and thresholds | investigation |
JUnit XML is ugly and almost universally understood, which makes it the safest choice when you do not know what will read your results on the customer's side in two years.
<testsuite name="checkout" tests="42" failures="1" errors="0" skipped="3" time="118.4">
<testcase classname="checkout.email" name="confirmation email" time="0">
<skipped message="No SMTP in the client-a profile"/>
</testcase>
</testsuite>Note the skip with a reason. In a closed environment some tests genuinely cannot run, and the difference between a documented skip and a silent absence is the difference between a report an auditor accepts and one they do not.
5What the header must carry

Six lines at the top of the report answer most of the questions that would otherwise become an email thread:
Suite: regression-suite @ 8f3a21c
Runner: 1.12.0 Chromium 129.0.6668.29 / Firefox 130.0 / WebKit 18.0
Environment: client-a (https://app.client-a.local) · app 4.2.1 · schema 118
Machine: 8 vCPU / 32 GB / shm 2 GB · 4 workers
Started: 2026-09-18 21:14 CEST Duration: 11 min 42 s
Result: 412 passed · 3 failed · 7 skipped (no SMTP, no SSO)Most reporting tools will not produce this for you. Generating it is a small job and it is the single highest-value addition you can make to a report that leaves your sight.
6What vallus does, and what it does not
Since this is our own area, the honest version.
It does: run suites on the customer's own infrastructure with no telemetry, no licence server and no call-home; serve the Playwright HTML report and the trace viewer per run; generate Allure when a JRE is present; keep run history and cross-run trends locally; hold everything on the customer's disk in formats that open without us.
It does not: replace a test-case management system - it keeps no manual test cases and no test plans. It does not yet link tests to requirements either: traceability by tag is planned, but today, if your auditor wants a traceability matrix, that lives in Xray or similar, and vallus files its runs there as test executions. It does not do cross-project analytics of the kind Allure TestOps offers.
What it costs to find out: there is a free tier - two users, one project, one run at a time, on a six-month key you renew. That is enough to run one team's suite for real rather than in a trial sandbox; what it leaves out are the integrations (Xray, SSO, Grafana, ReportPortal).
If your need is trend analytics across dozens of teams, a self-hosted TestOps-class tool is a better fit than we are. If your need is running suites reliably where nothing may leave the network, and getting reports that stand on their own, that is the case we were built for.
7Conclusion
The cloud dashboard is not the only shape reporting can take; it is the shape that happened to win in environments with a network. Remove the network and the question becomes what the report itself has to carry.
Three criteria worth holding any option to:
- It must open with nothing. Verify by disconnecting, not by reading the marketing.
- It must carry the context you will not be able to go and look up. Environment, versions, resources, evidence, explicit skips.
- It must be readable without the tool that made it. Archived artefacts outlive vendors, versions and sometimes the team.
Download the offline copyOne HTML file with everything inside - opens with no internet at all.
