Capacity Testing Harness with k6
Command-line load-testing harness built on Grafana k6 that ramps traffic until a target breaches critical thresholds, measures recovery time, and calculates safe concurrent capacity.
A load-testing harness built on Grafana k6 that finds a system's safe limit automatically and reports it in plain language. One command ramps the load, stops when the system starts to break, then checks whether it recovers.
Highlights
- It answers the question people actually ask — not "what is the p95 latency", but "how many concurrent users is my site still safe with, when does it start to slow down, when do errors appear, and does it come back afterwards". The final summary is written in those terms.
- The ramp stops itself — load rises in steps until the run crosses a critical threshold defined in the profile, so the operator does not have to watch a dashboard and decide when to abort. The stopping rule is data, not judgement.
- Recovery is measured, not assumed — a recovery probe runs automatically after the critical point, because a system that degrades and never comes back is a different failure from one that absorbs a spike and settles.
- Preflight before load — a lightweight validation pass confirms the target and the journey definitions are valid before any real traffic is generated, so a broken journey fails in seconds rather than halfway through a ramp.
- Guardrails are part of the tool — it is explicitly built to find limits on environments you are authorised to test, not to knock over public targets, and the profile is where that boundary is enforced.
- Every run is reproducible from its artefacts — each run writes a timestamped folder containing the preflight result, one JSON per load step, the recovery probes, a final summary and per-run logs, with the newest summary also copied to a stable path.
Key Features
- One-command capacity run — a PowerShell wrapper drives preflight, stepped ramp, critical stop and recovery probe end to end.
- Three execution modes in one k6 script —
preflight,capacity_stepandrecovery_probeshare the same journey definitions rather than drifting apart in separate files. - Mixed journeys — the profile describes several user paths at once, so the load resembles real traffic instead of one endpoint hammered in isolation.
- Configurable critical and recovery policy — thresholds for what counts as critical, and what counts as recovered, are declared in JSON rather than hard-coded.
- Readable output — safe concurrent users, the point where it turns critical, recommended safe load, average response time while safe and while critical, the dominant failure at the critical point, whether the system recovered, and how long it took.
- Structured results — per-step JSON plus a machine-readable final summary, suitable for diffing one run against another.
Tech Stack
- Grafana k6 (load generation and thresholds)
- JavaScript (
loadtest.js— the k6 engine script) - PowerShell (
run-capacity.ps1— the orchestration wrapper) - JSON configuration (
capacity-profile.json) +.envfor targets and secrets
Status & Maturity
Complete for its purpose and used to size real deployments. There is no separate unit-test layer, because the k6 scripts are themselves the executable tests — the thing under test is another system, and correctness here means the thresholds and journeys describe reality. The wrapper is PowerShell, so it is Windows-first; the k6 script itself is portable.
Measured Metrics
| Commits | 5 |
| Date Range | 7 Jun 2026 – 12 Jun 2026 |
| Lines of code | 1,420 |
| Automated tests | self-contained k6 test suites |