# Test strategy

## Purpose and truthful outcomes

Tests verify the website and its mapping to historical source. They do not provision GCP, rerun the release pipeline, establish remote branch rules, or prove current cluster status. Each check must propagate failure and retain enough detail for repair. Actual results belong in the final report and audit records; this strategy describes expected responsibilities rather than asserting they passed.

## Command layers

| Command | Responsibility |
| --- | --- |
| `npm ci` | Clean, lockfile-controlled dependency installation |
| `npm run lint` | Repository-specific source/content hygiene checks |
| `npm run check` | Astro/TypeScript/content type validation |
| `npm run test:unit` | URL/helper/schema contracts through Node tests |
| `npm run test:content` | Catalogue/page relationships plus negative content fixtures |
| `npm run build` | Prepare approved assets and create complete production output |
| `npm run audit:site` | Inspect generated links, fragments, assets, metadata, base handling, publication boundaries |
| `npm run test:e2e` | Playwright tests of the production site |
| `npm run audit:sources` | Explicit pinned-source hash verification using `SOURCE_REPO_PATH` for the separate implementation checkout |
| `npm run audit:performance` | Lighthouse lab measurements under a recorded setup |
| `npm run test:profiles` | Independently build/audit/E2E all three named deployment profiles, preserving separate output and audit artifacts |
| `npm run validate` | Orchestrate lint, check, unit, content and the three-profile production matrix with failure propagation |

Use Node satisfying the committed package's `engines` contract. Install Playwright Chromium and Firefox with `npx playwright install --with-deps chromium firefox` in an environment where dependency installation is allowed. Performance measurement and source hashing are explicit operations, separate from deterministic schema/helper checks. Source hashing reads the original checkout through `SOURCE_REPO_PATH`; routine installed-dependency builds do not fetch historical Git objects.

## Unit and content seams

Exercise base helpers with a project base, root base, external URLs, protocol-relative URLs, fragments, and already-prefixed paths. Exercise source URLs with a full SHA, unsafe/missing commit, and encoded file names. Schema tests cover IDs, calendar dates, duplicate records, bad source links, and category/status distinctions.

Catalogue tests check control/source/stage/evidence/reading-path references and route existence. Evidence association is bidirectional. Approval gates reject unapproved image display. Negative fixtures introduce a missing evidence ID, bad source reference, invalid metadata, and broken route in isolated data, then assert failure. They must not alter published content to make a validator look effective.

Content tests require substantive articles, complete required metadata, no unresolved template instructions, baseline-pinned source mapping, and active component coverage. Historical/optional material must remain classified. Code/source observations are not relabelled as executed tests or recorded cloud observations.

## Production integration

Build from locked dependencies; confirm routes, Pagefind assets, favicon, robots, sitemap, canonical URL, descriptions/titles, 404, diagrams, evidence, and approved reports. Audit every generated internal link and fragment after rendering. Internal paths must resolve at the configured project base, not merely at `/` in development. Output auditing must reject raw engineering/ledger/state publication.

Source mapping verifies snapshot hashes against the manifest revision in the separately selected implementation checkout. A missing `SOURCE_REPO_PATH` or missing baseline object is an unavailable source-audit prerequisite, not a routine rendering dependency. Explicitly preparing a source checkout may require network access; it must not import application history into the handbook repository. External link checks distinguish network unavailability from a bad local route.

## Browser matrix and interactions

Use the actual production output through preview. Representative coverage includes home architecture, release journey, evidence catalogue, long reference tables, operational command blocks, diagram pages, and engineering report links. Test desktop about 1440px, tablet 768px, mobile 390px, and 320px stress. Test light/dark, refresh/deep route, and 404.

Search queries are digest, keyless, unsigned, initContainer, provenance, runtime; include result navigation, no-result behavior, keyboard open/close, and base-prefixed chunk navigation. Test sidebar/mobile navigation, table-of-contents anchors below sticky headers, diagram scrolling, full-resolution evidence links, evidence combined filters/empty state, journey previous/next bounds, and code copy content without shell prompts/line numbers. Record uncaught errors, failed required requests, and page-wide overflow.

Run axe in distinct layouts and interactive states, inspect keyboard focus manually, and inspect screenshots rather than only producing them. No-JavaScript checks target readable content, links, native architecture details, and full catalogue/journey fallback; search can require JavaScript. Reduced-motion and print checks inspect the actual output. Fix serious/critical in-scope findings and record tested coverage rather than claiming certification.

## Performance and audit records

Measure representative production routes with Lighthouse, recording date, browser/version, viewport, throttling, route, score categories, LCP, and CLS. Targets are roughly 95 scores, LCP below 2.5s, CLS below 0.1 under the chosen lab configuration. They are not field guarantees and do not establish real-user INP. Record variability and unresolved trade-offs honestly.

Every audit finding records severity, route/path, evidence, correction, and retest. Source-project limitations remain separate from website defects. After final report and public design updates, rebuild and rerun relevant output/browser checks so new report links cannot invalidate the final delivered state. Preserve full reports internally and publish only sanitized summaries.

## Standalone repository migration

Initial validation/performance/screenshot artifacts describe the former `/gcp-supply-chain-security/` website base and remain historical records. The new private handbook repository uses `/gcp-security-handbook/`. Rerun the clean validation chain and production routing/search/output checks there; retain a distinct migration log and verification record. Do not transfer earlier measured results to the new base by editing their URLs or timestamps. The standalone repository has no application workflows, so its workflow tests concern only documentation validation and manual Pages publication.


Chromium runs the full 21-case interaction/viewport/accessibility suite. Firefox adds five mobile production-reading routes and one search/keyboard/deep-link case. This is bounded desktop-engine emulation, not testing on physical mobile devices or every browser. Workflow fixtures additionally reject privilege expansion and pull_request_target execution.

## Domain, image and identity addendum

The new profile matrix is separate from both initial and migration validation. `src/lib/deployment.mjs` resolves `DOCS_PROFILE=custom|github|legacy`, with optional validated `DOCS_SITE`/`DOCS_BASE_PATH` inputs. Unit fixtures must reject non-HTTPS/credential-bearing or path-bearing origins and malformed/unsafe base paths. An invalid environment fails early rather than silently producing a different publishing target. The three supported profiles are:

| Profile | Expected canonical origin and base | Required production behavior | Historical pre-gallery checkpoint |
| --- | --- | --- | --- |
| `custom` | `https://security.devsatym.xyz` + `/` | Root routes/assets/search; no stale repository prefix | Pre-gallery build/output/29 browser cases pass |
| `github` | `https://devsatym.github.io` + `/gcp-security-handbook` | Actual standalone repository's subpath routes/assets/search | Pre-gallery build/output/29 browser cases pass |
| `legacy` | `https://devsatym.github.io` + `/gcp-supply-chain-security` | Explicit requested compatibility prefix; never the current publishing repository | Pre-gallery build/output/29 browser cases pass |

`npm run test:profiles` executes `scripts/validate-profiles.mjs`, archives independent output under `.profile-builds/{custom,github,legacy}`, and stores build/output/browser evidence under `engineering/audits/profiles/{custom,github,legacy}`; screenshots belong to each profile's `screenshots/` directory. It removes ambient coordinate overrides from each named-profile subprocess and restores custom-domain `dist` afterward. Preserve the profile, exact resolved site/base, output SHA-256, actual revision/dirty state, command exit statuses and captured times. A previous profile's successful `dist` audit is not evidence for another profile. Rebuild if a different profile is selected for publication because a production output belongs to one configured origin/base.

For each production artifact, verify generated canonical/sitemap/robots/social URLs, favicon, public evidence/full-resolution links, diagram markup, reports, CSS, any emitted fonts, Pagefind chunks and deep search-result navigation. Browser checks require every rendered image to load with nonzero natural dimensions. Fetch the social PNG through production preview, assert its PNG type and 1200×630 dimensions, and keep it classified as editorial design. All fourteen requested originals now have completed inventory/visual privacy review and are copied unchanged locally; original SHA-256 and dimensions match for all fourteen. Optional redaction was not selected, so no derivative is fabricated. Fresh gallery checks must verify each original/full-resolution response, direct/context labels, retained transcripts, privacy disclosure, hash metadata and absence of new execution claims. Reject local filesystem paths, `/public/` browser paths, private/blob-page hotlinks and temporary attachment URLs.

Exercise identity/project components at 320/390/768/1440px in both themes with keyboard focus and mobile navigation. Verify absent avatar/email/resume/certifications produce useful layout without fabricated details; initials require no image request. Planned/unverified portfolio and project destinations must not render as live documentation links. Review supplied public copy and contribution statements against owner inputs and pinned source history, including retained upstream attribution. Re-run relevant axe/no-JavaScript/print checks after new components are added rather than transferring old screenshots or scores.

Review both documentation workflows for selected-profile propagation, read-only build permissions, exact main dispatch guard, artifact-only publication, bounded retention, standard runners and no cloud actions. The private-repository publication guard must fail clearly before privileged deployment; a workflow fixture must not change actual visibility to test it. Profile tests and local preview establish **LOCAL_VERIFIED** only when their fresh results pass. Public repository eligibility, Pages configuration, DNS ownership/TXT verification, TLS issuance, hosted 404 behavior and sharing-crawler delivery remain external checks, with **AUTHORIZATION_REQUIRED**, **DNS_PENDING** or **HTTPS_PENDING** recorded where applicable. Do not dispatch a remote workflow, modify DNS or relabel a local check **PUBLISHED_VERIFIED**.

The pre-gallery addendum execution is preserved in `history/2026-10-05-pre-gallery-checkpoint/audits/addendum-final-validation.log` and its archived profile records: clean installation added 478 packages and audited 479 with zero vulnerabilities; lint/type checks pass with zero Astro errors/warnings/hints; eight unit and sixteen content/workflow cases pass. Each profile has 48 HTML pages, 4,094 audited links/assets, 22 diagram instances, 149 output files, zero output errors and 29 passing browser cases (23 Chromium, six Firefox), with no skipped/flaky/unexpected cases. Explicit source hashing verifies 156/156 pinned records. These historical records bind dirty handbook HEAD `b925c622c21061781b090d3ee0d772454ab43d04` and source fingerprint `9960886fd599793c299614ee23cc727cbfd985d0d6eb9cd0b09c0e4f25f0e2df`.

Fresh pre-gallery custom-profile Lighthouse records preserved under `history/2026-10-05-pre-gallery-checkpoint/audits/profiles/custom/performance.json` give all four categories 100 on home, signatures and catalogue; LCP values are 256.908, 346.1325 and 362.6472ms, with CLS 0, 0 and 0.0106436095. These are the recorded 1440×1000 simulated desktop lab runs, not field/INP or hosted measurements. Initial and migration performance records remain unchanged.

The subsequent fourteen-image gallery is implemented locally with unchanged reviewed originals in the private repository. A read-only comparison against `audits/all-screenshot-inventory.json` verifies all fourteen local public-asset hashes and native dimensions. Fresh `audits/gallery-final-validation.log` records ten unit and eighteen content/workflow cases passing, then thirty browser cases per profile (24 Chromium, six Firefox), with zero skipped/flaky/unexpected results. This covers 48 authored pages and 22 custom components; each profile has 49 rendered HTML files, 4,335 audited link/asset references, 22 diagram instances, 162 output files and zero output errors. The matrix's source fingerprint is `8fafb9a4efc03f5ae409014f818da4a2d7efeda47ad6417026e7e4efc111a181` at the same dirty handbook HEAD. Artifacts and browser/output records live under `audits/profiles/{custom,github,legacy}`. Earlier 29-case/profile and Lighthouse measurements above remain pre-gallery historical results rather than substituted gallery measurements.

The first complete-gallery Lighthouse collector at fingerprint `8fafb9a4efc03f5ae409014f818da4a2d7efeda47ad6417026e7e4efc111a181` observed home/signatures performance 100, LCP 310.924/452.5289ms and CLS 0/0. Catalogue performance was **80**, LCP 382.351ms and **CLS 0.4658291306**, with its other three score categories at 100. Catalogue missed the recorded performance/CLS targets; this failed measurement remains archived rather than replaced by pre-gallery or repaired scores.

The layout repair updates `TwoColumnContent.astro`, `EvidencePanel.astro` and `theme.css`: desktop TOC stays sticky inside a bounded non-shrinking container, main pane can shrink with `min-width: 0`, and block image links/explicit aspect ratios reserve image space before decoding. The new production regression case intentionally holds PNG responses, verifies reserved image height, releases/decodes the image, requires panel height to remain stable within one pixel, and checks TOC viewport bounds and sticky position after scroll. It exercises the loading failure mode instead of merely mirroring CSS declarations.

`audits/gallery-layout-repair/performance.json` records the repaired frozen fingerprint `1ee38746d3c26eee1a3d3fe6f9882bdf38ea16ddf890fcb0e8872c594953993a`: all four categories 100 on home/signatures/catalogue; LCP 257.708/342.8466/362.340625ms and CLS 0/0/0.0080513528. This is fresh simulated desktop lab evidence, not field/INP or hosted performance. The failed raw run remains under `history/2026-10-05-gallery-layout-shift`; the first functional/gallery capture checkpoint is under `history/2026-10-05-gallery-functional-checkpoint`.

The definitive repaired `npm run validate` exits zero: lint/types, ten unit and eighteen content/workflow cases pass, then thirty-one browser cases per profile (25 Chromium, six Firefox), with zero skipped/flaky/unexpected results. Custom/GitHub/legacy durations are 74.994822/74.191673/80.787562 seconds. Each profile has 49 HTML files, 4,335 link references, 22 diagram instances, 162 output files and zero output errors. `audits/gallery-repaired-final-validation.log` and per-profile artifacts/browser records bind repaired fingerprint `1ee38746d3c26eee1a3d3fe6f9882bdf38ea16ddf890fcb0e8872c594953993a`, actual dirty handbook HEAD and selected inputs. Final captures use the frozen latest helper; earlier captures keep their original provenance. These fresh results verify the changed layout code rather than carrying earlier thirty-case results across it. Real hosted delivery remains unverified.

Final copied-report changes require rebuilding/output auditing each profile and proving non-report bytes remain equal to the applicable browser-tested artifact before carrying forward browser results. Gallery/code/asset changes require their fresh full validation and profile fingerprints; report-only comparison cannot substitute for it. Optional privacy masks remain private proposals; no unanswered choice authorizes image edits or remote publication. Only actual completed artifacts close pending outcome rows.
