Files
2026-09-18 22:37:40 +02:00

84 lines
19 KiB
Markdown

# PoC Results
Status: **success**. The full deployment chain described in `concept.md` was validated end-to-end and is verified working: a push to a dummy service repository builds two container images, publishes a SOPS-encrypted deployment artifact via ORAS, triggers a separate central deployment pipeline, which pulls the artifact, decrypts the secrets, transfers everything over SSH, and brings up the service via `docker compose` on a separate target host — with the decrypted secrets confirmed reaching the correct containers and no plaintext secrets left behind afterward.
## What was built
- Two Hetzner VMs provisioned by Ansible (`repo2cicd2deploy`, this repo): `ci` (Woodpecker server+agent, Zot registry, nginx+Let's Encrypt) and `target` (a plain Docker host acting as "production").
- `dummy-service` (Gitea repo): a two-container dummy service (`db` + `app`) with SOPS-encrypted dummy secrets, built and published by its own Woodpecker pipeline.
- `poc-deployment` (Gitea repo): the central deployment repository, containing only generic deployment logic (`bin/run-deployment`, `bin/deploy-compose`) and host aliases (`hosts/target1`), triggered by service repositories via Woodpecker's `deployment` event.
## Verification evidence
- `docker ps` on the target host shows both `poc-deploy-db-1` and `poc-deploy-app-1` running.
- `docker logs` on both containers confirms the decrypted secrets arrived correctly (`DB_ROOT_PASSWORD: received`, `DB_PASSWORD: received`, `APP_PASSWORD: received`).
- No plaintext `secrets.decrypted.env` remains on the target host after a deployment.
- `docker logout` runs after each deployment; the target host's `~/.docker/config.json` shows no lingering registry credentials between runs.
- Zot correctly serves both container images and arbitrary ORAS artifacts, retrievable by tag and by digest.
- The full trigger chain (`dummy-service` push → build → publish → trigger → `poc-deployment` → decrypt → deploy) was exercised repeatedly via real pipeline runs, not just isolated component tests.
## Deviations from `concept.md` / `plan.md`
- **Registry choice was settled before testing Gitea's own OCI registry.** `concept.md` originally planned to test Gitea's built-in registry first and fall back to Zot only if needed. Zot was chosen upfront instead, mainly for its retention-policy support, which Gitea's registry lacks. `concept.md` was updated during the scoping discussion to reflect this.
- **The central deployment repository is named `poc-deployment`, not `deployments`.** This was a naming choice made when the user created the Gitea repository; all references (in `dummy-service`'s pipeline and in the repository's own README) were updated to match.
- **Deployment logic lives in `bin/run-deployment` and `bin/deploy-compose`, not inlined in `.woodpecker/deploy.yaml`.** `concept.md`'s original repository layout already specified a `bin/deploy-compose` script, which the first implementation draft skipped in favor of inlining everything into the pipeline YAML. This was corrected twice during the PoC (first extracting `deploy-compose`, then extracting the remaining orchestration logic into `run-deployment` too) after repeated YAML-quoting issues made the inline approach unworkable. The final `deploy.yaml` step is under 10 command lines; all real logic is in the two scripts, which are independently testable via SSH without needing a pipeline run.
- **Restricted SSH, rootless Docker, per-repo host allow-listing and age-key rotation remain deferred**, exactly as scoped in `plan.md`. Nothing in the PoC changed the assessment that these are follow-up hardening items rather than PoC blockers.
## Bugs and quirks found (with fixes)
These were discovered by running real pipelines against real infrastructure, not by reading documentation alone. Several are undocumented or under-documented behaviors of Woodpecker and its plugins.
1. **`woodpeckerci/plugin-docker-buildx` is no longer privileged by default.** Needs `WOODPECKER_PLUGINS_PRIVILEGED=woodpeckerci/plugin-docker-buildx` set on the Woodpecker server, or image builds fail immediately with a linter error.
2. **`oras` has no Alpine `apk` package.** Any step needing it must download the binary release tarball directly (`curl` + `tar` + `install`).
3. **Woodpecker's `environment:` block does not support `${VAR}` string-substitution.** Only `settings:` blocks and `commands:` do. Variables meant to be computed from `${CI_COMMIT_SHA:0:8}`-style expressions must be `export`ed as the first commands in a step, not set via `environment:`.
4. **Shell-runtime variables referenced in `commands:` must be escaped with `$$` (e.g. `$${REGISTRY}`)** to prevent Woodpecker's own `${...}` pre-processor from evaluating them (and silently substituting empty strings) before the shell ever sees them. Genuine Woodpecker config-scope variables like `${CI_COMMIT_SHA:0:8}` are the only ones that should stay single-`$`.
5. **`woodpeckerci/plugin-trigger`'s `repositories:` entries require an explicit `@branch` suffix**, even when only using `deploy:` mode. Without it: `build no or branch must be mentioned for deploy`.
6. **`plugin-trigger`'s `deploy:` mode with a non-numeric `@branch` requires `last-successful: true`.** Without it: `for deploy build no must be numeric only or for branch deploy last_successful should be true`.
7. **Undocumented: `last-successful: true`'s lookup is hardcoded to match only `push`-event builds.** Confirmed by reading the plugin's Go source directly (`impl.go`'s `findFirstBuild`, filtering on `b.Event == woodpecker.EventPush`). It will never match a `deployment` or `manual` event build, regardless of the target pipeline's `when` filters. This is not mentioned anywhere in the plugin's documentation. The fix was to add a dedicated no-op step gated to `when: event: push` purely so a genuine push-event successful build exists for the lookup to find, alongside the real logic gated to `when: event: [deployment, manual]`.
8. **Woodpecker's "Allow deployments" project setting is off by default** and must be manually enabled on the repository being triggered *into*, or the deploy-trigger API call returns 403. The setting carries an explicit security warning: enabling it lets anyone with push access to that repository use `deploy` events to reach deploy-scoped secrets.
9. **`plugin-trigger`'s `params:` setting silently breaks with more than one list item.** Multiple `KEY=VALUE` entries in the YAML list get comma-joined into a single string by Woodpecker's settings-to-environment serialization, but the plugin's CLI argument parser then receives that as one list item rather than splitting it back apart — so the first parameter's value silently absorbs all subsequent parameters, and every parameter after the first is lost. The confirmed-working fix is to write parameters to a small `.env`-style file first (in its own step) and reference only that single file path in `params:`, which the plugin reads via `godotenv.Read()`.
10. **A specific Woodpecker command-parsing bug/quirk causes "unterminated quoted string" errors** on certain `commands:` list items containing embedded double-quoted arguments. This is a confirmed real bug in Woodpecker's own command handling (not a real shell syntax error — every affected command was independently verified as valid POSIX shell via `sh -n` and in a real Alpine container), acknowledged by a Woodpecker maintainer in a public GitHub discussion. Partial fixes (YAML literal block scalars, removing unnecessary quoting) reduced but did not fully eliminate the issue; it also reproduced intermittently on syntactically identical lines across different runs, suggesting a log-streaming or command-dispatch race condition rather than a deterministic parsing bug tied to specific syntax. The most effective fix was architectural: consolidating what had been many separate `ssh`/`scp` `commands:` list items into a single external shell script (`bin/run-deployment`), invoked from the pipeline with one `commands:` line. This reduces the number of Woodpecker command-dispatch boundaries to a minimum and has been reliable since.
11. **`docker compose pull` fails with "no basic auth credentials"** if the target host's Docker daemon has never authenticated against the private registry. `docker login` must run on the target host as part of every deployment (or the credential would otherwise need to persist between runs, which was avoided by adding `docker logout` after each deployment instead).
12. **`nginx`'s reload signal (SIGHUP) did not reliably pick up newly-added virtual host server blocks** in the same Ansible play run that installed nginx fresh. Not fully root-caused; switching the relevant Ansible handler from `reload` to a full `restart` resolved it reliably.
13. **The Woodpecker server and agent container images run as a non-root user (UID/GID 1000).** Bind-mounted host directories for persistent data must be `chown`ed to match, or the server fails to start with a misleading "no such file or directory" error from its SQLite driver.
14. **`oras` 1.3.4 fails with 400 Bad Request on monolithic blob uploads over plain HTTP from a workstation over the public internet**, because Go's HTTP client sends `Transfer-Encoding: chunked` instead of `Content-Length`, which Zot's blob-upload endpoint rejects (confirmed by comparison against a raw `curl` upload, which succeeds). This did not reproduce when `oras` was run from the CI host itself over `localhost` — the actual path real pipelines use — and was also confirmed fixed once TLS termination (nginx + Let's Encrypt) was added in front of Zot. Documented here as a known friction point for local/manual testing against a plain-HTTP registry, not a blocker for the pipeline itself.
## Open items deferred beyond this PoC
Unchanged from `plan.md`'s original list, still not needed to validate the core architecture:
- Restricted SSH (forced-command / dedicated non-shell user) for the deployment account on production hosts.
- Rootless Docker on production hosts.
- Per-repository allow-list for which host aliases a service repository may deploy to.
- Age key rotation process and procedure for onboarding/offboarding DevOps team members.
- A dedicated bot Gitea/Woodpecker account for the cross-repo trigger token (a personal token was used for the PoC).
## Security analysis
This section evaluates the trust model actually implemented in the PoC, not just the one described in `concept.md`. Several gaps below are PoC-specific shortcuts; others are structural properties of the design that would need attention before onboarding real services. Severity is relative to a self-hosted homelab/small-team context, not an enterprise threat model.
### Secrets exposure
- **`ZOT_USERNAME`/`ZOT_PASSWORD` are available to every step of `dummy-service`'s pipeline, including `build-db`/`build-app`.** These steps run `woodpeckerci/plugin-docker-buildx`, a third-party (if official) plugin, with full access to the registry-push credential. Any change to the Dockerfile or build context that exfiltrates environment variables (e.g. a build stage that `curl`s them to an external host) would leak the shared org-level registry credential. Since this credential is an **org-level** Woodpecker secret, compromising it from *any* one service repo's build step exposes push access to *every* repo under that org's Zot namespace, not just the compromised one.
- **The SOPS age private key and the production SSH private key are both Woodpecker repo-level secrets on `poc-deployment`, gated to `deployment`/`manual` events only** (not `push`), which is the correct application of concept.md's trust boundary — a service repo's build pipeline never sees these two credentials directly. This isolation held up correctly throughout the PoC.
- **The Zot registry credential is also needed inside `poc-deployment`'s deploy step**, for both `oras login` and `docker login` on the target host. It is currently the *same* org-level secret used by every service's build pipeline, meaning the deploy pipeline and every service's build pipeline share one registry credential. A compromise of any single service repo's pipeline configuration is therefore sufficient to obtain a credential that is also trusted by the central deployment pipeline's registry access, even though it does not grant SSH or SOPS access directly.
- **The production SSH private key is transiently written to a plaintext file (`/tmp/ssh/id_deploy`) inside the `show-deployment-request` step's container.** It is never written to the target host's disk — only the corresponding *public* key lives there. The private key file inside the ephemeral build container is deleted along with the container when the step finishes; it does not persist on the Woodpecker agent host beyond the step's lifetime, but it is written to disk (not held only in memory) during that window. This is a reasonable PoC-level trade-off, not a severe issue, but is worth naming explicitly since "plaintext should exist only transiently" is exactly the principle applied to `secrets.decrypted.env`, and the same standard now also applies to the SSH key.
- **The registry credential now correctly does not persist on the target host** between deployments (`docker logout` added after `docker compose up`), and the decrypted application secrets file is deleted from the target host after each deployment. Both were gaps found and fixed during the PoC (see bug list above) rather than being correct from the start.
### Malicious pull requests and untrusted contributions
- **Neither pipeline restricts secrets from pull-request events**, and Woodpecker's project-level "Require approval" setting was left at its default for both repos rather than being explicitly reviewed or configured. Woodpecker's own documentation states secrets are *not* exposed to `pull_request` events by default unless explicitly enabled per secret — none of this PoC's secrets have that opt-in set, so this specific PoC is not currently exposed to the classic "malicious PR reads secrets" attack as implemented. However, this protection was never deliberately verified or exercised during the PoC (no pull request was opened against either repo), so it should be treated as an assumption inherited from Woodpecker's defaults, not as a tested control.
- **This PoC has a single contributor (the repository owner) with push access to both repos.** The realistic multi-contributor threat model — an external or lower-trust contributor opening a pull request against `dummy-service` — was not exercised at all. Before onboarding a real service repository with more than one contributor, the "Require approval for" and "Allow pull requests" settings on that repository should be deliberately reviewed, and any secret that must be available to a `pull_request` build (there are none in this PoC) should be treated as a specific, individually justified exception.
### Cross-repository attack surface
- **The Woodpecker personal access token used for `woodpecker_trigger_token` is unscoped and grants exactly the same permissions as the user account it belongs to** (Woodpecker tokens have no scope/permission model at all — confirmed directly against the API specification during this PoC). Since a personal token was used rather than the dedicated bot account recommended in `concept.md`, any pipeline that can read this secret can trigger a pipeline run on *any* repository the token's owner (here, the instance's admin account) can access — not just `poc-deployment`. This is a materially larger blast radius than intended and is the single highest-priority item to fix before this design is used for anything beyond a PoC.
- **Enabling "Allow deployments" on `poc-deployment` was required to make the trigger work, and this setting carries Woodpecker's own explicit warning**: any user with push access to that repository can use `deploy` events to reach that repository's deploy-scoped secrets (the SOPS key and the production SSH key). In this PoC that risk is theoretical, since only the repository owner has push access — but it means the security boundary protecting the two most sensitive credentials in the whole system ultimately rests on Gitea's push-access control for one repository, not on anything Woodpecker-specific.
- **`poc-deployment`'s `deploy.yaml` accepts `TARGET` from the triggering pipeline and resolves it only against a fixed set of files under `hosts/`,** correctly preventing an arbitrary or attacker-controlled SSH destination from being reached even if `ARTIFACT`/`TARGET` were manipulated — this part of concept.md's design (fail closed on an unknown alias) worked exactly as intended and was verified with the deliberate `hosts/$TARGET` existence check.
- **There is currently no allow-list restricting which service repository may request which `TARGET` alias.** Per `plan.md`'s explicit scope decision, this was deferred deliberately; any repository that can trigger `poc-deployment` can currently request deployment to `target1`. With only one service repo and one target host in the PoC this has no practical effect, but it is a real gap the moment a second, less-trusted service repository or a second production host is introduced.
- **`ARTIFACT` is passed by the triggering pipeline and used directly in an `oras pull` command on the deployment pipeline's side, with no verification that the referenced artifact actually originated from the repository that triggered the deployment.** Any pipeline capable of triggering `poc-deployment` at all (which, given the token scoping issue above, is currently broader than just `dummy-service`) could in principle request the pull and deployment of an artifact pushed by a different, possibly less-trusted repository, as long as that artifact resolves on the shared Zot registry. This was not exploitable within the PoC's single-tenant setup, but is a structural gap worth closing (e.g. binding `ARTIFACT` to the expected repository/namespace) before this design serves multiple independently-trusted services.
### Summary assessment
The core trust boundary from `concept.md` — production SSH credentials and the SOPS decryption key never reaching a service repository's build pipeline — held up correctly throughout the PoC and was the one design property never compromised or worked around, even under significant debugging pressure. The gaps found are concentrated in three places: the breadth of the org-level registry credential (shared across all service builds and the deploy pipeline), the unscoped personal trigger token (broader blast radius than the bot-account design called for), and the absence of any binding between a triggering repository and the artifact/target it is allowed to request. None of these blocked the PoC's goal of proving the mechanism works, but all three should be addressed — the token scoping first, since it is both the cheapest fix and the largest blast-radius reduction — before any real service is onboarded to this pipeline.