# Proposal: separating build from deploy
This is a proposal to change how we deploy Docker services, moving away from giving CI direct access to the production Docker socket. It has been validated with a working proof of concept. This document focuses on the decision itself: what changes, why, and what it costs us.
## The problem with today's setup
Right now, a service's CI pipeline builds the image and deploys it in the same step, by mounting the production host's Docker socket straight into the build container. Whoever can edit a repository's pipeline file and has that repo marked "Trusted" can run any Docker command against that host's daemon. There's no registry in between. The image that ends up running is just whatever the build step happened to produce, with no separate, addressable record of what got deployed and when.
This has worked for us and is simple to reason about. The trade-off is that a mistake or a compromised dependency in any one service's build can reach directly into that service's production host.
## What changes
We introduce a registry ([Zot](https://zotregistry.dev/)) between build and deploy, and a second, central pipeline that owns all production credentials. A service's own pipeline only builds an image and publishes it, plus a small encrypted bundle describing how to run it. It then asks the central pipeline, by name, to deploy that specific bundle. The central pipeline is the only thing that ever holds SSH access to production hosts or the key to decrypt secrets.
### Before
```mermaid
flowchart LR
Dev["Developer"] -->|push| Repo["Service repo"]
Repo -->|triggers| CI["CI pipeline"]
CI -->|Docker socket
mounted directly| Host["Production host"]
CI -->|docker compose
up --build -d| Host
```
### After
```mermaid
flowchart LR
Dev["Developer"] -->|push| Repo["Service repo"]
Repo -->|triggers| BuildCI["Service's CI pipeline"]
BuildCI -->|push image| Registry["Zot registry"]
BuildCI -->|push encrypted
deploy bundle| Registry
BuildCI -->|"trigger deploy
(artifact + target)"| DeployCI["Central deploy pipeline"]
Registry -->|pull image + bundle| DeployCI
DeployCI -->|decrypt secrets| DeployCI
DeployCI -->|SSH| Host["Production host"]
DeployCI -->|docker compose
pull, up -d| Host
```
## Benefits
- **No service pipeline ever touches a production Docker socket or SSH key.** The blast radius of a compromised or careless service repository is limited to that service's own build, not to the host it runs on.
- **What's deployed is a specific, addressable thing.** Every deployment pulls an exact image and bundle by digest from the registry, not "whatever the build produced this time." We can see exactly what's running and roll back to a previous digest.
- **Secrets live encrypted in the service's own repository**, not typed into a CI settings page. They're versioned, reviewable, and travel with the code. For this, we use [SOPS](https://getsops.io/) using [age](https://age-encryption.org/) keys.
- **Adding a new service doesn't touch the deployment machinery.** The central pipeline is generic; onboarding a service is just adding its own repo with a build pipeline and a target.
- **The registry gives us retention and a real audit trail** for what was built and deployed, which we don't have today.
## Downsides and added complexity
- **More moving parts.** We now run a registry and a second pipeline that didn't exist before. That's more infrastructure to operate and monitor.
- **A new secrets workflow.** Encrypting secrets with SOPS and managing age keys is an extra step compared to typing a value into a CI settings page. It needs a short onboarding for anyone touching a service's secrets, and a real process for rotating keys when someone leaves.
- **The CI engine itself has real, sharp edges.** Building this took working around several undocumented or buggy behaviors in Woodpecker and its plugins (see `POC-RESULTS.md`). None were blockers, but they cost real debugging time and mean the deployment scripts are less obvious to read than a plain `docker compose up`.
- **A few things still need hardening before this goes further than a PoC:**
- The token that lets a service pipeline trigger the central one is currently a personal, unscoped credential rather than a purpose-built one. This needs a dedicated bot account with the narrowest access possible.
- Any service can currently ask the central pipeline to deploy to any target host, and there's no check that the artifact being deployed actually came from the service that's asking. Both need an allow-list before we onboard more than one service.
- Production SSH access is currently a normal key, not restricted to running only the deploy script. That's a reasonable next step, not done yet.
- **It's a bigger change for people to understand.** Today's flow is one file, one step. The new flow spans two repositories and needs a mental model of "build here, deploy there" that not everyone will have by default.
## Where this leaves us
The core idea - build pipelines never get production access - held up through everything I threw at it during testing, including cases where it would have been easier to just reach for the Docker socket. The remaining gaps are about tightening who can ask the central pipeline to do what, not about the fundamental design. Given the added operational weight, this is worth doing for services we care about protecting, not necessarily as a blanket replacement for every service we run today.