Skip to content

Production Promotion and Deployment Notifications

This guide explains how to use the manual production promotion process and deployment success notifications in the CI/CD pipeline.

Overview

The CI/CD pipeline includes two key features for production deployments:

  1. Manual Production Promotion: Keeps one rolling draft PR in cluster-gitops that a human must mark ready, a code owner must approve, and an administrator merges to promote services from staging to production
  2. Deployment Success Notifications: Automatically notifies the source repository when ArgoCD successfully deploys services

Manual Production Promotion

How It Works

After staging deployment succeeds:

  1. Automatic Trigger: The promote-to-production job automatically starts after successful staging promotion
  2. Update the rolling draft PR: Job copies versions only for services with an existing, reviewed production config. If that changes production configuration, it creates or updates the single rolling draft PR in cluster-gitops (fixed branch production-promotion)
  3. Workflow Completes: CI/CD workflow finishes successfully; the PR stays a draft
  4. Ready for review (explicit human step): A person marks the draft PR ready for review
  5. Code-owner approval: A code owner approves (CODEOWNERS on syrf/environments/production/ plus the Main Protection ruleset, cluster-gitops#1555)
  6. Manual Merge: An administrator merges the PR. Merging is the production deployment
  7. ArgoCD Sync: Production cluster automatically syncs with new versions after merge

Key Point: No GitHub Environment configuration required. The explicit approval happens in cluster-gitops: the PR is a draft until a human marks it ready, and it cannot merge without a code-owner approval. CI never marks it ready, approves it or merges it.

Test gate: staging promotion, and therefore the production PR, runs only when the release-gate job succeeded. That job fails unless change detection succeeded and test-dotnet, test-web and validate-docs-indexes each passed or were skipped because the change set did not need them. A failed or cancelled test stops image publication, retags, Lambda packaging, tags and both promotions. See Release Gate.

Production-only integration gate: promote-to-production also requires the production-gate job. It fails unless change detection succeeded and test-dotnet-integration passed or was skipped because no relevant .NET change was detected. A failed, cancelled or timed-out integration run still promotes to staging but holds the production PR. The lane is often cancelled at its 25-minute timeout (28 of the last 60 main pushes on 2026-10-03), so expect held production PRs until that lane is stabilised. Re-run the failed jobs to release a held run. Because the production PR copies all of staging, a later run that passes the gate also carries a held run's versions, so review the production PR's version table before merging. See Production Gate.

Staging-only services are skipped. The promotion workflow never creates a missing syrf/environments/production/{service}/config.yaml; production enablement requires its own reviewed GitOps change before automatic promotion can include that service.

If no effective production config changes remain—because no staging services have a reviewed production config or every opted-in config already matches—the job reports a successful no-op. It does not create or update the rolling PR, and it closes an open rolling PR (keeping its branch), because that PR could only propose stale content.

The Rolling Draft Promotion PR

There is at most one open production promotion PR. Each main push that passes the release and production gates updates it in place instead of opening a new PR. The job uses peter-evans/create-pull-request@v8 with branch: production-promotion, draft: always-true and delete-branch: false, and runs it only when the copied production configuration differs from cluster-gitops main.

Situation when the job runs What happens
Production differs from main; no open rolling PR The branch is pushed (an old branch is reset onto main) and a new draft PR is opened
Production differs; open PR with different content The branch is force-pushed and the title and description are regenerated. A PR someone marked ready is converted back to draft, and the ruleset dismisses earlier approvals
Production differs; open PR already has this content Only the title and description are regenerated. Draft state and approvals are unchanged
Production already matches staging; open rolling PR The action does not run. The job closes the PR with a comment and keeps the branch
Production already matches staging; no open PR Nothing happens
release-gate or production-gate failed The job does not run. The open PR keeps its last gate-passing content
Rolling PR merged or closed by a person The next run that finds a production difference opens a new draft PR on the same branch. Closing the PR does not stop future proposals

The action's own empty-diff behaviour is never used. Its documented behaviour when the diff with main becomes empty is to reset the branch to main so that the open PR closes automatically (and to delete the branch when delete-branch is true). The job checks for a production difference first and handles the no-change case itself, with an explanatory comment.

Because draft: always-true re-drafts the PR whenever its content changes, the Ready for review click always refers to the content currently in the PR.

Workflow Overview

CI/CD Build → Staging Promotion → Rolling Draft PR Created/Updated → Workflow Complete
                                                        ↓
              (Human marks ready → code owner approves in cluster-gitops)
                                                        ↓
                    PR Merged (= deployment) → ArgoCD Syncs Production

Production Promotion Process

Step 1: Automatic Rolling Draft PR

When code is pushed to main and staging deployment succeeds:

  1. CI/CD runs: All services build and deploy to staging
  2. Production job triggers: promote-to-production job starts automatically
  3. Draft PR created or updated: The rolling draft PR on the production-promotion branch in cluster-gitops is created, or updated in place, with the labels production, requires-review and ci-cd
  4. Workflow completes: CI/CD workflow shows success

Step 2: Review the Production PR

  1. Navigate to cluster-gitops:
  2. Go to https://github.com/camaradesuk/cluster-gitops/pulls
  3. Open the draft PR from the production-promotion branch (label requires-review): gh pr list --repo camaradesuk/cluster-gitops --head production-promotion

  4. Review the changes:

  5. Check which services are being updated
  6. Verify versions match what's in staging
  7. Review the checklist in PR description

  8. Checklist to verify (included in PR):

  9. All services have been tested in staging
  10. No critical issues reported in staging
  11. Release notes reviewed (if applicable)
  12. Stakeholders notified of production deployment

Step 3: Mark Ready, Approve and Merge

Merging is the production deployment, so each of these is an explicit human step:

  1. Mark ready for review: CI keeps the PR a draft, and converts it back to draft whenever its content changes
  2. Code-owner approval: a code owner for syrf/environments/production/ approves. A later push to the PR dismisses that approval
  3. Merge: an administrator merges; ArgoCD then syncs production

Option 1: GitHub UI: click Ready for review, approve as a code owner, then Merge pull request.

Option 2: CLI

gh pr list --repo camaradesuk/cluster-gitops --head production-promotion
gh pr view <PR_NUMBER> --repo camaradesuk/cluster-gitops   # Review changes
gh pr ready <PR_NUMBER> --repo camaradesuk/cluster-gitops  # Explicit human step
# A code owner approves in the UI or with: gh pr review <PR_NUMBER> --approve
gh pr merge <PR_NUMBER> --repo camaradesuk/cluster-gitops --squash   # Deploys to production

Keep the production-promotion branch. The next promotion reuses it, and deleting it is harmless but unnecessary.

Step 4: Monitor Deployment

After merging:

  1. Check ArgoCD sync:
kubectl get applications -n argocd | grep production
  1. Watch deployment progress:
# For specific service (e.g., API)
kubectl get pods -n syrf-production -l app.kubernetes.io/name=syrf-api -w
  1. Verify new version:
kubectl get deployment api-production-syrf-api -n syrf-production \
  -o jsonpath='{.spec.template.spec.containers[0].image}'

Example Production PR

The rolling PR's description is regenerated on every update. It opens with a note that it is the single rolling draft PR, that marking it ready is an explicit step, that merging needs a code-owner approval, and that merging deploys. A simplified older example:

## Production Promotion - Manual Review Required

This PR promotes services from staging to production.

**⚠️ MANUAL REVIEW REQUIRED**: This PR requires manual approval and merge by an administrator.

**Source Run**: [42](https://github.com/camaradesuk/syrf-test/actions/runs/19340633801)
**Source Commit**: 6b47f42a
**Created By**: platform-bot

### Services Updated

- API: `api-v8.21.0`
- Web: `web-v5.0.1`
- Docs: `docs-v1.2.3`

### Review Checklist

Before merging, please verify:

- [ ] All services have been tested in staging
- [ ] No critical issues reported in staging
- [ ] Release notes reviewed (if applicable)
- [ ] Stakeholders notified of production deployment

### Deployment

Once merged, ArgoCD will automatically sync the changes to the production cluster.

Deployment Success Notifications

How It Works

After ArgoCD successfully syncs a service to staging or production:

  1. PostSync Hook: ArgoCD triggers a Kubernetes Job after successful sync
  2. GitHub App Authentication: Job uses GitHub App credentials from cluster secrets
  3. API Call: Creates a commit status on the source repository
  4. Status Details:
  5. Context: argocd/deploy-{environment} (e.g., argocd/deploy-staging)
  6. State: success
  7. Description: "Deployed {service}:{version} to {environment}"
  8. Target URL: Link to deployed service

Enabling Deployment Notifications

Prerequisites:

  • GitHub App credentials must be available in the cluster
  • Secret github-app-credentials must exist in the service namespace

Configuration Overview

Deployment notification configuration follows DRY principle with three levels of inheritance:

  1. Environment Shared Values (cluster-gitops/environments/{env}/shared-values.yaml):
  2. Contains common configuration for ALL services in that environment
  3. Already configured with defaults (githubOrg, githubRepo, credentials, etc.)
  4. Services inherit these values automatically

  5. Service-Specific Values (if needed):

  6. Override only what's unique to the service
  7. Typically only need to set enabled: true

Step 1: Enable Notifications for a Service

Edit the service's environment-specific values file in cluster-gitops/environments/{env}/services/{service}.yaml:

# Example: environments/staging/services/api.yaml
service:
  name: api
  chartTag: api-v8.21.0

  # Enable deployment notifications - inherits config from shared-values.yaml
  deploymentNotification:
    enabled: true  # This is all you need!
    # commitSha will be populated by CI/CD during promotion

That's it! The service inherits all other configuration from environments/staging/shared-values.yaml:

  • githubOrg: camaradesuk
  • githubRepo: syrf-test
  • credentialsSecret: github-app-credentials
  • serviceAccount: default
  • createReleaseNote: false (staging) or true (production)

Step 2: Ensure GitHub App Credentials Exist

Verify the secret exists:

kubectl get secret github-app-credentials -n syrf-staging

Expected output:

NAME                      TYPE     DATA   AGE
github-app-credentials    Opaque   2      10d

If missing, the extra-secrets chart should create it automatically from Google Secret Manager.

Step 3: (Optional) Override Configuration

If a service needs custom configuration different from shared values, you can override in the service file:

# Example: environments/staging/services/api.yaml
service:
  name: api
  chartTag: api-v8.21.0

  deploymentNotification:
    enabled: true
    # Override only if needed (inherits from shared-values.yaml by default)
    githubRepo: different-repo  # Only if this service uses a different repo
    createReleaseNote: true     # Only if you want releases in staging (unusual)

Note: In most cases, you won't need any overrides - shared values cover all services.

Viewing Deployment Notifications

Commit Statuses

  1. Go to the commit in GitHub (e.g., https://github.com/camaradesuk/syrf-test/commit/{sha})
  2. Scroll to Checks section
  3. Look for statuses like:
  4. ✅ argocd/deploy-staging - Deployed api:8.21.0 to staging
  5. ✅ argocd/deploy-production - Deployed api:8.21.0 to production

GitHub Releases (Optional)

If createReleaseNote: true:

  1. Go to Releases tab in GitHub
  2. Look for releases named like: api deployed to staging
  3. Each release includes:
  4. Service name and version
  5. Environment (staging/production)
  6. Deployment timestamp
  7. URL to deployed service

Troubleshooting Deployment Notifications

PostSync Job Not Running

Check if the job exists:

kubectl get jobs -n syrf-staging -l argocd.argoproj.io/hook=PostSync

If not running, check ArgoCD Application status:

kubectl get application api-staging -n argocd -o yaml | grep -A10 status

Job Fails to Authenticate

Check job logs:

POD=$(kubectl get pods -n syrf-staging -l job-name -o name | tail -1)
kubectl logs $POD -n syrf-staging

Common issues:

  • Missing github-app-credentials secret
  • Invalid GitHub App private key
  • Insufficient GitHub App permissions

Status Not Appearing on GitHub

Verify the commit SHA is correct:

kubectl get application api-staging -n argocd -o yaml | grep commitSha

Check GitHub App installation:

# Get installation ID
curl -H "Authorization: Bearer {JWT}" \
  -H "Accept: application/vnd.github+json" \
  "https://api.github.com/orgs/camaradesuk/installation"

Best Practices

Production Promotion

  1. Always test in staging first: Never bypass staging for production deployments
  2. Review PR changes: Check the cluster-gitops PR before merging to production
  3. Monitor deployments: Watch ArgoCD sync status after merging
  4. Rollback plan: Know how to revert production deployments quickly
  5. Communication: Notify stakeholders before merging production PRs

Deployment Notifications

  1. Start disabled: Enable notifications after testing in staging
  2. Use commit statuses: Don't create releases for every deployment (noise)
  3. Monitor job failures: Set up alerts for failed PostSync jobs
  4. Clean up old jobs: Jobs auto-delete after 5 minutes (TTL)

Security Considerations

GitHub App Permissions

The GitHub App needs these permissions:

  • Repository permissions:
  • statuses: write - Create commit statuses
  • contents: write - Create releases (if enabled)

  • Organization permissions:

  • members: read - Verify installation

Production PR Protection

Controls on the cluster-gitops repository (applied 2026-10-03):

  • Draft until a human acts: CI creates the rolling PR as a draft and re-drafts it on every content change; it never marks it ready, approves or merges it
  • CODEOWNERS: syrf/environments/production/ (and other production paths) are owned, cluster-gitops#1555
  • Main Protection ruleset: requires a code-owner approval and dismisses stale approvals when new commits are pushed
  • Workflow tokens cannot approve PRs
  • Audit logs: Review who merged production PRs

Troubleshooting

Production Promotion Fails

Symptom: promote-to-production job fails after staging success

Solutions:

  1. Check PR creation errors:
gh run view {run_id} --log-failed
  1. Verify GitHub App token:
  2. App ID and private key are correct
  3. App is installed on cluster-gitops repository
  4. App has contents: write and pull_requests: write permissions

  5. Check YAML validation:

  6. Ensure all service files are valid YAML
  7. Check for syntax errors in updated files

Deployment Notification Job Timeout

Symptom: PostSync job exceeds 5-minute TTL

Solutions:

  1. Increase job timeout (not recommended):
spec:
  activeDeadlineSeconds: 600  # 10 minutes
  1. Simplify notification logic:
  2. Remove release creation
  3. Use commit statuses only

  4. Check GitHub API rate limits:

curl -H "Authorization: Bearer {token}" \
  https://api.github.com/rate_limit

Future Enhancements

Planned improvements:

  1. Slack notifications: Send deployment notifications to Slack channels
  2. Deployment metrics: Track deployment frequency and success rates
  3. Automated rollback: Trigger rollback on failed health checks
  4. Progressive delivery: Canary deployments with automatic promotion