We are looking for a Cloud & DevOps Automation Engineer to own the platform layer beneath our automation practice - one of the core service categories on the REWORK platform. Client automations, agents, and pipelines have to run somewhere: you will provision that somewhere with infrastructure-as-code, wire the CI/CD that ships changes safely, containerize and deploy workloads across serverless and VPS targets (including self-hosted n8n and agent runtimes), and build the observability that catches failures before clients do. This is an SRE-minded role: least-privilege IAM and secrets management, backups and restore drills, cost budgets with alerting, and runbooks another engineer can execute at 2 a.m.
- Provision client and internal environments with infrastructure-as-code (Terraform) - reproducible, reviewable, destroyable
- Build CI/CD pipelines (GitHub Actions) with environment promotion, rollbacks, and secret hygiene
- Containerize and deploy workloads across serverless (Cloud Run, Lambda) and VPS targets, including self-hosted n8n and agent runtimes
- Instrument logging, metrics, uptime checks, and alerting so failures page before clients notice
- Design backup/restore and disaster-recovery paths - and actually run the restore drills
- Enforce least-privilege IAM, rotate credentials, and manage secrets across environments
- Track and optimize cloud spend with budgets, alerts, and right-sizing - automation margins live and die on infra cost
- Production infrastructure you provisioned, deployed, and kept running for real workloads
- 1+ years hands-on with at least one major cloud (GCP or AWS), Docker, and a CI/CD system, plus solid Linux fundamentals
- Thinks in failure modes: blast radius, rollback paths, idempotent deploys, and graceful degradation
- Comfortable debugging across the stack - DNS and TLS, container networking, IAM denials, and noisy-neighbor performance
- Writes runbooks, architecture diagrams, and incident notes another engineer can execute without you
- Nice to have: Kubernetes, Fly.io / Railway, Firebase / Cloud Functions ecosystems, SOC 2-style hardening experience
Environments provisioned as code and reproducible from scratch
Deploys shipping through CI/CD with rollback paths that have been exercised
Monitoring and alerting that catches failures before clients report them
Documented runbooks, backups verified by restore drills, and infra cost kept in budget
Terraform / IaC
GitHub Actions CI/CD
Docker
GCP (Cloud Run / Functions)
AWS (Lambda / ECS)
Linux & VPS ops
n8n self-hosting
Nginx / Caddy
Grafana / CloudWatch
Secrets management
IAM & least privilege
Bash / Python
- Work asynchronously in a remote-first environment (Slack, email, documented reports)
- Operate highly independently - evaluated on proof-of-work and reliability of deliverables
- Collaborate during core hours, 10:00 AM – 4:00 PM Eastern Time
Keep infrastructure, pipelines, and runbooks clear, well-documented, and easy to navigate