Key Responsibilities
We're hiring an experienced DevOps Engineer to join our engineering team. This role suits someone who has already spent real time operating cloud-native infrastructure at scale - not someone looking to break into the field. You'll bring a GitOps mindset, hands-on Kubernetes depth, and the independence to research and learn new tools quickly in a fast-moving, large-scale environment.
- Design, operate, and troubleshoot Kubernetes clusters (EKS/AKS) at scale
- Manage application delivery and infrastructure using GitOps tools (ArgoCD/Flux)
- Build and maintain Helm charts and Kustomize overlays for multi-environment deployments
- Provision and manage cloud infrastructure using Crossplane
- Own and optimize CI/CD pipelines (GitHub Actions, GitLab CI) for build, test, and deployment workflows
- Maintain and extend observability stacks, including Prometheus, Grafana, alerting rules, and dashboards
- Write automation scripts and tooling in Python, Go, or Bash to streamline operations
- Develop internal developer portals, AI agents, and other engineering productivity tools
- Support AWS infrastructure across multiple accounts and regions, including networking, IAM, compute, and storage services
- Participate in on-call rotations, troubleshoot production incidents, and drive root-cause analysis
- Collaborate with Platform, Security, and Application Engineering teams on infrastructure design and reliability improvements
Qualifications
- 4+ years of experience in DevOps, Platform Engineering, or a related role
- Experience with Kubernetes and Helm in production environments
- Experience with Crossplane and/or Terraform for Infrastructure as Code (IaC)
- Proficiency in AWS
- Experience with GitOps tools such as ArgoCD or Flux
- Ability to script in Python, Go, or Bash for automation and tooling
- Experience with Prometheus and Grafana for monitoring and alerting
- Experience building and maintaining CI/CD pipelines (GitHub Actions, GitLab CI)
- Strong self-learning ability, including independently researching unfamiliar technologies, reading documentation, and applying findings without hand-holding
- Strong troubleshooting skills across networking, compute, and distributed systems
Nice to Have
- Proficiency in Go, or Python beyond scripting
- Experience with service mesh technologies (Istio or similar)
- Experience working in multi-cloud environments