About TeraCyte AnalyticsTeraCyte develops advanced imaging and data-processing systems combining microscopy, large-scale image processing, cloud infrastructure, and AI/ML. We are looking for a Software Infrastructure Engineer to design, build, and maintain the infrastructure and platforms powering our software, data, and ML systems. This is a hands-on engineering role with a strong focus on cloud infrastructure, data platforms, and MLOps.
Role OverviewThis is a hands-on engineering role with a strong focus on cloud infrastructure, data platforms, and MLOps. You will design, build, and maintain the infrastructure and platforms powering our software, data, and ML systems — owning CI/CD, observability, reliability, scalability, and infrastructure automation, while working closely with software, data, and ML engineers to improve development and production workflows.
Key Responsibilities- Build and evolve infrastructure and internal platforms for software, data, and ML workloads.
- Design and operate cloud infrastructure across Microsoft Azure and GCP.
- Build and maintain Kubernetes and container-based environments.
- Develop infrastructure services, automation, tooling, and APIs.
- Support large-scale data and image-processing pipelines.
- Build and improve MLOps infrastructure, including model training, deployment, inference, and lifecycle management.
- Own CI/CD, observability, reliability, scalability, and infrastructure automation.
- Work closely with software, data, and ML engineers to improve development and production workflows.
Requirements- 3+ years of experience in Software Infrastructure, Platform Engineering, DevOps, Backend Infrastructure, or a similar role.
- Strong software engineering skills, preferably Python.
- Hands-on experience with Azure and/or GCP.
- Strong experience with Kubernetes and Docker.
- Experience with CI/CD and Infrastructure as Code.
- Experience working with data-intensive or distributed systems.
- Experience with databases such as PostgreSQL and MongoDB.
- Strong Linux, networking, debugging, and problem-solving skills.
Preferred Experience- Experience with MLOps and ML infrastructure.
- Experience with GPU workloads and Kubernetes-based ML/compute environments.
- Experience with workflow orchestration such as Argo Workflows.
- Experience building large-scale data processing pipelines.
- Experience with Azure ML, Vertex AI, or similar ML platforms.
- Experience with messaging and distributed systems such as RabbitMQ, Azure Service Bus, or Kafka.
- Experience with Helm, Terraform, Grafana, Prometheus, and OpenTelemetry.