Build and own the infrastructure and pipelines used to train, evaluate, package, deploy, and operate machine learning models in production.
Develop reliable MLOps capabilities across experiment tracking, model and data versioning, reproducibility, orchestration, automated testing, monitoring, and controlled model rollouts.
Partner with Data Science, Engineering, and Infrastructure teams to productionize models and continuously improve the scalability, reliability, observability, and cost efficiency of our ML platform.
What We’re Looking For
5–7+ years of experience in Machine Learning Engineering, MLOps, ML Infrastructure, Platform Engineering, or a related production engineering role.
Strong hands-on experience with Python, cloud infrastructure, Docker, Kubernetes, CI/CD, infrastructure-as-code, workflow orchestration, and production observability.
Proven experience building and operating production ML systems, including training pipelines, experiment tracking, model registries, versioning, monitoring, data-quality checks, staged deployments, and rollback mechanisms.
Experience managing GPU-based training workloads and/or distributed training infrastructure, and cloud cost optimization.
Experience working in a fast-growing startup, with the ability to operate in a dynamic, fast-paced environment.