Join an advanced technology unit operating in a complex environment, in a role that combines professional and managerial leadership of the DevOps / Platform domain.
The position includes responsibility for designing, building, automating, and operating modern infrastructure that supports Data and AI pipelines across Cloud and On-Prem environments, while working with distributed systems, high workloads, and demanding performance and availability requirements.
This is a key role connecting Development, AI, Data Engineering, and mission-critical systems, combining technological leadership, hands-on work, architectural decision-making, and team development.
Key Responsibilities
- Manage and professionally lead a DevOps / Platform team.
- Lead infrastructure architecture for AI and Data systems across Cloud and On-Prem environments.
- Design, build, and operate scalable, distributed, highly available infrastructure.
- Lead CI/CD processes for models, Data Pipelines, and application services.
- Manage and scale distributed clusters, including Kubernetes / Ray.
- Lead automation of Deployment, Scaling, Provisioning, and Monitoring.
- Implement MLOps processes and support the full model lifecycle — Training, Deployment, Monitoring, and Optimization.
- Manage containerized environments and Docker, while optimizing CPU / GPU resources.
- Lead the Observability domain, including Logs, Metrics, and Tracing.
- Take responsibility for system stability, availability, performance, and High Availability.
- Work closely with AI, Data Engineering, and Development teams to improve Delivery and deployment processes.
- Lead the resolution of complex production issues, perform Root Cause Analysis, and improve preventive mechanisms.
- Define Best Practices, standards, and working methodologies for the DevOps / Platform domain.
- Mentor and professionally develop team members, assign tasks, and manage priorities.
- Participate in architectural decision-making and in planning the medium- and long-term infrastructure roadmap.
Requirements:
- Significant experience as a DevOps Engineer / Platform Engineer.
- Experience in professional leadership or management of a DevOps / Platform team.
- Hands-on experience with Kubernetes and Docker.
- Significant experience designing and managing CI/CD Pipelines.
- Experience with at least one major cloud environment: AWS / GCP / Azure.
- Experience working with Linux and managing Production environments.
- Experience writing scripts and automation in Python and/or Bash.
- Strong understanding of Distributed Systems, Scalability, and High Availability.
- Experience designing and implementing Infrastructure as Code.
- Strong troubleshooting capabilities and experience resolving complex Production issues.
- Ability to lead multiple interfaces and work effectively in a dynamic, multidisciplinary environment.
- Ability to combine deep hands-on technical expertise with professional and managerial leadership.
Advantages
- Experience with AI / ML systems or MLOps processes.
- Experience with Ray, Airflow, or Prefect.
- Experience managing GPU workloads and high-performance compute infrastructure.
- Experience with Real-Time or Mission-Critical systems.
- Experience in operational or defense-related environments.
- Experience designing an internal Platform or Developer Platform for R&D teams.