A Defense-Tech company developing advanced AI and real-world technology is looking for a
Senior DevOps Engineer (Backend Oriented) to take ownership of its infrastructure and deployment environment. The role combines hands-on DevOps with real backend development, Docker, and production on-premises environments.
What You’ll Do:
- Design, build, and own production infrastructure across on-premises and cloud environments, including customer deployments, upgrades, remote management, and recovery.
- Build and maintain Dockerized services and multi-container deployments using Docker Compose.
- Own CI/CD, Infrastructure-as-Code, release processes, and deployment strategy from development through production.
- Build observability across the stack, including monitoring, logging, alerting, and operational practices.
- Manage GPU-based workloads and ensure AI workloads run reliably on real hardware.
- Contribute to backend services and APIs, improving performance, data processing, database access, and real-time event handling.
- Own networking, secrets management, security hardening, and infrastructure reliability.
- Work closely with backend, ML, and product teams to improve scalability, deployment, and operational excellence.
Requirements:
- 5+ years of experience in DevOps, SRE, Platform, or Infrastructure Engineering.
- Strong hands-on backend development experience, preferably with Python.
- Deep hands-on production experience with Docker and Docker Compose.
- Real-world experience operating production systems on-premises, ideally on physical servers.
- Strong experience with AWS, GCP, or Azure, and with CI/CD and IaC tools such as Terraform or Ansible.
- Strong Linux, networking, and system-level troubleshooting skills.
- Production experience with PostgreSQL and observability tools such as Prometheus, Grafana, or Loki.
- Strong understanding of system design, reliability, scalability, and security.
- B.Sc. in Computer Science or a related degree, or equivalent academic background.
- Proactive, hands-on, and comfortable taking ownership in a fast-moving startup environment.
Bonus Points For:
- On-prem / air-gapped environments, edge devices, GPU/NVIDIA.
- Real-time video, streaming, distributed systems, or strict performance constraints.
- Python/FastAPI, ML infrastructure, LLMs, vector DBs, or AI agents.
- Defense-tech, security, robotics, autonomous systems, or physical-world AI.
- 0→1 / early-stage startup experience.