Site Reliability Engineer (SRE) On-Prem
Dream Security
שכר לא צויןTLV - ISR, Tel-Aviv, Israel, מהמשרדמידמשרה מלאה
משרה חיצונית, ההגשה באתר החברהאושר שהמשרה פתוחה לפני 13 שעות
Own the reliability, automation, and deployment of DREAM’s platform across customer sites, including on-prem and hybrid environments. Work with Product, R&D, and customer technical teams to operate Kubernetes deployments, automate infrastructure, and resolve complex deployment and networking issues.
מה תעשו
- Deliver and maintain on-prem and hybrid platform deployments at customer sites.
- Deploy, configure, operate, and improve Kubernetes environments and Helm charts.
- Build and maintain deployment automation with Ansible.
- Manage configurations, manifests, and playbooks in Git; use AWS EC2 and S3 for hybrid components.
- Troubleshoot Linux networking, Kubernetes, and deployment issues as a Tier-3 escalation point.
- Improve delivery pipelines, bootstrap processes, and technical documentation.
דרישות
- 3–5 years of hands-on experience in enterprise infrastructure deployment, systems engineering, or on-prem operational reliability.
- Production-grade Kubernetes experience, including architecture, troubleshooting, and CNI networking; experience creating and managing Helm charts.
- Experience writing Ansible roles and playbooks for automation and infrastructure provisioning.
- Experience with Git and AWS resources, specifically EC2 and S3; strong Linux (Ubuntu) and networking knowledge.
- Experience with storage protocols and GPU-enabled Kubernetes nodes; air-gapped deployment experience.
יתרון
- Additional citizenship, such as EU citizenship.
- Experience with MongoDB, PostgreSQL, Neo4j, or RabbitMQ.
- Experience with Jira or Monday.
- Valid Israeli security clearance.
תנאי סף
- 3–5 years of relevant hands-on experience
- High-level English proficiency
- Willingness to travel to customer sites at least 30%
KubernetesHelmAnsibleGitAWS EC2LinuxLinux networkingGPU-enabled Kubernetes