Senior Site Reliability / DevOps Engineer– AI Products
Navan
שכר לא צויןTel-Aviv, Israel, היברידיסניורמשרה מלאה
משרה חיצונית, ההגשה באתר החברהאושר שהמשרה פתוחה לפני 17 שעות
You will build and operate the infrastructure that supports Navan’s user-facing generative AI features. You will improve performance, reliability, and resource use for AI systems serving millions of users.
מה תעשו
- Design, build, and scale infrastructure for AI applications and inference engines.
- Optimize latency, including Time-to-First-Token and request round-trip time.
- Manage and scale GPU clusters within Kubernetes.
- Build fallback systems, circuit breakers, and rate limiting for API failures and traffic spikes.
- Implement observability for AI workloads, including token usage, model drift, and GPU memory saturation.
דרישות
- 4+ years of experience in SRE, DevOps, or Production Engineering supporting high-traffic, user-facing applications.
- Experience with AI workloads and AI infrastructure, such as vLLM, TGI, Bedrock, or OpenAI.
- Strong expertise in Kubernetes and infrastructure-as-code with Terraform.
- Hands-on experience with inference servers and vector databases.
- Proficiency in Python and Go.
- Deep experience managing cloud compute resources, including specialized GPU instances.
- Experience with observability tools such as OpenTelemetry, Prometheus, Datadog, or Grafana.
יתרון
- Experience building semantic caching layers to reduce LLM API costs.
- Active contribution to open-source LLMOps or MLOps projects.
תנאי סף
- 4+ years of experience in SRE, DevOps, or Production Engineering
- Strong expertise in Kubernetes
- Proficiency in Python and Go
Site Reliability EngineeringKubernetesTerraformPythonGovLLMPrometheusGPU infrastructure