חזרה

Senior Site Reliability / DevOps Engineer– AI Products

Navan
שכר לא צויןTel-Aviv, Israel, היברידיסניורמשרה מלאה
משרה חיצונית, ההגשה באתר החברהאושר שהמשרה פתוחה לפני 17 שעות

You will build and operate the infrastructure that supports Navan’s user-facing generative AI features. You will improve performance, reliability, and resource use for AI systems serving millions of users.

מה תעשו

  • Design, build, and scale infrastructure for AI applications and inference engines.
  • Optimize latency, including Time-to-First-Token and request round-trip time.
  • Manage and scale GPU clusters within Kubernetes.
  • Build fallback systems, circuit breakers, and rate limiting for API failures and traffic spikes.
  • Implement observability for AI workloads, including token usage, model drift, and GPU memory saturation.

דרישות

  • 4+ years of experience in SRE, DevOps, or Production Engineering supporting high-traffic, user-facing applications.
  • Experience with AI workloads and AI infrastructure, such as vLLM, TGI, Bedrock, or OpenAI.
  • Strong expertise in Kubernetes and infrastructure-as-code with Terraform.
  • Hands-on experience with inference servers and vector databases.
  • Proficiency in Python and Go.
  • Deep experience managing cloud compute resources, including specialized GPU instances.
  • Experience with observability tools such as OpenTelemetry, Prometheus, Datadog, or Grafana.

יתרון

  • Experience building semantic caching layers to reduce LLM API costs.
  • Active contribution to open-source LLMOps or MLOps projects.

תנאי סף

  • 4+ years of experience in SRE, DevOps, or Production Engineering
  • Strong expertise in Kubernetes
  • Proficiency in Python and Go
Site Reliability EngineeringKubernetesTerraformPythonGovLLMPrometheusGPU infrastructure