חזרה

Team8- Fintech Stealth Startup- Senior AI Researcher, Evaluation & Agent Performance

Team8
שכר לא צויןTel Aviv, Israel, Tel Aviv-Yafo, Israel, מהמשרדסניורמשרה מלאה
משרה חיצונית, ההגשה באתר החברהאושר שהמשרה פתוחה לפני 9 שעות

Own evaluation for AI agents, from designing benchmarks and performance metrics to diagnosing failures and building CI regression suites. Lead research projects and set the research agenda for agent performance.

מה תעשו

  • Own evaluation for code Q&A, change-impact analysis, incident response, and MCP tool use by coding agents.
  • Design evaluation harnesses and benchmarks using merged PRs, production traces, and code as ground truth.
  • Define and track agent performance metrics, including recall, precision, cost, tool-call efficiency, latency, and failure modes.
  • Run controlled comparisons across models, prompts, retrieval strategies, and baselines.
  • Diagnose agent failures and build CI regression suites to score changes before release.
  • Lead external research, own research papers, and set the research agenda.

דרישות

  • 7+ years in ML/NLP research or applied AI, including evaluating LLMs or agents.
  • Track record designing benchmarks, including task sampling, contamination control, scoring methods, and power analysis.
  • Deep knowledge of agent architectures, including tool use, retrieval, multi-step planning, and MCP or similar protocols.
  • Expert Python and experience shipping production-quality evaluation infrastructure.
  • Experience leading research projects end to end, from design to publication or product.

יתרון

  • Publications at NeurIPS, ICLR, ACL, ICSE, or FSE on LLM evaluation, code intelligence, or agents.
  • Experience with repo-level code benchmarks such as SWE-bench.
  • Experience with LLM-as-judge methods and their known failure modes.
  • Background in distributed systems, observability data, or large codebases.

תנאי סף

  • 7+ years in ML/NLP research or applied AI
  • Expert Python

הטבות

  • Enterprise ground truth: thousands of repos, live traces, and production incidents.
  • A founding role in the research function.
PythonLLM evaluationAgent architecturesBenchmark designRetrievalMCPStatistical analysis