Back to All Positions

Senior AI Engineer

EngineeringSF Bay Area or Remotefull-time$180,000 – $240,000 + equity

About the Role

CodeConductor lets anyone — from a citizen developer to a Fortune 500 IT team — go from a prompt to a real, production-grade application in minutes. The AI layer that powers that transformation is the heart of our platform, and we're hiring a Senior AI Engineer to make it dramatically better.

You'll own the systems that turn natural language into working code, schemas, integrations, and AI agents. You'll work on prompt orchestration, fine-tuning, evaluation pipelines, retrieval, agent loops, and the model-routing logic that decides when to call which model and why. You'll ship to customers every week.

What you'll do

  • Design and ship the LLM systems that generate apps, agents, and workflows from natural-language prompts
  • Build evaluation harnesses that measure quality, latency, cost, and safety on every release
  • Tune retrieval and context strategies to make code generation accurate at scale
  • Work with the platform team to expose AI features as first-class platform primitives
  • Partner with customer-facing teams to turn real failure modes into eval suites and fixes
  • Help us decide when to use frontier models vs. smaller fine-tuned models — and prove it with data

You're a good fit if you

  • Have shipped production LLM systems that real users depend on (not just demos)
  • Are comfortable across the stack: prompt design, fine-tuning, RAG, agents, eval, observability
  • Can read a 10,000-line codebase and figure out what to change to make a model 20% better
  • Care about latency, cost, and reliability as much as you care about capability
  • Want to work on a product where AI quality is the difference between a happy customer and a refund

Requirements

  • 5+ years of software engineering experience, including 2+ years building production AI/LLM systems
  • Deep familiarity with at least one major LLM provider (Anthropic, OpenAI, Bedrock, Azure OpenAI) and their evaluation tooling
  • Strong Python or TypeScript skills; comfortable with backend services, queues, and observability
  • Experience designing evals, golden datasets, and offline/online quality measurement
  • Track record of debugging model behavior in production using traces, metrics, and user feedback
  • Bonus: experience with code generation, agentic systems, or developer tools

Nice to have

  • Open-source contributions to LLM tooling, eval frameworks, or developer platforms
  • Background in compilers, programming languages, or static analysis
  • Experience working at an early-stage startup

Benefits

  • Competitive salary and meaningful equity in a profitable, bootstrapped company
  • Comprehensive health, dental, and vision coverage
  • Flexible PTO and a culture that respects it
  • Remote-friendly with optional time at our SF Bay Area HQ
  • A small, senior team where your work ships to customers every week
  • The chance to build the AI layer of a platform that's already profitable and growing fast

Apply for this Position

By submitting, you agree to our Terms and Privacy Policy.