← Careers

Member of Technical Staff, Platform

San Francisco or Seattle · Full time · GPU compute and inference platform (YC S26)

Before you apply: a demo video is optional, and everything you write and record must be in English. You will need to rate your English 1 to 6, and we are looking for 4 or higher.

About OpenRelay

OpenRelay is GPU compute and LLM inference for people who ship. Customers rent GPU machines and call inference APIs; providers plug their hardware into our network and get paid for it. We run a distributed fleet on real hardware, not a reseller skin over someone else's cloud.

We are a small team backed by Y Combinator. Everyone here owns large, important systems that are live in production.

About the role

Our platform turns a customer's API call into a running GPU workload and keeps it running. You will build and own core pieces of it: the services that schedule work, the workflows that have to survive any failure, and the cloud infrastructure underneath.

This is a high-autonomy role. There is no ticket queue. You will find the most important problem, decide how to solve it, ship it to production, and own the result. You will work directly with the founders and help set our technical direction.

Responsibilities

  • Design, build, and operate the backend services that run customer GPU workloads
  • Own long-running workflows that must finish correctly through crashes, retries, and partial failures
  • Manage cloud infrastructure as code, and keep it secure, reliable, and cost-efficient
  • Make architectural decisions that other engineers build on
  • Own the reliability of what you build: monitoring, alerting, incident response, and the fix that stops it happening again
  • Decide what to build next, and say no to work that does not matter

Minimum qualifications

  • 3-4+ years of industry experience building and running backend or infrastructure systems in production
  • Strong Go
  • Production experience with AWS, Terraform, and Docker
  • Experience with Temporal or another durable workflow engine
  • A track record of independently scoping and shipping complex, ambiguous projects
  • You do your best work with high autonomy: you scope your own projects, ship them, and own the outcome
  • Clear written communication in English

Strong candidates may also have

  • HashiCorp tools beyond Terraform, such as Nomad, Consul, or Vault
  • Experience with GPU or high-performance compute infrastructure
  • Experience at an early-stage startup

Representative projects

  • Design the workflow that takes a GPU request from API call to running machine, and make it correct through every failure along the way
  • Make provisioning faster and more predictable, and prove it with production numbers
  • Build the tooling that lets a small team run a growing fleet without growing the on-call load
  • Lead an incident review, then ship the change that removes the whole class of failure

The video

Optional, five minutes or less, in English. If you record one: your face on camera introducing yourself, and a screen recording walking through a system you built and run and the hardest problem it gave you. Casual is fine, content matters more than polish.

Logistics

  • Location: San Francisco or Seattle
  • Education: no degree required, equivalent experience counts
  • Compensation: competitive salary and equity

We encourage you to apply even if you do not meet every qualification listed. Strong candidates rarely match all of them.