← Careers

Member of Technical Staff, Inference

San Francisco or Seattle · Full time · GPU compute and inference platform (YC S26)

Before you apply: a demo video is optional, and everything you write and record must be in English. You will need to rate your English 1 to 6, and we are looking for 4 or higher.

About OpenRelay

OpenRelay is GPU compute and LLM inference for people who ship. Customers rent GPU machines and call inference APIs; providers plug their hardware into our network and get paid for it. We run a distributed fleet on real hardware, not a reseller skin over someone else's cloud.

We are a small team backed by Y Combinator. Everyone here owns large, important systems that are live in production.

About the role

We serve open-weight models on GPUs we run, and throughput per GPU is our margin. You will own that serving layer: which engine runs a model, how it is configured, how requests reach it, and how it holds up when traffic spikes.

This is a high-autonomy role. There is no ticket queue. You will find the most important problem, decide how to solve it, ship it to production, and own the result. You will work directly with the founders and help set our technical direction.

Responsibilities

  • Own the performance, cost, and reliability of every model we serve
  • Bring new open-weight models to production quickly, including ones that current engine releases do not yet support
  • Tune serving for each model and GPU: parallelism, batching, caching, and quantization
  • Design how requests are routed and scheduled across replicas and nodes
  • Profile and benchmark under realistic traffic, fix what the numbers point to, and send general fixes upstream
  • Decide what to build next, and say no to work that does not move the numbers

Minimum qualifications

  • 3+ years of industry experience, with production LLM serving a real part of it
  • Hands-on production experience with vLLM, SGLang, and NVIDIA Dynamo
  • Strong Python and PyTorch, and the ability to read and change large codebases you did not write
  • A track record of measurable wins in throughput, latency, or cost, with numbers you can defend
  • You do your best work with high autonomy: you scope your own projects, ship them, and own the outcome
  • Clear written communication in English

Strong candidates may also have

  • Contributions to an open-source inference engine
  • GPU kernel or low-level performance work (CUDA, ROCm, or Triton)
  • Experience serving mixture-of-experts models or running multi-node inference
  • Experience at an early-stage startup

Representative projects

  • Get a newly released open model serving in production within days of its release, then make it fast
  • Cut cost per token on a high-traffic model by retuning parallelism, batching, and caching, and prove it with benchmarks
  • Route repeat prompts to the replicas that already hold their KV cache
  • Trace a tail-latency regression from the request down to a scheduler or kernel, and fix it

The video

Optional, five minutes or less, in English. If you record one: your face on camera introducing yourself, and a screen recording walking through a serving system you built or tuned and the numbers it hit. Casual is fine, content matters more than polish.

Logistics

  • Location: San Francisco or Seattle
  • Education: no degree required, equivalent experience counts
  • Compensation: competitive salary and equity

We encourage you to apply even if you do not meet every qualification listed. Strong candidates rarely match all of them.