The best candidates would be top 1% at multiple parts of the inference stack.
work on PD disaggregation research
Morph builds the inference infrastructure behind the fastest open models. Our stack spans kernels, model serving, routing, autoscaling, and capacity. We are hiring a performance engineer to make the entire system faster, cheaper, and more reliable.
You will work directly with the founders on problems that determine how efficiently frontier-scale models can be served. Small team, enormous compute, immediate production impact.
2 day work trial
Morph builds specialized code-generation models and serves them on a custom inference stack.
Technical work involves autoresearch for kernels and custom speculative-decoding models.