E

Inference Infrastructure Engineer, Serving

Elorian Palo Alto, California, United States Full-time 5 days ago
AI & ML Engineering

$200,000 - $400,000 USD yearly

About Us

We are a well-funded, early-stage AI lab focused on building the next generation of frontier multimodal AI models. Founded by former DeepMind researchers, including Andrew Dai, who was previously a leader on Gemini. Our team currently consists of 20 world-class scientists and engineers. We recently raised $55M in seed funding from Striker Ventures, Menlo Ventures, Altimeter Capital, and NVIDIA. We are tackling some of the hardest problems in artificial intelligence, and we are growing fast.

Β 

About the Role

We're looking for an infrastructure engineer to design, optimize, and scale the systems that serve our large multimodal models. Your work will make inference faster, more cost-effective, and more reliable, so our teams can focus on advancing model capabilities rather than managing bottlenecks.

Our focus is on performant, efficient inference, both to power real-world applications and to accelerate research. This role owns the infrastructure that ensures every deployment and evaluation runs smoothly at scale for our visual foundation models.

What You Will Do

  • Build low-latency, high-throughput inference serving systems for our large multimodal models

  • Design and implement techniques that improve latency, throughput, and efficiency, including quantization, batching, speculative decoding, and KV cache management

  • Optimize our codebase and GPU fleet to fully utilize hardware FLOPs, bandwidth, and memory

  • Implement multi-GPU and multi-node model parallelism for serving (tensor or pipeline parallel)

  • Build autoscaling and load balancing for production ML services

  • Establish standards for reliability, observability, and reproducibility across the inference stack

  • Collaborate with researchers to enable high-performance inference for novel architectures

Skills and Qualifications

Minimum qualifications:

  • 3+ years of experience building low-latency, high-throughput inference serving systems for large models

  • Strong knowledge of inference optimization techniques (quantization, batching, speculative decoding, KV cache management)

  • Hands-on experience with serving frameworks such as vLLM, TensorRT-LLM, Triton, or SGLang

  • Experience with multi-GPU/multi-node model parallelism for serving (tensor or pipeline parallel)

  • Strong systems programming skills; C++/CUDA a plus alongside Python

  • Experience with autoscaling and load balancing for production ML services

  • A track record of GPU cost optimization at scale

Preferred qualifications (strong candidates may have some, not all):

  • Experience serving multimodal (vision + language) models

  • Contributions to open-source ML or systems infrastructure projects (e.g., vLLM, SGLang, TensorRT-LLM, Triton)

  • A bias for action and comfort working across stacks and teams in an early-stage environment

Logistics

  • Location: This role is based on-site in Palo Alto, California.

  • Compensation: Depending on background, skills, and experience, the expected annual base salary range for this position is $200,000 - $400,000 USD, plus equity and benefits.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: We offer health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Elorian AI is an equal opportunity employer. We are committed to building a diverse team and inclusive environment.


E

Elorian

Apply now
Palo Alto, California, United States
Full-time
$200,000 - $400,000 USD yearly
5 days ago

Share this job

Similar Jobs