[Remote] Generative AI Inference Engineer
Note: The job is a remote job and is open to candidates in USA. Stability AI is seeking a passionate Generative AI Inference Engineer to join their Inference team, focusing on creative applications of generative AI models. The role involves leading the design and development of customer-facing multi-modal ML inference systems and optimizing inference techniques for generative models.ResponsibilitiesLead efforts to drive the design, development of customer-facing multi modal ML inference systemsWork with the Platform and Inference teams on building inference systems for the next generation of models, where you will work on areas such as optimization, model tuning and deploymentPartner with leading cloud providers to deliver hosted Stability AI inference solutionsBe a strategic thought partner for leaders across the organization on driving business impact through machine learningBe part of the team to bring new Stability models and pipelines into existencePrototype and productionize inference platform improvements and new featuresSkills7+ years working on productionizing machine learning systems, including inference pipeline developmentExpert level knowledge on writing and running python services at scale5+ years working on python scientific stack, pyTorch and at least one high-performance inference framework (e.g. Triton and TensorRT)Deep understanding of Diffusion ArchitectureExperience profiling and optimizing deep neural networks on Nvidia GPUs, using profiling tools such as NVIDIA NsightExperience with python-based image manipulation/encoding/decoding frameworks, such as OpenCVExperience deploying to cloud orchestration systems such as Kubernetes and cloud providers such as AWS, GCP, and AzureExperience with DockerAbility to rapidly prototype solutions and iterate on them with tight product deadlinesStrong communication, collaboration, and documentation skillsExperience with the open-source ML ecosystem (HuggingFace, W&B, etc.)Company OverviewStability AI is an artificial intelligence company focused on developing open-source generative AI models. It is a sub-organization of Stability AI. It was founded in 2019, and is headquartered in London, England, GBR, with a workforce of 51-200 employees. Its website is https://stability.ai.