HPC Engineer - Research Infrastructure @ Luma AI

Remote Full-time
Help Luma build some of the biggest & fastest AI supercomputing clusters in the world! As a High-Performance Computing engineer, you’ll work at the intersection of hardware and software, designing systems that deliver the maximum possible performance for running large-scale AI models. We work at the very cutting edge of speed and scale, combining the traditions of High-Performance Computing (HPC) in a modern cloud environment. For this role, it’s important you understand how to combine CPU’s, GPU’s, and network devices into systems that are then deployed at a large scale to peak efficiency. You understand the lowest levels of the software platforms that sit on top of this hardware, including how to best optimize the Linux kernel and user-space code. You are capable of writing code to automate the monitoring and healing of these systems, commanding a large number of servers with few people.ResponsibilitiesIn this role, you will work closely with and directly accelerate machine learning researchers, but don't need to be a machine learning expert yourself. We value people who can quickly obtain a deep technical understanding of new domains and enjoy being self-directed and identifying the most important problems to solve. You’ll be managing training HPC clusters at Luma from provisioning to performance tuning.Areas of work will include observability, distributed job tracing, GPU diagnostics, software environment management and additional tooling plus work on the actual code to enable necessary features.We believe that increasing compute is a huge lever to AI progress. You will have a direct impact on our ability to grow to an unprecedented scale and likewise produce unprecedented results.Experience8+ years experience as infrastructure engineer or Devops in large and complex distributed systems.Deep understanding of networking, bonus points for experience in HPC networking.Experience developing high-quality software in a general-purpose programming language, preferably including Python.Excellent problem-solving skills and…

Apply Now
Apply Now →

Similar Jobs

Experienced Registered Behavior Technician for In-Home ABA Therapy - Atlanta, GA

Remote

Immediate Hiring: Experienced Registered Behavioral Technician (RBT) for Clinic-Based ABA Therapy Services

Remote

Experienced Registered Behavioral Technician (RBT) - ABA Therapy for Children with Autism Spectrum Disorder

Remote

Experienced Registered Nurse - Telehealth: Providing Remote Care Coordination and Patient Support

Remote

Experienced Substitute Teacher for Riverside County Schools - Join Scoot Education's Innovative Team

Remote

Experienced Substitute Teacher for San Bernardino County - Flexible Schedules & Competitive Pay

Remote

Experienced School Year Instructional Coach for High-Dosage Tutoring Programs in Edgewater Park, NJ

Remote

Experienced School Year Tutor for K-8 Students in Math and Literacy - Mickleton, NJ

Remote

Experienced Secondary Social Studies Teacher for Kansas - Flexible Hybrid Remote Arrangement

Remote

USPS Office Helper

Remote

[Remote] Marketing Coordinator - (Remote in Indiana)

Remote

Entry-Level Live Chat Customer Service Representative – Remote Work Opportunity with Professional Growth and Development

Remote

Apply Now: Adjunct Faculty - Medical Coding

Remote

Speech Therapist (ST)

Remote

Sales Engineer

Remote

Head of Asset Management & AI Transformation

Remote

**Experienced Customer Chat Support Specialist – Hospitality Industry Expert**

Remote

CyberSecurity Analyst I

Remote

Jr. Salesforce Business Analyst [REMOTE JOB] – Amazon Store

Remote

Entry Level Fedex data entry job(Work At Home)

Remote
← Back