All jobs

Senior HPC Cluster Engineer

Nebius Sourced

On-site Full-time Not specified

About the role

<div class="content-intro"><p><strong>About Nebius:</strong></p> <p>Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.</p> <p>Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.</p> <p>Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&amp;D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&amp;D.</p></div><h3><strong><span data-contrast="auto">The role</span></strong></h3> <p>We’re looking for a&nbsp;Senior HPC Cluster Engineer&nbsp;to join our team and play a key role in the development of our cutting-edge hyperscaler platform. The&nbsp;GPU &amp; InfiniBand team&nbsp;is responsible for enhancing and optimizing the core components of our Cloud platform, with a specific focus on&nbsp;GPU computing,&nbsp;InfiniBand networks, and the&nbsp;KVM/QEMU stack. You’ll work closely with hardware virtualization and device emulation technologies, ensuring high performance and security in multi-GPU, HPC environments. The role involves analyzing, troubleshooting, and improving infrastructure to support new hardware, fine-tuning system performance, and automating fault detection and resolution in a complex system.</p> <p>&nbsp;</p> <p><strong>In this position, you will be responsible for:</strong></p> <ul> <li><strong>Tuning the performance</strong> of GPU clusters and InfiniBand networks to&nbsp;ensure optimal operation in HPC and GPU-based environments.
</li> <li><strong>Analyzing and troubleshooting</strong> the root cause of issues related to GPUs and InfiniBand networks, and proposing corrective actions.
</li> <li><strong>Integrating new hardware</strong> into the existing infrastructure, including&nbsp;support for new GPU hardware through software stacks like Kubernetes, QEMU, and KVM.
</li> <li><strong>Enhancing automation systems </strong>for proactive monitoring, detecting, and resolving issues in GPU and InfiniBand environments.
</li> <li><strong>Configuring and managing GPU devices</strong> and InfiniBand fabrics, ensuring efficient and reliable operation.<strong><br></strong></li> </ul> <p>&nbsp;</p> <p><strong>We expect you to have:</strong></p> <ul> <li>5+ years of professional experience in&nbsp;<strong>system-level software development</strong>&nbsp;(focused on performance optimization, low-level programming).
</li> <li>3+ years of hands-on experience with&nbsp;<strong>Linux systems</strong>&nbsp;(administration, troubleshooting, and performance tuning).
</li> <li><strong>In-depth understanding</strong> of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing (HPC) systems.
</li> <li>Strong proficiency in one or more&nbsp;<strong>performance-oriented programming languages</strong>&nbsp;(C/C++, Go, Python).</li> </ul> <p>&nbsp;</p> <p><strong>It would be a plus if you have:</strong></p> <ul> <li>Experience with&nbsp;<strong>GPU end-to-end testing</strong>&nbsp;in a&nbsp;<strong>cluster environment</strong>&nbsp;using InfiniBand networking.
</li> <li>Proven track record of analyzing and optimizing the performance of&nbsp;<strong>HPC workloads</strong>&nbsp;(e.g., simulations, data analysis, AI/ML workloads).
</li> <li>Familiarity with&nbsp;<strong>RDMA, RoCE, and InfiniBand</strong>&nbsp;protocols for high-performance communication.
</li> <li>Background in&nbsp;<strong>Software-Defined Networking</strong>&nbsp;(SDN) and experience with&nbsp;<strong>HPC cluster networking</strong>.
</li> <li>Understanding of&nbsp;<strong>QEMU/KVM virtualization</strong>&nbsp;and ma

Skills

System Engineers

Apply to Senior HPC Cluster Engineer at Nebius

Hyrovo matches you to jobs worldwide and helps you apply. Browsing, matching, and applying are free; AI-written CVs and cover letters are pay-as-you-go.

Apply now