All jobs

Senior ML Engineer (Token Factory)

Nebius Sourced

Czech Republic; Remote - Europe; United Kingdom, United Kingdom Full-time Not specified

About the role

<div class="content-intro"><p><strong>About Nebius:</strong></p> <p>Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.</p> <p>Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.</p> <p>Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&amp;D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&amp;D.</p></div><h2 id="The-role" data-renderer-start-pos="1"><strong data-renderer-mark="true">The role</strong></h2> <p data-renderer-start-pos="11">Token Factory is a part of Nebius Cloud, one of the world’s largest GPU clouds, running tens of thousands of GPUs. We are building an inference &amp; fine-tuning platform that makes every kind of foundation model — text, vision, audio, and emerging multimodal architectures — fast, reliable, and effortless to train &amp; deploy at massive scale.</p> <div class="ewa-rteLine"><strong>Some directions we currently working on and which you can be a part of:</strong></div> <ul class="ak-ul" data-indent-level="1"> <li> <div class="ewa-rteLine"> <p><strong>Advanced Fine-Tuning:</strong> Enhancing fine-tuning methodologies - both LoRA-based and full-parameter - for cutting-edge LLMs (e.g., GPT-OSS, Kimi K2.5, DeepSeek V3.1/V3.2, GLM-4.7), focusing on both model quality and training efficiency.</p> </div> </li> <li> <div class="ewa-rteLine"><strong>Inference Optimization:</strong> Identifying LLM inference bottlenecks to drive production speedups. This involves building model training and evaluation pipelines in JAX for speculative decoding, experimenting with architectures (dense/MoE, auto-regressive/parallel), and deriving scaling laws to guide resource allocation.&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</div> </li> <li><strong>Low&nbsp;</strong><strong>Precision Training &amp; Inference:</strong> Investigating low-precision (FP8, NVFP4/MXFP4) methodologies for supervised fine-tuning and reinforcement learning - spanning both inference and training - optimized for modern hardware</li> </ul> <p><strong data-renderer-mark="true">We expect you to have:</strong></p> <ul class="ak-ul" data-indent-level="1"> <li> <p data-renderer-start-pos="927">A profound understanding of theoretical foundations of machine learning and reinforcement learning.</p> </li> <li> <p data-renderer-start-pos="1030">Deep expertise in modern deep learning for language processing and generation</p> </li> <li> <p data-renderer-start-pos="1111">Experience with training large models on multiple computational nodes</p> </li> <li> <p data-renderer-start-pos="1196">Reasonable understanding of performance aspects of large neural network training (sharding strategies, custom kernels, hardware features etc.)</p> </li> <li> <p data-renderer-start-pos="1342">Strong software engineering skills (we mostly use Python)</p> </li> <li> <p data-renderer-start-pos="1403">Deep experience with modern deep learning frameworks (we use JAX)</p> </li> <li> <p data-renderer-start-pos="1472">Proficiency in contemporary software engineering approaches, including CI/CD, version control and unit testing</p> </li> <li> <p data-renderer-start-pos="1586">Strong communication and leadership abilities</p> </li> </ul> <p><strong data-renderer-mark="true">Nice to have:</strong></p> <ul class="ak-ul" data-inden

Skills

ML

Apply to Senior ML Engineer (Token Factory) at Nebius

Hyrovo matches you to jobs worldwide and helps you apply. Browsing, matching, and applying are free; AI-written CVs and cover letters are pay-as-you-go.

Apply now