Machine Learning Performance Engineer at Jane Street | New York, New York, US | Rezi

Machine Learning Performance Engineer at Jane Street

Machine Learning Performance Engineer

Jane Street · New York, New York, US

3 days ago

Machine Learning Performance Engineer

Jane Street · New York, New York, US

3 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Join our growing ML team to optimize the performance of machine learning models for training and inference. This role involves improving CUDA, storage systems, networking, and host/GPU considerations to ensure efficient large-scale training and low-latency, high-throughput inference.

Responsibilities

  • Optimize model performance for training and inference.
  • Improve CUDA performance.
  • Address storage systems, networking, and host/GPU considerations.
  • Analyze and improve system efficiency at the lowest level, including cache performance and throughput vs. goodput.
  • Debug performance issues end-to-end.
  • Utilize networking technologies like Infiniband, RoCE, GPUDirect, PXN, rail optimisation, and NVLink for GPU cluster linking.
  • Understand collective algorithms for distributed GPU training in NCCL or MPI.

Requirements

  • Experience in low-level systems programming and optimization.
  • Systems knowledge to debug training run performance end-to-end.
  • Experience debugging and optimizing using tools like CUDA GDB, NSight Systems, NSight Compute, and Nsight-Compute.
  • Background in Infiniband, RoCE, GPUDirect, PXN, rail optimisation and NVLink.
  • Understanding of collective algorithms supporting distributed GPU training in NCCL or MPI.
  • An inventive approach and willingness to question existing methods and tools.
  • Fluent in English.

Skills

  • Modern ML techniques and toolsets
  • Low-level GPU knowledge (PTX, SASS, warps, cooperative groups, Tensor Cores, memory hierarchy)
  • Triton
  • CUTLASS
  • CUB
  • Thrust
  • cuDNN
  • cuBLAS
  • CUDA graph launch latency and throughput characteristics
  • Tensor core arithmetic
  • Warp-level synchronization
  • Asynchronous memory loads
  • Infiniband
  • RoCE
  • GPUDirect
  • PXN
  • Rail optimisation
  • NVLink
  • NCCL
  • MPI

About the Company

  • Machine learning is a critical pillar of Jane Street's global business.
  • Our trading environment serves as a unique, rapid-feedback platform for ML experimentation.
  • We encourage individuals with curious minds and a passion for solving interesting problems.