About the Role
Join our growing ML team to optimize the performance of machine learning models for training and inference. This role involves improving CUDA, storage systems, networking, and host/GPU considerations to ensure efficient large-scale training and low-latency, high-throughput inference.
Responsibilities
- Optimize model performance for training and inference.
- Improve CUDA performance.
- Address storage systems, networking, and host/GPU considerations.
- Analyze and improve system efficiency at the lowest level, including cache performance and throughput vs. goodput.
- Debug performance issues end-to-end.
- Utilize networking technologies like Infiniband, RoCE, GPUDirect, PXN, rail optimisation, and NVLink for GPU cluster linking.
- Understand collective algorithms for distributed GPU training in NCCL or MPI.
Requirements
- Experience in low-level systems programming and optimization.
- Systems knowledge to debug training run performance end-to-end.
- Experience debugging and optimizing using tools like CUDA GDB, NSight Systems, NSight Compute, and Nsight-Compute.
- Background in Infiniband, RoCE, GPUDirect, PXN, rail optimisation and NVLink.
- Understanding of collective algorithms supporting distributed GPU training in NCCL or MPI.
- An inventive approach and willingness to question existing methods and tools.
- Fluent in English.
Skills
- Modern ML techniques and toolsets
- Low-level GPU knowledge (PTX, SASS, warps, cooperative groups, Tensor Cores, memory hierarchy)
- Triton
- CUTLASS
- CUB
- Thrust
- cuDNN
- cuBLAS
- CUDA graph launch latency and throughput characteristics
- Tensor core arithmetic
- Warp-level synchronization
- Asynchronous memory loads
- Infiniband
- RoCE
- GPUDirect
- PXN
- Rail optimisation
- NVLink
- NCCL
- MPI
About the Company
- Machine learning is a critical pillar of Jane Street's global business.
- Our trading environment serves as a unique, rapid-feedback platform for ML experimentation.
- We encourage individuals with curious minds and a passion for solving interesting problems.
