About the Role
We are seeking an engineer with expertise in low-level systems programming and optimization to enhance the performance of our machine learning models, focusing on both training and inference within our dynamic trading environment.
Responsibilities
- Optimize the performance of ML models for training and inference.
- Improve CUDA performance.
- Apply a whole-systems approach to optimization, considering storage, networking, and host/GPU levels.
- Analyze and ensure the efficiency of throughput and actual useful output (goodput).
- Investigate low-level performance details, such as cache latency.
Requirements
- Experience in low-level systems programming and optimization.
- Systems knowledge to debug end-to-end performance of training runs.
- Background in Infiniband, RoCE, GPUDirect, PXN, rail optimization, and NVLink.
- Understanding of collective algorithms for distributed GPU training in NCCL or MPI.
- An inventive approach and willingness to question existing methods and tools.
- Fluency in English.
Skills
- Modern ML techniques and toolsets
- Low-level GPU knowledge (PTX, SASS, warps, cooperative groups, Tensor Cores, memory hierarchy)
- Debugging and optimization using tools like CUDA GDB, NSight Systems, NSight Compute
- Library knowledge of Triton, CUTLASS, CUB, Thrust, cuDNN, cuBLAS
- Intuition about latency and throughput characteristics of CUDA graph launch, tensor core arithmetic, warp-level synchronization, and asynchronous memory loads
Location
- Remote
Work Type
- Full-time
Experience Level
- Mid-level
- Senior
About the Company
- Machine learning is a critical pillar of Jane Street's global business.
- Our trading environment serves as a unique, rapid-feedback platform for ML experimentation.
- We encourage individuals with a curious mind and a passion for solving interesting problems.
