Machine Learning Performance Engineer, Training at Tower Research Capital | NY, US | Rezi

Machine Learning Performance Engineer, Training at Tower Research Capital

Machine Learning Performance Engineer, Training

Tower Research Capital · NY, US

Today

Machine Learning Performance Engineer, Training

Tower Research Capital · NY, US

3 hours ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Machine Learning Performance Engineer, Training role.

Rezi rewrites your resume against Tower Research Capital's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Machine Learning Performance Engineer, Training posting at Tower Research Capital — free, in seconds.

About the Role

Bridge the gap between quantitative research and high-performance computing, building and optimizing systems for training machine learning models at scale. Focus on accelerating the end-to-end training lifecycle to enable researchers to iterate more quickly across complex models and datasets.

Responsibilities

  • Benchmark model-training workloads across CPUs, GPUs, and other accelerator platforms to identify bottlenecks and guide compute infrastructure decisions.
  • Develop performance models and standardized benchmarks for measuring throughput, utilization, scalability, and time to convergence.
  • Design and optimize distributed training strategies, including data, tensor, pipeline, and model parallelism.
  • Improve communication efficiency across multi-GPU and multi-node environments by optimizing collective operations, topology awareness, and computation–communication overlap.
  • Analyze and improve the full training pipeline, including data loading, preprocessing, memory management, forward and backward passes, optimizer execution, checkpointing, and experiment recovery.
  • Identify bottlenecks across compute, memory, storage, networking, and interconnects to increase accelerator utilization and researcher productivity.
  • Develop and optimize GPU kernels and performance-critical framework components for quantitative machine learning workloads.
  • Integrate specialized libraries, compilers, and execution techniques to improve throughput, memory efficiency, and numerical performance.
  • Apply techniques such as mixed-precision training, gradient accumulation, activation checkpointing, operator fusion, and memory-efficient optimizers.
  • Evaluate tradeoffs among training speed, numerical stability, reproducibility, model quality, and infrastructure cost.
  • Optimize workload scheduling, resource allocation, observability, fault tolerance, and reproducibility across shared compute environments.
  • Help define the architecture and tooling required to support large-scale experimentation across on-premises and cloud-based infrastructure.
  • Work closely with ML Researchers, Quantitative Researchers, HPC Engineers, Systems Engineers, and hardware specialists to translate research requirements into highly efficient training systems.

Requirements

  • 3+ years of experience optimizing machine learning training workloads in high-performance, distributed, or large-scale computing environments.
  • Deep knowledge of machine learning frameworks such as PyTorch or JAX, including their execution models, compilation paths, autograd systems, and distributed-training capabilities.
  • Strong programming skills in Python and C++, with experience developing or optimizing performance-critical systems.
  • Proven experience with GPU kernel development and optimization using technologies such as CUDA, Triton, CUTLASS, cuBLAS, cuDNN, or related libraries.
  • Strong understanding of GPU architecture, including streaming multiprocessor execution, warp scheduling, tensor cores, and the memory hierarchy from registers through HBM.
  • Experience with distributed-training technologies and communication libraries such as NCCL, FSDP, DeepSpeed, Megatron-LM, XLA, or equivalent systems.
  • Proficiency with performance-analysis tools such as Nsight Systems, Nsight Compute, PyTorch Profiler, or comparable tracing and profiling platforms.
  • Understanding of high-performance networking, storage, and accelerator interconnects, including technologies such as InfiniBand, RDMA, NVLink, or NVSwitch.
  • Demonstrated ability to benchmark heterogeneous compute platforms and make rigorous, data-driven recommendations about performance, scalability, and cost.
  • Experience optimizing training workloads for transformer-based, time-series, reinforcement-learning, or other computationally intensive models.
  • Experience with cluster orchestration and scheduling technologies such as Kubernetes, Slurm, Ray, or similar platforms.
  • Familiarity with fault-tolerant distributed training, large-scale checkpointing, experiment reproducibility, and GPU-cluster observability.
  • Practical experience with specialized accelerators, custom hardware, or compiler technologies for machine learning.
  • Prior experience in financial trading is not required.

Skills

  • PyTorch
  • JAX
  • Python
  • C++
  • CUDA
  • Triton
  • CUTLASS
  • cuBLAS
  • cuDNN
  • NCCL
  • FSDP
  • DeepSpeed
  • Megatron-LM
  • XLA
  • Nsight Systems
  • Nsight Compute
  • PyTorch Profiler
  • InfiniBand
  • RDMA
  • NVLink
  • NVSwitch
  • Kubernetes
  • Slurm
  • Ray

Location

  • New York

Work Type

  • Hybrid

Experience Level

  • 3+ years

Salary/Compensations

  • $200,000

Benefits

  • Generous paid time off policies
  • Savings plans and other financial wellness tools available in each region
  • Hybrid working opportunities
  • Free breakfast, lunch and snacks daily
  • In-office wellness experiences and reimbursement for select wellness expenses
  • Volunteer opportunities and charitable giving
  • Social events, happy hours, treats and celebrations throughout the year
  • Workshops and continuous learning opportunities

About the Company

  • Tower Research Capital is a leading quantitative trading firm founded in 1998, built on a high-performance platform and independent trading teams with a 25+ year track record of innovation.
  • Tower empowers portfolio managers to build teams and strategies independently while providing the economies of scale of a large, global organization.
  • Engineers develop electronic trading infrastructure at a world-class level, solving challenging problems in low-latency programming, FPGA technology, hardware acceleration, and machine learning.
  • Ongoing investment in engineering talent and technology ensures the platform remains unmatched in functionality, scalability, and performance.
  • Business Support teams are essential to building and maintaining the platform that powers everything, enabling trading and engineering teams to perform at their best.
  • Employees will find a stimulating, results-oriented environment where highly intelligent and motivated colleagues inspire each other.
  • Tower’s headquarters are in the historic Equitable Building in NYC’s Financial District, with a global impact and over a dozen offices worldwide.
  • Tower fosters a culture where smart, driven people thrive without egos, with an open concept workplace, casual dress code, and well-stocked kitchens reflecting a friendly, collaborative environment.
  • A collaborative and welcoming culture, a diverse team, and a workplace that values both performance and enjoyment, with no unnecessary hierarchy or ego.

Equal Opportunity

  • Tower Research Capital is an equal opportunity employer.