About the Role
G-Research is seeking an exceptional NLP Performance Engineer to join their NLP Engineering team, focusing on large-scale LLM inference performance. This role involves designing and implementing techniques to improve the performance, cost-efficiency, and capabilities of inference workloads on cutting-edge compute infrastructure, enabling researchers and engineers to maximize the use of current and future systems.
Responsibilities
- Profiling, benchmarking and optimising large-scale LLM inference workloads across our compute infrastructure
- Ensuring efficient deployment of the latest models across a range of GPU architectures, adapting the inference stack as hardware evolves
- Designing and implementing inference optimisations while maintaining output quality
- Developing reference implementations, libraries and tooling to improve the efficiency and reliability of NLP workloads
- Collaborating with researchers, senior stakeholders and engineers to design optimised solutions
- Working with systems, architecture and platform teams to evolve the compute stack and influence long-term platform decisions
Requirements
- Proven experience profiling, benchmarking and optimising large-scale LLM inference workloads
- A scientific, evidence-led approach to performance optimisation, using rigorous benchmarking and reproducible measurement
- Deep understanding of transformer inference, including prefill versus decode, KV-cache behaviour, attention variants and performance bottlenecks
- Hands-on experience with LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM or TGI, and the PyTorch ecosystem
- Experience with inference optimisation techniques, including quantisation, speculative decoding and model parallelism across modern GPU architectures
- Strong software engineering skills, including Python, CUDA and building reliable systems for machine learning workloads
- Strong communication skills, with the ability to collaborate across research, infrastructure and engineering teams
Skills
- LLM inference
- Natural Language Programming (NLP)
- Performance optimisation
- Transformer inference
- LLM serving frameworks (vLLM, SGLang, TensorRT-LLM, TGI)
- PyTorch ecosystem
- Quantisation
- Speculative decoding
- Model parallelism
- Python
- CUDA
- Building reliable systems for machine learning workloads
- Communication
Location
- London
Work Type
- Full-time
Experience Level
- Quantitative Developer
Education Level
- Bachelor’s, Master’s or PhD in computer science, or equivalent experience
Salary/Compensations
- Highly competitive compensation plus annual discretionary bonus
Benefits
- Lunch provided (via Just Eat for Business)
- Dedicated barista bar
- 35 days’ annual leave
- 9% company pension contributions
- Informal dress code
- Excellent work/life balance
- Comprehensive healthcare
- Life assurance
- Cycle-to-work scheme
- Monthly company events
About the Company
- G-Research tackles complex problems in quantitative finance by bringing scientific clarity to financial complexity.
- They unite world-class researchers and engineers in an environment that values deep exploration and methodical execution.
- They are building a world-class platform to amplify their teams' most powerful ideas.
Equal Opportunity
- G-Research is committed to cultivating and preserving an inclusive work environment.
- We are an ideas-driven business and we place great value on diversity of experience and opinions.
- We want to ensure that applicants receive a recruitment experience that enables them to perform at their best. If you have a disability or special need that requires accommodation please let us know in the relevant section
