Senior Research HPC Engineer at Science and Technology Facilities Council | GB | Rezi

Senior Research HPC Engineer at Science and Technology Facilities Council

Senior Research HPC Engineer

Science and Technology Facilities Council · GB

5 days ago

Senior Research HPC Engineer

Science and Technology Facilities Council · GB

6 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Senior Research HPC Engineer role.

Rezi rewrites your resume against Science and Technology Facilities Council's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Senior Research HPC Engineer posting at Science and Technology Facilities Council — free, in seconds.

About the Role

The Senior Research HPC Engineer will develop and deliver Scientific Computing infrastructure, software provisions, and resilient services in line with strategic and scientific goals. This role involves working closely with researchers to translate requirements into scalable computational solutions, supporting HPC utilization, and building/maintaining reproducible runtime environments.

Responsibilities

  • Develop and deliver Scientific Computing infrastructure, software provisions, and resilient services in line with strategic and scientific goals.
  • Work closely with researchers to develop reusable tools and translate experimental and analytical requirements into scalable computational solutions.
  • Support researchers in utilising HPC, including triaging user issues and translating common pain points into platform improvements.
  • Help build and maintain reproducible runtime environments, container images, and workflow-supporting services for scientific computing workloads.
  • Assume responsibility for coordination and provision of HPC user training and onboarding.
  • Ensure best practices in project management are maintained for HPC projects.
  • Monitor, report on, and promote the Institute’s Scientific Computing provisions.
  • Design and manage communication strategies for effective stakeholder engagement.
  • Develop and deliver training, documentation, and onboarding materials.
  • Network with similar organisations to benchmark services and activities.
  • Research developing trends to ensure future provision and cost-effective delivery of scientific goals.
  • Assist with the scoping, provisioning, configuration, scaling, and validation of compute, storage, networking, and platform services within a high-performance computing infrastructure.
  • Support and implement the packaging, versioning, and validation of scientific software.
  • Establish policies and systems to ensure that workloads placed on the HPC systems are managed, prioritised, and run to achieve optimal service levels and uptime.
  • Install, configure, and manage Linux applications manually and using deployment software.
  • Resolve workflow performance issues on the cluster, and help to design HPC jobs.
  • Implement and maintain management and monitoring tools, report usage levels and service level data.
  • Streamline and automate maintenance, deployment, and configuration tasks.
  • Assume responsibility for documenting and implementing relevant disaster recovery processes.
  • Ensure awareness of likely points of failure, proactive maintenance, and repair of HPC resources.
  • Maintain awareness of emerging HPC software security vulnerabilities and apply relevant security patches.
  • Support the Head of IT in implementing and maintaining HPC related usage policies.
  • Support LMS IT staff for issues and queries relating to standalone Linux/Unix systems.
  • Other duties commensurate with the grade of the post as directed by the supervisor.

Requirements

  • A degree in a computing/scientific research/engineering subject with a significant computational component, or equivalent skills and experience.
  • Significant experience of working at a high-level within a relevant area.
  • Hands-on experience supporting and administering Linux-based systems in an HPC, research, academic, or production environment.
  • Experience in configuration and maintenance of multi-queue job scheduling systems (e.g. SLURM).
  • Experience in the use of scientific software compilation and deployment systems (e.g. Spack, EasyBuild, Lmod, conda).
  • Virtualisation and containerisation deployment and management (e.g. Docker, Singularity).
  • Ability to work closely with multidisciplinary research teams, understand scientific computing needs, and deliver practical services that advance scientific goals.
  • Knowledge of scripting (e.g., Bash) and version control systems (e.g., Git/GitHub) to support reproducible and collaborative research.
  • Integration of heterogeneous Linux/Windows/Mac environments (e.g. Active Directory).
  • Ability to troubleshoot issues across systems, networking, storage, identity, containers, schedulers, and user workloads, and to follow problems through to a reliable operational fix.
  • Experience with automation tooling (e.g. Salt, xCAT).
  • Maintain awareness of HPC cyber security principles, best practices, and emerging threats.
  • High attention to detail and accuracy to effectively analyse and interpret complex data and use it solve to complex technical problems quickly and effectively.
  • Excellent verbal and written communication skills.
  • Able to self-motivate when working independently, on projects, and collaboratively within a team.
  • Effectively plans, multitask, and prioritises workload to achieve results, adapting quickly while maintaining control in challenging situations.
  • Ensures appropriate engagement of colleagues with relevant stakeholders and issues.
  • Ability to develop and encourage cross-boundary working relationships and to work collaboratively.
  • Contribute positively to IT and cross-functional teams and shares knowledge and skills.

Skills

  • HPC
  • Scientific Computing
  • Linux
  • Containerisation (Docker, Singularity)
  • Scripting (Bash)
  • Version Control (Git/GitHub)
  • Automation Tooling (Salt, xCAT)
  • Job Scheduling Systems (SLURM)
  • Software Compilation and Deployment (Spack, EasyBuild, Lmod, conda)
  • Project Management
  • Troubleshooting
  • Cyber Security Principles
  • Communication
  • Teamwork
  • GPU Computing
  • AI/ML
  • Bioinformatics
  • Scientific Workflow Frameworks (Nextflow, Snakemake, WDL/Cromwell)
  • CI/CD pipelines
  • IaC tools (Ansible, Terraform)
  • Large-scale biological data types
  • Identity, access, and security controls
  • C programming
  • R programming
  • Python programming
  • Monitoring and Optimising Workloads (Grafana, Prometheus, Arbiter2)

Location

  • Hammersmith London

Work Type

  • Permanent
  • Full Time

Experience Level

  • Senior
  • Significant relevant experience

Education Level

  • Degree in a computing/scientific research/engineering subject with a significant computational component, or equivalent skills and experience.
  • Industry-standard Linux certifications (e.g. RHCA, RHCSA) - Desirable
  • Current qualifications in project management (e.g. PRINCE2) - Desirable
  • Knowledge, experience, or qualification in cloud computing - Desirable
  • Knowledge or education to a higher level within a scientific field - Desirable

Salary/Compensations

  • £57,623 - £61,622 plus London Allowance £5,560 per annum
  • £61,622 - £67,693 plus London Allowance of £5,560 per annum (for candidates with significant relevant experience)

Benefits

  • Defined benefit pension scheme
  • Excellent holiday entitlement
  • Employee shopping/travel discounts
  • Salary sacrifice cycle to work scheme

About the Company

  • UKRI is an organisation that brings together the seven disciplinary research councils, Research England and Innovate UK, building an independent organisation with a strong voice and vision ensuring the UK maintains its world-leading position in research and innovation.
  • Supporting some of the world’s most exciting and challenging research projects, we develop and operate some of the most remarkable scientific facilities in the world.
  • The Laboratory Institute of Medical Sciences (LMS) is a Medical Research Council funded research institute at the Hammersmith Hospital campus (London W12) of Imperial College London, with around 400 biomedical researchers undertaking multidisciplinary and internationally-competitive science.
  • The MRC is a unique working environment where our researchers are rewarded by world class innovation and collaboration opportunities.
  • The MRC is an excellent place to develop yourself further and a range of training & development opportunities will be available.

Equal Opportunity

  • At UKRI, we believe that everyone has a right to be treated with dignity and respect, and to be provided with equal opportunities to thrive and succeed in an environment that enables them to do so.
  • We also value diversity of thought and experience within inclusive groups, organisations and the wider community.
  • As users of the disability confident scheme, any candidate who opts into the scheme and best meets the essential criteria, will be shortlisted for interview.
  • We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment.
  • Our success is dependent upon our ability to embrace diversity and draw on the skills, understanding and experience of all our people.
  • We welcome applications from all sections of the community irrespective of gender, race, ethnic or national origin, religion or belief, sexual orientation, disability or age.
  • As "Disability Confident" employers, we guarantee to interview all applicants with disabilities who meet the minimum criteria for the vacancy.