About the Role
Aion Silicon is seeking a Senior EDA & HPC Infrastructure Engineer to own and manage the platforms powering their engineering workloads. This hands-on role requires deep expertise in Linux-based EDA and HPC environments, combined with strong AWS cloud engineering experience, to ensure the stability, performance, and continuous improvement of critical engineering infrastructure.
Responsibilities
- Own and maintain AWS-based engineering infrastructure across a multi-account, multi-VPC environment.
- Operate, optimize, and troubleshoot cloud environments supporting engineering workloads.
- Support cloud-based job scheduling using technologies like AWS ParallelCluster and SLURM.
- Balance platform performance, capacity, and cost.
- Maintain and extend Infrastructure as Code and automation for secure, repeatable builds.
- Own, operate, and continuously improve the Linux-based EDA and HPC compute farm.
- Manage license servers, grid services, production job scheduling, and shared storage.
- Optimize job throughput, queue behavior, license availability, and storage performance in collaboration with engineering teams.
- Manage EDA tooling and licensing arrangements with vendors.
- Provide clear visibility of tool availability and license utilization.
- Diagnose and resolve networking issues across on-premises and cloud environments.
- Work across multi-VLAN, multi-account, and multi-VPC infrastructure.
- Apply security, identity, and compliance best practices.
- Support the wider IT team with service issues and infrastructure incidents.
- Cross-train colleagues in AWS and cloud-native ways of working.
- Help build shared ownership of the platform and reduce single points of knowledge.
Requirements
- Substantial experience administering production Linux-based EDA compute platforms within a semiconductor or comparable engineering environment.
- Hands-on responsibility for licensing, job scheduling, and platform performance.
- Several years of practical AWS experience.
- Experience operating cloud-first infrastructure at scale, including migrated workloads.
- End-to-end EDA license management, including license servers, feature pools, denials, and utilization analysis.
- Experience with production job schedulers such as SLURM, LSF, or Grid Engine.
- Linux administration at scale using RHEL, Rocky Linux, or similar distributions.
- Experience with OS and kernel tuning, patching, and lifecycle management.
- Experience with storage performance tuning for I/O-intensive and metadata-heavy workloads, including NFS and cloud-native alternatives.
- Experience with production AWS environments using EC2, EBS, S3, VPC networking, and IAM.
- Experience working within a multi-account AWS estate.
- Experience with Infrastructure as Code using Terraform, CloudFormation, or similar technologies.
- Experience with automation using Python, Bash, or equivalent scripting languages.
- Experience with source-controlled infrastructure and automation integrated into CI pipelines.
- Experience with Layer 2 and Layer 3 network fault diagnosis across VLANs, routing, and switching.
- Experience with monitoring and alerting across compute, storage, and license utilization.
- Ability to communicate technical issues and trade-offs clearly to engineers, vendors, and senior stakeholders.
- Willingness to support both routine service requests and complex infrastructure incidents.
Skills
- Linux-based EDA compute platforms
- High-performance computing (HPC) environments
- AWS cloud engineering
- AWS ParallelCluster
- SLURM
- Infrastructure as Code
- Automation
- Python
- Bash
- Terraform
- CloudFormation
- CI pipelines
- NFS
- EC2
- EBS
- S3
- VPC networking
- IAM
- Layer 2 networking
- Layer 3 networking
- Monitoring
- Alerting
Location
- Remote
Work Type
- Full-time
- Permanent
Experience Level
- Senior
About the Company
- Aion Silicon is a company focused on building world-class engineering solutions.
