About the Role
Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.
Responsibilities
- Design, deploy, and maintain large-scale bare-metal GPU clusters running Slurm and Kubernetes for distributed training.
- Configure, tune, and optimize Slurm schedulers, dynamic job queuing, priority topologies, and autoscaling.
- Manage Kubernetes clusters and hybrid Slurm-K8s environments for serving, data processing, and research.
- Optimize high-speed interconnects and high-throughput distributed storage to ensure continuous GPU saturation.
- Build observability pipelines to proactively detect hardware and system failures.
Requirements
- Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on experience in HPC or infrastructure engineering.
- Deep hands-on expertise administering Linux-based HPC clusters running Slurm or Kubernetes.
- Strong troubleshooting skills in low-level Linux networking, kernel parameters, hardware diagnostics, and storage systems.
- Proficiency in automation and infrastructure-as-code tools.
- Background supporting distributed deep learning frameworks.
Skills
- Slurm
- Kubernetes
- InfiniBand
- RoCE
- NCCL
- NVLink
- Lustre
- WEKA
- Ceph
- Prometheus
- Grafana
- DCGM
- Ansible
- Terraform
- Helm
- Python
- Bash scripting
- DeepSpeed
- Ray
Education Level
- Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on experience
About the Company
- Veeda AI is building the next generation of multimodal foundation world models for Physical AI.
- Tackling challenging problems at the intersection of AI, robotics, and embodied intelligence.
