About the Role
Lead the deployment and scaling of AI solutions for financial workflows, ensuring flawless operation in demanding environments. Provide technical leadership and hands-on expertise to maintain secure, reliable, and high-performing infrastructure. Optimize deployments for massive scale across cloud and on-prem environments, implement security controls, and manage AI/ML workload challenges. Collaborate with engineering, product, and customer teams to streamline continuous delivery. Play a central role in maintaining, optimizing, and supporting systems across Google Cloud Platform, Microsoft Azure, Amazon Web Services, and on-premises environments.
Responsibilities
- Architect and maintain scalable, secure software-stack and infrastructure on major cloud providers (AWS, GCP, Azure) and on-premises environments.
- Build and support high-performance computing clusters for model training and inference.
- Design and implement robust CI/CD pipelines (Argo CD, GitHub Actions) to streamline the delivery of AI models and applications.
- Define SLOs/SLIs and implement comprehensive monitoring and alerting systems (Datadog, Prometheus, Grafana) to ensure high availability.
- Enforce DevSecOps best practices, managing IAM policies, network security, and compliance automation for regulated financial environments.
- Oversee the deployment and maintenance of production databases, including Postgres, Vector Stores and Graph Databases.
Requirements
- 6+ years of DevOps or SRE experience, with a strong background in supporting distributed systems at scale.
- Expert-level knowledge of Kubernetes (EKS, GKE, AKS) and Docker.
- Deep proficiency with at least one major cloud provider (AWS, GCP, or Azure) and hybrid/on-prem deployments.
- Advanced skills in Terraform or Ansible for reproducible infrastructure.
- Experience building complex pipelines with tools like Jenkins, Argo CD, or GitLab CI.
- Strong scripting skills in Bash and Python.
- Hands-on experience with modern monitoring stacks (Datadog, ELK, Prometheus/Grafana) and distributed tracing.
- Solid understanding of network security, IAM, VPC peering, and encryption standards.
- Excellent communication and collaboration abilities.
Skills
- Kubernetes
- Docker
- AWS
- GCP
- Azure
- Terraform
- Ansible
- Jenkins
- Argo CD
- GitLab CI
- Bash
- Python
- Datadog
- ELK
- Prometheus
- Grafana
- Network Security
- IAM
- VPC Peering
- Encryption
- Postgres
- Vector Stores
- Graph Databases
Location
- US
Work Type
- Full-time
Experience Level
- Senior
- 6+ years
Salary/Compensations
- Competitive compensation structure, including salary, performance-based bonuses, and additional components based on experience.
Benefits
- Comprehensive benefits
About the Company
- Domyn is a company specializing in the research and development of Responsible AI for regulated industries, including financial services, government, and heavy industry.
- It supports enterprises with proprietary, fully governable solutions based on a composable AI architecture — including LLMs, AI agents, and one of the world’s largest supercomputers.
