About the Role
Partner with engineering teams to accelerate the development and maturity of the Scale Generative AI Platform (SGP). Own strategic alignment and end-to-end execution of critical infrastructure initiatives, serving as the communication backbone between platform engineering, product teams, and executive leadership. Translate architectural complexities into clear execution strategies, unblock engineering bottlenecks, mitigate deployment risks, and ensure reliable, performant, and secure systems for global-scale deployment.
Responsibilities
- Lead strategic planning and high-velocity execution for SGP core capabilities.
- Manage features from technical scoping and architecture design through production launch.
- Drive execution and manage complex technical dependencies across systems engineering, Core ML, Research, and Product teams.
- Translate complex infrastructure metrics into actionable roadmaps.
- Map demands like multi-tenancy, data privacy, and isolation into platform features.
- Proactively identify, track, and mitigate technical risks unique to massive-scale GenAI infrastructure and global SGP deployments.
- Establish lightweight agile processes that empower engineers to ship fast without breaking core systems.
- Define and enforce clear SLOs and performance benchmarks to guarantee production-grade reliability.
- Track and report on SGP adoption metrics, system reliability, delivery forecasts, and engineering bottlenecks to executive leadership.
Requirements
- 5+ years of experience as a Technical Program Manager, Product Manager, or Software Engineer with a proven track record of building and shipping technical products or platforms from scratch.
- 3+ years of dedicated experience managing programs focused on core engineering infrastructure, cloud-native ecosystems (AWS/GCP), container orchestration (Kubernetes), or distributed systems.
- Foundational understanding of the infrastructure required for the Generative AI lifecycle, including high-throughput data pipelines, GPU/CPU cluster utilization, or model training/evaluation setups.
- Proven track record of presenting to and influencing executive-level stakeholders.
- Ability to translate complex technical/architectural challenges into clear business impacts.
- Advanced proficiency with iterative development methodologies and modern project management tooling (Linear, Jira, etc.) applied to foundational infrastructure environments.
- Strong software engineering fundamentals, with prior professional experience as a Software Engineer, DevOps Engineer, or Data Developer.
- Proven success driving the internal adoption of technical platforms, SDKs, or APIs.
- Direct experience working with large-scale data quality pipelines, distributed vector databases, or specialized AI inference engines (e.g., Triton, Ray).
Skills
- Technical Program Management
- Product Management
- Software Engineering
- Platform Development
- Developer Tooling
- Distributed Systems
- Infrastructure Initiatives
- Communication
- Agile Methodologies
- Project Management Tooling (Linear, Jira)
- Cloud-Native Ecosystems (AWS/GCP)
- Container Orchestration (Kubernetes)
- Generative AI Infrastructure
- Data Pipelines
- GPU/CPU Cluster Utilization
- Model Training/Evaluation
- Executive Stakeholder Management
- Technical Risk Mitigation
- SLO Definition
- Performance Benchmarking
- Platform Adoption
- Data Quality Pipelines
- Vector Databases
- AI Inference Engines (Triton, Ray)
Experience Level
- 5+ years of experience as a Technical Program Manager, Product Manager, or Software Engineer
- 3+ years of dedicated experience managing programs focused directly on core engineering infrastructure, cloud-native ecosystems (AWS/GCP), container orchestration (Kubernetes), or distributed systems.
About the Company
- At Scale, our mission is to develop reliable AI systems for the world's most important decisions.
- Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact.
- We work closely with industry leaders like Meta, Ernst & Young, Mayo Clinic, Time Inc., the Government of Qatar, and U.S. government agencies including the Army and Air Force.
- We are expanding our team to accelerate the development of AI applications.
Equal Opportunity
- We believe that everyone should be able to bring their whole selves to work, which is why we are proud to be an inclusive and equal opportunity workplace.
- We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability status, gender identity or Veteran status.
- We are committed to working with and providing reasonable accommodations to applicants with physical and mental disabilities.
- If you need assistance and/or a reasonable accommodation in the application or recruiting process due to a disability, please contact us at accommodations@scale.com.
- Please see the United States Department of Labor's Know Your Rights poster for additional information.
- We comply with the United States Department of Labor's Pay Transparency provision.
