About the Role
As a Principal Applied AI Developer on the Foundation Model ML Infrastructure team, you will help define and accelerate the roadmap for the Autodesk Machine Learning Platform. You will design, build, and evolve resilient, secure, scalable, observable, and cost-effective platform services that support model training, inference, evaluation, deployment, and serving at global scale. This is a principal-level technical leadership role for someone who combines deep hands-on engineering expertise with the ability to define technical direction, lead complex initiatives across teams, mentor senior developers, and raise the bar for engineering excellence.
Responsibilities
- Define and drive technical strategy for Foundation Model ML Infrastructure capabilities within the Autodesk Machine Learning Platform
- Lead the design and implementation of large-scale platform services that support the full lifecycle of Autodesk’s ML models, including training, inference, serving, evaluation, deployment, monitoring, and operations
- Architect highly resilient, secure, observable, scalable, and cost-effective infrastructure for large-scale AI and ML workloads
- Build and evolve developer-facing APIs, tools, workflows, and self-service capabilities that enable researchers and ML developers to move quickly and safely
- Work hands-on with Kubernetes, Ray, SageMaker, AWS, and related cloud-native technologies to support distributed training, scalable inference, and production model serving
- Identify, frame, and prioritize high-impact technical problems aligned with product, research, and platform strategy
- Translate ambiguous AI research goals, product needs, and business requirements into practical technical designs and executable engineering plans
- Lead complex cross-team technical initiatives, align stakeholders, and influence technical direction without requiring direct authority
- Drive reliability, scalability, performance, security, quality, and cost improvements across training, inference, and serving workloads
- Establish and evolve platform standards for production readiness, observability, SLAs/SLOs, incident response, release quality, model deployment, versioning, lineage, and governance
- Partner with researchers, ML developers, product managers, architects, security, privacy, and platform teams to define quality bars and safe production deployment practices, including Trusted AI requirements
- Improve developer productivity through CI/CD, automated testing, infrastructure as code, contract testing, quality gates, documentation, and platform automation
- Lead root-cause analysis for systemic production issues and implement durable, platform-level improvements
- Act as a technical authority for critical decisions, guiding trade-offs across performance, reliability, security, cost, scalability, and developer experience
- Mentor senior developers, elevate engineering standards, and foster a culture of ownership, quality, action, and accountability
- Actively participate in Agile, Kanban, or other modern development methodologies to deliver high-quality outcomes incrementally
Requirements
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Machine Learning, or equivalent practical experience
- 8+ years of professional software engineering experience, including significant experience with large-scale, cloud-native, distributed, platform, or machine learning infrastructure systems
- Experience using AI-assisted development tools, coding agents, and AI-powered automation to improve engineering productivity, with practical understanding of context design, human review, testing, secure usage, and integration into developer workflows
- Deep hands-on experience designing, building, and operating production-grade services that support model training, inference, serving, evaluation, deployment, or observability
- Strong experience with Kubernetes and cloud-native infrastructure
- Experience with distributed compute, ML infrastructure, or model serving technologies such as Ray, SageMaker, distributed training platforms, inference serving platforms, or equivalent systems
- Proven ability to lead complex technical initiatives across teams and influence technical direction without direct authority
- Experience translating ambiguous research, product, or business requirements into practical technical designs and executable engineering plans
- Strong experience designing and operating resilient, secure, observable, and cost-effective production systems using CI/CD, automated testing, infrastructure as code, monitoring, alerting, and production operations practices
- Demonstrated ability to mentor developers, elevate engineering standards, and act as a strong technical voice for excellence
- Strong written and verbal communication skills, with the ability to influence technical and non-technical stakeholders
Skills
- Kubernetes
- Ray
- SageMaker
- AWS
- Cloud-native technologies
- Distributed compute
- ML infrastructure
- Model serving technologies
- CI/CD
- Automated testing
- Infrastructure as code
- Monitoring
- Alerting
- Production operations
Location
- Canada
Work Type
- Full-time
Experience Level
- Principal
- 8+ years
Education Level
- Bachelor's degree
- Master's degree
Salary/Compensations
- $153,000 - $224,400
Benefits
- Annual cash bonuses
- Commissions for sales roles
- Stock grants
- Comprehensive benefits package
About the Company
- Autodesk creates software tools for designing buildings, machines, products, infrastructure, and entertainment, empowering creative individuals worldwide.
- Autodesk is building cloud-scale software, data platforms, and AI-enabled capabilities that help customers design, build, and operate the world around them.
- Autodesk fosters a culture of belonging where everyone can thrive.
- Amazing things are created every day with Autodesk's software – from the greenest buildings and cleanest cars to the smartest factories and biggest hit movies.
- Autodesk helps innovators turn their ideas into reality, transforming not only how things are made, but what can be made.
