About the Role
We are seeking a Senior Machine Learning Infrastructure Engineer to design, build, and scale infrastructure for massive-scale data, modeling, and analysis platforms. This role is critical in shaping a high-performance, production-grade ML ecosystem to support rapid experimentation with diverse datasets. You will have significant ownership over the ML R&D platform, working closely with domain experts to architect cloud infrastructure, data pipelines, and modeling flows, ultimately enabling cutting-edge models for neuroscientific discovery and neural decoding.
Responsibilities
- Create flexible and performant ML infrastructure
- Design and build ML cloud infrastructure for massive-scale modeling and analytics
- Support diverse model exploration, hyperparameter optimization, pretraining, fine-tuning, and evaluation processes
- Design and optimize scalable distributed training pipelines, with support for model sharding, cross-GPU communication, and real-time training monitoring
- Create, operate, and maintain robust ML platforms and services across the model lifecycle
- Make informed architecture decisions balancing performance, cost, reliability, and scalability
- Build diverse and scalable data platforms
- Design, build, and optimize massive-scale databases and data pipelines for scalable, flexible, and reliable data access
- Explore research-driven, tailored data solutions using existing and simulated data
- Create infrastructure and pipelines for ingesting internal and external datasets with varied shapes, formats, and associated metadata
- Design and assess custom data formats for efficient storage and slicing of high-dimensional time-series data
- Enable efficient data movement, preprocessing, and artifact management for data lineage and modeling reproducibility
- Meet company standards for delivered solutions
- Establish best practices for reliability, observability, reproducibility, and operational excellence across the ML ecosystem
- Make informed and collaborative decisions with domain experts across software & ML teams
- Foster visibility and reproducibility by maintaining extensive documentation of design decisions, evaluations, and pipeline assessments
- Support ML R&D operations while preparing for eventual incorporation into product pipelines
Requirements
- Bachelor's degree in Computer Science, Electrical Engineering, or a related technical discipline
- 5+ years of industry experience in software engineering, large-scale data infrastructure, or systems ML
- Extensive proficiency in Python
- Familiarity with PyTorch
- Experience designing, building, and maintaining high-throughput data pipelines for large and diverse datasets
- Experience working with distributed-training frameworks (e.g. FSDP, DeepSpeed, Megatron-LM, Ray, etc.)
- Experience building or optimizing ML training pipelines for transformers or other large neural-network models
- Demonstrated ability to partner closely with research and modeling teams to productionize workflows
- Excellent communication and collaboration skills to work effectively on cross-functional and interdisciplinary teams
- Experience having technical ownership over at least one successfully implemented collaborative project
- Advanced degree (MS or PhD) in Computer Science, Electrical Engineering, or a related technical discipline (Preferred)
- Proficiency in C++, Go, CUDA, Rust, and/or Java (Preferred)
- Experience in data engineering and systems ML for time-series data (Preferred)
- Deep understanding of the fundamentals of distributed systems, including scalability, fault tolerance, monitoring, observability, scheduling, performance tuning, and resource management (Preferred)
- Experience with cloud-native environments and orchestration (Kubernetes, Docker, etc.) (Preferred)
- Experience scaling foundation-model training infrastructure or multi-cluster computing environments (Preferred)
Skills
- Python
- PyTorch
- Distributed-training frameworks (FSDP, DeepSpeed, Megatron-LM, Ray)
- ML training pipelines for transformers or large neural-network models
- C++
- Go
- CUDA
- Rust
- Java
- Data engineering
- Systems ML for time-series data
- Distributed systems fundamentals
- Cloud-native environments
- Orchestration (Kubernetes, Docker)
Location
- Remote
Work Type
- Full-time
Experience Level
- Senior
- 5+ years of industry experience
Education Level
- Bachelor's degree in Computer Science, Electrical Engineering, or a related technical discipline
- Advanced degree (MS or PhD) in Computer Science, Electrical Engineering, or a related technical discipline (Preferred)
Benefits
- Competitive compensation, including stock options
- Comprehensive benefits package
- 401(k) program with matching contributions
About the Company
- Echo Neurotechnologies is an exciting new startup in the Brain-Computer Interface (BCI) space, driving innovation through advanced hardware engineering and AI solutions.
- Our mission is to deliver cutting-edge technologies that restore autonomy to people living with disabilities and improve their quality of life.
- Join a small, dedicated team of knowledgeable and motivated professionals.
- Our early-stage environment offers the opportunity to take ownership of broad decisions with significant and long-lasting impact.
- We emphasize continuous learning and growth, fostering cross-functional collaboration where your contributions are vital to our success.
Equal Opportunity
- Echo Neurotechnologies is an Equal Opportunity Employer (EOE). We celebrate diversity and are committed to creating an inclusive environment for all employees.
