About the Role
This role removes the bottleneck between frontier research and production reality, enabling researchers to ship faster, demos to launch quicker, and customers to experience models at their best.
Responsibilities
- Turn research checkpoints into production-ready inference services
- Design and maintain high-performance APIs serving millions of requests
- Optimize inference latency and throughput across GPU infrastructure
- Build scalable serving architectures that handle unpredictable traffic
- Improve reliability, monitoring, and observability across model-serving systems
- Prototype and ship demos that showcase new capabilities in days, not weeks
- Collaborate closely with researchers to move from idea to live endpoint rapidly
Requirements
- Built and operated systems at meaningful scale
- Understand the difference between a research prototype and a production system
- Comfortable navigating ambiguity, making tradeoffs, and improving systems under real-world constraints
- Strong judgment around performance, reliability, and cost tradeoffs
- Experience scaling APIs or ML systems under load
- Comfort working in fast-moving, research-adjacent environments
- Ownership from system design through debugging and deployment
Skills
- Python
- FastAPI
- async systems
- GPU infrastructure
- CUDA
- inference optimization
- Docker
- Kubernetes
- Redis
- Postgres
- distributed task queues
- Cloud platforms (AWS, GCP, or Azure)
- Observability stacks (metrics, logging, tracing)
- Building and operating ML inference services in production
- Designing scalable API architectures with async processing
- Optimizing GPU workloads (batching, quantization, compilation, CUDA)
- Managing distributed systems and task queues under variable load
- Implementing monitoring and observability for production ML systems
- Debugging performance bottlenecks across model, infrastructure, and network layers
- Real-time or low-latency inference systems
- TensorRT
- reduced precision
- layer fusion
- model compilation techniques
- Frontend demo tooling (Streamlit, Gradio, React)
- CI/CD and automated testing for ML systems
- Security best practices for API and model serving
Location
- Freiburg
- San Francisco
- Remote
Work Type
- Hybrid
- Remote
Experience Level
- Meaningful scale
- Production systems
- Scaling APIs or ML systems under load
- Fast-moving, research-adjacent environments
- System design through debugging and deployment
- ML inference services in production
- Scalable API architectures with async processing
- GPU workloads optimization
- Distributed systems and task queues management
- Monitoring and observability implementation
- Performance bottleneck debugging
Salary/Compensations
- $180,000–$300,000 USD
About the Company
- We're the team behind Latent Diffusion, Stable Diffusion, and FLUX—foundational technologies that changed how the world creates images and video.
- We’re creating the generative models that power how people make images and video—tools used by millions of creators, developers, and businesses worldwide.
- Our FLUX models are among the most advanced in the world, and we're just getting started.
- Headquartered in Freiburg, Germany with a growing presence in San Francisco, we're scaling fast while staying true to what makes us different: research excellence, open science, and building technology that expands human creativity.
- We're based in Europe and value depth over noise, collaboration over hero culture, and honest technical conversations over hype.
- Our models have been downloaded hundreds of millions of times, but we're still a ~50-person team learning what's possible at the edge of generative AI.
