About the Role
We are seeking an individual to make implicit system design decisions explicit and ensure they are maintained. This involves codifying golden paths, building shared primitives and APIs for product engineers, owning the system's architecture, and establishing reliability standards before incidents occur. Platform at Cortea bridges software engineering, DevOps, and SRE, with engineers as its customers. This role emphasizes application code over YAML, with infrastructure experience valued for its contribution to effective system design.
Responsibilities
- Make the easiest way for product engineers to do something also the most secure, reliable and scalable way by default.
- Build shared primitives, libraries, and APIs that hide complexity and carry quality and observability standards.
- Own the architecture and core stack, defining standardization, reuse, and future system direction.
- Implement infrastructure as code for cloud resources, dashboards, and alerts.
- Provision, run, and tune Kubernetes clusters and cloud footprint.
- Keep CI fast as the deploy rate grows.
- Define SLIs for critical workloads and build SLOs on top of them.
- Ensure alerting is high-signal and trusted.
- Drive down AI and infrastructure spend.
- Keep internal foundations secure by default, including IAM, dependency management, and secrets.
- Own the authentication and authorization stack, including ReBAC models for humans and agents.
- Implement SOC 2 and ISO 27001 controls efficiently.
- Ensure documentation and agent guidelines are in place for product engineers and agents.
- Shorten the path from design to implementation with standardized automations and development environments.
- Manage and distribute evolving agent configurations with auditability and tenant isolation.
- Develop harnesses for AI agents, abstracting over rapidly changing use cases.
- Build centralized progress tracking for parallel agent executions.
- Implement usage-based tenant billing with quotas and rate limiting.
- Define SLIs for the background job system and fix exposed bottlenecks.
- Develop a document pipeline for various file formats and large uploads under strict reliability requirements.
Requirements
- Designed, built, and operated distributed systems end-to-end, with a deep understanding of system interactions.
- Strong backend software engineering and DevOps/SRE skills, viewing them as integrated disciplines.
- Experience writing design docs, RFCs, ADRs, and postmortems.
- A preference for simple, pragmatic solutions, distinguishing essential from accidental complexity, and knowing when to cut corners.
- Ability to approach complex problems from alternative perspectives.
- Proven experience scaling relational transactional databases like Postgres under load.
- Experience running production Kubernetes.
- Experience building observability (SLIs, SLOs, distributed tracing) rather than inheriting it.
- Experience designing a system from scratch, running it in production, and documenting design decisions.
Skills
- Backend software engineering
- DevOps
- SRE
- Design docs
- RFCs
- ADRs
- Postmortems
- Distributed systems
- Kubernetes
- Observability
- SLIs
- SLOs
- Distributed tracing
- Postgres
- Relational transactional databases
- Infrastructure as code
- IAM
- Dependency management
- Secrets management
- Authentication
- Authorization
- ReBAC models
- SOC 2
- ISO 27001
- AI dev tooling
- Agent configurations
- Workflow orchestrators
- Temporal
- Background job processing
- Queue processing
- LLM-based products
- Agentic systems
- Audit
- Finance
- Compliance
- Azure
- GCP
Location
- Berlin
Work Type
- On-site
Experience Level
- Senior
Salary/Compensations
- Competitive salary
Benefits
- Meaningful equity
- Work face-to-face with experienced founders
- Learn directly from customer insights
- Real influence on strategy
- High autonomy
- Shape product direction
- Fast, ambitious, and fun team environment
- Rapid experimentation
- Share in the upside
About the Company
- We ship fast and are growing our system.
- We run AI agents over customer documents in a high-stakes domain.
- Platform at Cortea sits between software engineering, DevOps, and SRE.
- We use AI heavily.
- We are building intelligent systems for a $200bn industry.
- Based in Berlin, with AI at the core.
