About the Role
As a Lead Data Engineer, you will build and lead the data infrastructure powering an intelligent AI assistant that automates operational workloads for property managers, letting agents and build-to-rent teams. You will architect and scale the data systems behind the AI products, covering real-time data pipelines, analytics infrastructure, vector databases and machine learning data workflows. As the first senior data hire, you will play a key role in defining the data architecture, technology stack, engineering standards and ways of working, and help shape the data team.
Responsibilities
- Architect and build scalable data pipelines and infrastructure to support AI and product systems.
- Design and maintain data ingestion, transformation and storage architectures for operational and AI workloads.
- Develop and manage batch and real-time data pipelines.
- Build and optimise systems for vector search, retrieval and machine learning data pipelines.
- Ensure data reliability, security and governance across the platform.
- Collaborate with AI and backend engineering teams to support training, inference and product features.
- Implement monitoring, observability and data quality frameworks.
- Optimise the performance of large-scale datasets and query systems.
- Contribute to technical architecture decisions and long-term data strategy.
- Act as the founding data hire, defining culture, standards and the hiring bar for the data function as it scales.
- Partner directly with founders and product leadership to translate data capabilities into product decisions.
Requirements
- 7+ years of professional experience, with the majority in dedicated data engineering roles.
- Strong experience designing and building data pipelines and distributed data systems.
- Experience working with relational databases, with PostgreSQL preferred.
- Experience working with NoSQL databases.
- Experience with vector databases used in modern AI systems.
- Strong programming experience in Python.
- Demonstrated ability to make and justify architectural decisions.
- Experience building scalable backend systems.
- Experience designing data models and storage architectures.
- Strong understanding of data processing performance and optimisation.
- Experience with Apache Spark, Apache Airflow, Kafka, and Elasticsearch or OpenSearch is highly desirable.
- Experience with PostgreSQL, MongoDB, and vector databases such as Qdrant, Milvus or pgvector is highly desirable.
- Experience with Python data-processing libraries such as Pandas or Polars is highly desirable.
- Experience working on AI or machine learning platforms is a plus.
- Familiarity with stream processing and event-driven architectures is a plus.
- Experience with cloud infrastructure such as GCP, AWS or Azure is a plus.
- Experience working in high-growth startups or early-stage companies is a plus.
Skills
- Python
- PostgreSQL
- MySQL
- NoSQL databases
- Vector databases
- Apache Spark
- Apache Airflow
- Kafka
- Elasticsearch
- OpenSearch
- MongoDB
- Qdrant
- Milvus
- pgvector
- Pandas
- Polars
- GCP
- AWS
- Azure
Experience Level
- 7+ years of professional experience
About the Company
- At Smart Working, we believe your job should not only look right on paper but also feel right every day.
- This isn’t just another remote opportunity - it’s about finding where you truly belong, no matter where you are.
- From day one, you’re welcomed into a genuine community that values your growth and well-being.
- Our mission is simple: to break down geographic barriers and connect skilled professionals with outstanding global teams and products for full-time, long-term roles.
- We help you discover meaningful work with teams that invest in your success, where you’re empowered to grow personally and professionally.
- Join one of the highest-rated workplaces on Glassdoor and experience what it means to thrive in a truly remote-first world.
