About the Role
As our Data Engineer, you’ll architect and maintain pipelines that transform high-frequency time-series, lab, and historian data into a scalable Lakehouse architecture, usable for both deep learning models and real-time LLMs. You’ll be working across AWS and Databricks/PySpark, ensuring data is contextualised, synchronised, and optimised for AI workloads. This role involves solving problems at the intersection of control systems, industrial data engineering, and AI enablement.
Responsibilities
- Ingest data from OPC UA servers, process historians, IoT sensors, LIMS systems, alarms/events, and P&IDs.
- Map signals to their physical processes for interpretability in AI pipelines.
- Build pipelines for real-time streaming and batch ingestion into the Lakehouse.
- Manage data synchronisation between historian archives, unstructured files, and AWS storage.
- Orchestrate Databricks Lakeflow/Connectors for integrating data into Lakebase/Lakehouse.
- Handle secure, high-throughput data transfers between historian archives and environments.
- Detect and manage schema changes, signal drift, and inconsistencies over time.
- Implement lineage and audit trails across Spark/Databricks and AWS pipelines.
- Build and maintain dual pipelines for training (historical data prep) and inference (real-time pipelines).
- Support heterogeneous AI workloads including time-series forecasting and retrieval-augmented LLMs.
- Tune PostgreSQL and Spark for high-throughput time-series workloads.
- Optimise pipelines for fast analytical queries and efficient model training.
- Deploy and manage data pipelines in AWS EKS with persistent EBS-backed storage.
Requirements
- Deep expertise in PostgreSQL (partitioning, indexing, query optimisation, storage design).
- Strong proficiency in Python for data processing, scripting, and pipeline orchestration.
- Hands-on experience with AWS (EKS, S3, EBS, IAM, KMS, CloudWatch, etc.) for secure and scalable data pipelines.
- Proven ability to work with Databricks and PySpark for large-scale distributed data processing.
- Familiarity with time-series industrial data (control systems, DCS/SCADA logs, process historians).
- Experience in unstructured data sync and management within hybrid cloud/on-prem environments.
- Experience working as a data engineer in oil and gas or energy environments is a bonus.
- Knowledge of streaming frameworks (Kafka, Flink, Spark Streaming) or MLOps stacks for data versioning and lineage is a bonus.
Skills
- PostgreSQL
- Python
- AWS
- EKS
- S3
- EBS
- IAM
- KMS
- CloudWatch
- Databricks
- PySpark
- Time-series data
- Control systems
- DCS/SCADA logs
- Process historians
- Unstructured data management
- Hybrid cloud
- On-premise environments
- Kafka
- Flink
- Spark Streaming
- MLOps
- Data versioning
- Data lineage
Location
- UK
Work Type
- Full-time
Experience Level
- Deep expertise
- Strong proficiency
- Hands-on experience
- Proven ability
- Familiarity
- Experience
About the Company
- Applied Computing was founded in 2024 to build Orbital, a physics-informed foundation model for energy operations.
- We’re live across oil and gas, refineries, and petrochemicals, working towards our mission: sustainable abundance for a growing planet.
- The hydrocarbon industry keeps the world running, but its complexity has left operators tied to legacy systems, making critical decisions on less than 10% of available data.
- We built Orbital to change that. It’s a foundation model built specifically for energy that lets companies use AI at scale, harnessing all of their operational data and optimising in real time for any metric.
- Decisions get faster, operations get safer, and carbon intensity falls.
- We’ve raised over $32 million, including one of the largest seed rounds for an AI company in the UK.
- We’re just getting started.
