About the Role
This is a broad, full-stack data engineering role where you will design and run production data pipelines, model data for cross-team usability, and build the warehouse and tooling to transform raw data into actionable insights. As the second data engineer, you will have significant ownership from day one, a direct impact on company operations, and numerous challenging problems to solve.
Responsibilities
- Design, build, and maintain reliable production data pipelines, including ingestion, transformation, and orchestration.
- Model and structure data in the warehouse to be clean, documented, and useful for engineering, product, and research teams.
- Own data infrastructure, including schema design, migrations, performance, and cost, alongside the current data engineer.
- Provide data support to GTM engineers by delivering curated, enriched company and account data and the necessary modeling layer.
- Integrate and route diverse data sources, from internal product events to third-party enrichment feeds, into the warehouse.
- Build data-extraction jobs, including web scraping, when sources are not otherwise available.
- Improve data quality, observability, and documentation to enable the team to move quickly and reliably.
- Diagnose and fix pipeline issues proactively.
- Contribute to data and infrastructure tasks as needed to enhance robustness.
Requirements
- Solid experience as a data engineer building and running production pipelines and data warehouses.
- Strong SQL and Python skills.
- Comfort designing data models for others to build upon.
- Hands-on experience with a modern data stack, including relational databases, columnar/analytics stores, orchestration, and transformation tooling.
- Experience with PostgreSQL (or similar) and efficient data structuring and querying.
- A production mindset focused on reliability, migrations, performance, and cost.
- Ability to work independently and own problems end-to-end in a fast-paced, early-stage environment.
- Clear communication skills.
- Ability to work in English.
Skills
- SQL
- Python
- Data Modeling
- Data Warehousing
- Production Data Pipelines
- PostgreSQL
- Relational Databases
- Columnar/Analytics Stores
- Orchestration Tooling
- Transformation Tooling
- Streaming/Eventing (Kafka)
- Dataset Versioning
- Lightweight Data Pipelines
- dbt
- Fivetran
- Airbyte
- dlt
- AWS (Athena, Glue, S3)
- GTM/CRM Data Workflows
- Crunchbase
- Clay
- Web Scraping
- Systems Language (Rust)
Location
- Paris (hybrid)
Work Type
- Hybrid
- Full-time
Experience Level
- Early-stage
- Senior
Benefits
- Paid time off in line with local regulations
- Relocation package
- Best medical insurance in France
- All necessary hardware, tools, and services
- Covered subscriptions for AI agents and IDEs
- Team off-sites twice a year
About the Company
- White Circle is an AI Safety company building the safety, reliability, and optimization layer for AI systems.
- At the core of our platform are policies – simple natural-language rules that define what an AI model should and shouldn’t do.
- We automatically test, enforce, and continuously improve these policies at scale.
- We’ve raised $11M from top funds, founders, and senior leaders at OpenAI, Anthropic, HuggingFace, Mistral, DeepMind, Datadog, Sentry, and others.
- We process over one hundred million API calls every month.
- We fine-tune and train our own LLMs so they run faster and cheaper than any open or proprietary model.
- We’re a small, highly focused team looking for individuals who want to work deeply on hard problems, see their work ship to production quickly, and influence how AI safety is actually built.
