Research Engineer, Privacy and Anonymization at Clera | CA, US | Rezi

Research Engineer, Privacy and Anonymization at Clera

Research Engineer, Privacy and Anonymization

Clera · CA, US

6 days ago

Research Engineer, Privacy and Anonymization

Clera · CA, US

7 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Research Engineer, Privacy and Anonymization role.

Rezi rewrites your resume against Clera's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Research Engineer, Privacy and Anonymization posting at Clera — free, in seconds.

About the Role

This Research Engineer role focuses on building privacy and anonymization systems to make sensitive data safe and useful for AI training. You will own the full pipeline for protecting privacy without destroying data value, operating at the intersection of applied research and production engineering.

Responsibilities

  • Build systems to detect PII, quasi-identifiers, credentials, and other sensitive information, designing transformations based on data type and downstream use case.
  • Develop and benchmark detection approaches combining rules, statistical models, classifiers, and LLM-based methods.
  • Build production pipelines that anonymize raw data before downstream processing, training, evaluation, or synthetic data generation.
  • Create evaluation frameworks to measure privacy risk and retained data utility, including recall-weighted metrics, leakage tests, and adversarial re-identification attempts.
  • Design systems robust to new data sources, schema drift, unusual formats, and sensitive information in unexpected fields.
  • Collaborate with engineering, research, operations, and customers to translate privacy requirements into practical technical policies and safeguards.

Requirements

  • 2+ years building production data or ML systems in Python, with strong proficiency.
  • Hands-on experience with information extraction, named-entity recognition, classification, or related methods for detecting sensitive or rare content.
  • Strong experimental instincts and ability to compare approaches across recall, precision, latency, cost, and downstream data utility.
  • Solid understanding of redaction, masking, pseudonymization, anonymization, and synthetic data generation, and their appropriate use.
  • Experience designing systems robust to schema drift, unusual data formats, and edge cases.
  • End-to-end experience building data processing pipelines without a fully prescribed roadmap.
  • Familiarity with privacy-enhancing technologies (e.g., differential privacy, k-anonymity, secure aggregation, format-preserving encryption) is a plus.
  • Experience with low-latency or high-throughput ML inference and data-processing systems is a plus.
  • Prior work with sensitive data in healthcare, finance, security, or related domains is a plus.

Skills

  • Python
  • Information Extraction
  • Named-Entity Recognition
  • Classification
  • Rules-based Models
  • Statistical Models
  • LLM-based Methods
  • Data Anonymization
  • Privacy Risk Assessment
  • Data Utility Measurement
  • Schema Drift Handling
  • Edge Case Handling
  • Data Processing Pipelines
  • Differential Privacy
  • K-anonymity
  • Secure Aggregation
  • Format-Preserving Encryption
  • Low-latency ML Inference
  • High-throughput ML Inference

Location

  • San Francisco, California

Work Type

  • On-site

Experience Level

  • 2+ years