Data Scientist at Sunset | NY, US | Rezi

Data Scientist at Sunset

Data Scientist

Sunset · NY, US

1 months ago

Data Scientist

Sunset · NY, US

a month ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Data Scientist role.

Rezi rewrites your resume against Sunset's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Data Scientist posting at Sunset — free, in seconds.

About the Role

Sunset turns sensitive internal enterprise data into de-identified datasets without destroying the structure and meaning that make the data valuable. As Sunset's first Data Scientist focused on evaluation, you will establish how we know whether that data is actually getting better. You will build the datasets, experiments, quality measures, and feedback loops that expose hidden failures, accelerate model and pipeline improvement, and give the team confidence in what it delivers. This is a hands-on, zero-to-one role at the intersection of data science, AI, and a real production system. You will write Python and SQL, construct evaluation corpora, study failure patterns, design comparisons, calibrate human and model-based judgments, and turn the result into a clear decision. The questions are scientifically difficult, but the output must be practical enough to change what the team builds and ships. You will work closely with Machine Learning, Product Engineering, Data Engineering, Security, Quality, domain experts, and the team making delivery decisions. Machine Learning Engineers own changing model behavior. You own the credibility of the evidence used to decide whether a model, pipeline, or delivery change actually made the data safer or more useful.

Responsibilities

  • Define what high-quality and safe-to-deliver data mean across de-identification, structure preservation, semantic coherence, and customer utility
  • Design representative samples and build golden, adversarial, replay, and production-like corpora with explicit provenance, labeling policy, agreement, adjudication, and versioning
  • Turn ambiguous concepts such as “useful,” “clean,” or “safe” into measurable claims with known uncertainty and clear decision consequences
  • Evaluate detectors, models, prompts, judges, thresholds, review workflows, and pipeline changes using comparisons that can support a real decision
  • Break aggregate results into the modalities, providers, entity classes, customer contexts, languages, formats, and risk tiers that reveal consequential failures
  • Connect local measures to escaped sensitive information, avoidable over-redaction, preserved data utility, review burden, rework, and delivery acceptance
  • Build reproducible analysis, evaluation pipelines, and high-fidelity environments using Python, SQL, synthetic data, historical replay, seeded failures, and programmatic verifiers
  • Establish holdout and evaluation practices that keep the evidence trustworthy while model and product teams iterate quickly
  • Use modern AI tools deeply for analysis, corpus development, coding, review, and hypothesis generation while independently verifying their output

Requirements

  • You have at least three years of professional experience in applied science, data science, machine learning, quantitative research, or a closely related role
  • You have designed evaluations or experiments that changed a product, model, release, or operational decision
  • You understand sampling, uncertainty, precision, recall, F1, calibration, agreement, class imbalance, distribution shift, and imperfect labels
  • You can investigate messy, multi-stage data systems and determine where an apparent gain or loss actually came from
  • You are comfortable writing Python and SQL and building reproducible technical artifacts rather than handing requirements to an engineering team
  • You can protect the independence of an evaluation while collaborating closely with the people whose work it evaluates
  • You have startup experience and enjoy broad ownership, changing context, and building the measurement foundation while decisions are already moving quickly
  • You use AI tools fluently but do not confuse an articulate model output with valid evidence
  • You communicate uncertainty and difficult findings directly, without hiding behind false precision
  • Experience evaluating NER, entity resolution, information extraction, document understanding, multimodal, retrieval, or LLM systems
  • Experience with privacy, de-identification, data quality, model risk, safety, or other high-trust decision systems
  • Experience designing human-review, adjudication, weak-supervision, or active-learning systems
  • Experience building adversarial corpora, replay systems, simulation environments, programmatic verifiers, or model-judge evaluations
  • Experience connecting offline measures to escaped defects, customer outcomes, review effort, or preserved data utility
  • Experience measuring quality across multi-stage batch or data pipelines

Skills

  • Python
  • SQL
  • Data Science
  • Machine Learning
  • Quantitative Research
  • Sampling
  • Uncertainty
  • Precision
  • Recall
  • F1
  • Calibration
  • Agreement
  • Class Imbalance
  • Distribution Shift
  • Imperfect Labels
  • AI Tools

Experience Level

  • 3+ years of professional experience

About the Company

  • Sunset was founded to help founders, initially supporting startups through shutting down, and has expanded into unlocking a new revenue stream for all types of businesses.
  • The company's insight in 2025 was that the data generated daily through collaboration, communication, and building is valuable training data for AI models.
  • Sunset partners directly with frontier AI labs, providing a primary source of real, proprietary data.
  • The company has scaled from $0 to a multi-eight-figure run rate in months.
  • Sunset has raised funding from investors including Floodgate, Afore, Ludlow, and Hustle Fund.