Research Scientist, Performance Engineering at The Biological Computing Co. | San Francisco, CA, USA | Rezi

Research Scientist, Performance Engineering at The Biological Computing Co.

Research Scientist, Performance Engineering

The Biological Computing Co. · San Francisco, CA, USA

2 weeks ago

Research Scientist, Performance Engineering

The Biological Computing Co. · San Francisco, CA, USA

15 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

TBC is building next-generation AI systems at the intersection of biological computing, generative models, and large-scale AI infrastructure. This role focuses on improving the efficiency, latency, throughput, and deployability of large models, particularly LLMs, diffusion models, video generation models, and world-model systems, to turn research systems into scalable, customer-ready products.

Responsibilities

  • Optimize inference for LLMs, diffusion models, video models, and world-model systems
  • Improve serving efficiency through techniques such as KV caching, batching, quantization, distillation, speculative decoding, and memory optimization
  • Build and optimize high-throughput inference pipelines for large models running on GPU clusters
  • Profile model performance across latency, throughput, memory usage, GPU utilization, and cost
  • Implement custom kernels or low-level optimizations using Triton, CUDA, PyTorch, or related systems
  • Improve training and fine-tuning efficiency for large generative models, including distributed training, checkpointing, parallelism, and data loading
  • Work with research teams to identify bottlenecks in model architecture, inference paths, and deployment workflows
  • Translate model performance improvements into clear customer-facing benchmarks and technical proof points
  • Evaluate trade-offs across model quality, latency, cost, memory, and deployability

Requirements

  • Strong background in machine learning systems, model optimization, or high-performance AI infrastructure
  • Hands-on experience optimizing LLMs, diffusion models, video generation models, or other large generative systems
  • Experience with inference optimization
  • Experience with KV caching / attention optimization
  • Experience with Triton or CUDA kernel development
  • Experience with quantization, pruning, distillation, or model compression
  • Experience with distributed training / fine-tuning efficiency
  • Experience with GPU profiling and performance debugging
  • Strong PyTorch experience and comfort working close to the model/runtime boundary
  • Ability to reason about trade-offs between quality, latency, throughput, memory, and cost
  • Comfortable working across research code, production systems, and benchmarking infrastructure
  • Excited to work in an ambiguous, early-stage environment where optimization work directly shapes product feasibility

Skills

  • PyTorch
  • Triton
  • CUDA

Education Level

  • PhD, MS, or equivalent industry experience in Computer Science, Machine Learning, Systems, Robotics, or related field

About the Company

  • TBC is building next-generation AI systems at the intersection of biological computing, generative models, and large-scale AI infrastructure.