Research Engineer - Post-Training at Pluralis Research | AU | Rezi

Research Engineer - Post-Training at Pluralis Research

Research Engineer - Post-Training

Pluralis Research · AU

3 weeks ago

Research Engineer - Post-Training

Pluralis Research · AU

23 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Research Engineer - Post-Training role.

Rezi rewrites your resume against Pluralis Research's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Research Engineer - Post-Training posting at Pluralis Research — free, in seconds.

About the Role

Pluralis Research is pioneering Protocol Learning, enabling decentralized training and serving of large models on consumer devices. This role focuses on making RL post-training work within a unique, decentralized system, adapting algorithms and building the end-to-end stack for decentralized model releases.

Responsibilities

  • Build the RL training loop end-to-end, including rollout ingestion, reward computation, policy updates, and weight distribution.
  • Adapt standard RL algorithms for asynchronous, high-latency, and partially trusted generation environments.
  • Develop evaluation metrics to demonstrate model improvement.
  • Manage the release of the first decentralized post-trained models from development to public artifact.

Requirements

  • Hands-on experience with RL post-training on large language models (RLHF, RLVR, or reasoning-focused RL).
  • Experience with the systems layer of RL post-training, including rollout generation, asynchronous training loops, and weight synchronization.
  • Production-quality Python and PyTorch skills, including concurrency, failure handling, and performance profiling.
  • Ability to defend research work, demonstrated through publications or detailed unpublished work.
  • Alignment with the mission of Protocol Learning for collective, trustless, and sovereign AI.
  • Professional-level English proficiency (written and spoken).

Skills

  • RL post-training
  • Large language models
  • RLHF
  • RLVR
  • Reasoning-focused RL
  • Systems engineering
  • Rollout generation
  • Asynchronous training loops
  • Weight synchronization
  • Python
  • PyTorch
  • Concurrency
  • Failure handling
  • Performance profiling
  • Research
  • Algorithm adaptation
  • Decentralized systems
  • Federated learning
  • Reward modeling
  • P2P networking
  • NAT traversal

Location

  • Remote
  • Australia
  • North America
  • US

Work Type

  • Remote-First
  • Full-time

Experience Level

  • Senior

Salary/Compensations

  • High base salary

Benefits

  • Significant ownership (equity)
  • Flexible work environment
  • Optional full visa sponsorship
  • Relocation support

About the Company

  • Pluralis Research is developing Protocol Learning, a novel approach to decentralized training and serving of large AI models on consumer devices.
  • The company has achieved significant advances, including the Agora run which pretrained an 8B model from scratch on distributed consumer GPUs.
  • Pluralis is backed by Union Square Ventures and other tier-1 investors.
  • The team consists of world-class, deeply technical ML researchers.
  • Pluralis is ideologically driven, believing Protocol Learning offers a better path for AI development.