Machine Learning Engineer

Pluralis Research
RemoteFull-timePosted Aug 31, 2026

About the role

About the Company

Pluralis Research conducts foundational research in Protocol Learning, an approach to training foundation models across multiple participants without requiring any single participant to hold a complete copy of the model.

The company is developing technology designed to enable community-trained and community-owned frontier models through decentralized machine learning infrastructure and sustainable economic models.

Pluralis Research is backed by leading investors and brings together experienced machine learning researchers and distributed systems engineers from major technology companies and innovative startups.

About the Role

Pluralis Research is seeking Senior and Staff-level Machine Learning Engineers with strong experience in distributed systems and large-scale machine learning training.

In this role, you will help build a novel infrastructure layer for distributed ML training designed to operate across consumer-grade internet connections and heterogeneous computing environments.

You will work at the intersection of distributed systems, machine learning infrastructure, networking, GPU optimization, and decentralized computing.

Responsibilities

Distributed Training Architecture & Optimization

Design and implement large-scale distributed machine learning training systems.

Optimize distributed training for heterogeneous hardware operating in low-bandwidth and high-latency environments.

Develop model-parallel training strategies, including data, tensor, and pipeline parallelism.

Implement custom sharding techniques to reduce communication overhead.

Optimize GPU utilization, memory efficiency, and overall compute performance.

Develop reliable checkpointing and state synchronization mechanisms.

Build recovery systems for long-running and fault-prone training workloads.

Develop monitoring and metrics systems to track training progress, model quality, and infrastructure bottlenecks.

Decentralized Networking & Resilience

Architect resilient distributed training systems capable of handling node failures and network partitions.

Design systems that support participants dynamically joining or leaving the network.

Develop peer-to-peer communication and coordination mechanisms.

Implement peer discovery, NAT traversal, dynamic routing, and connection lifecycle management.

Analyze and optimize communication patterns across distributed nodes.

Reduce network latency and bandwidth requirements in multi-participant training environments.

Build systems capable of operating reliably across non-co-located infrastructure.

Required Qualifications

5+ years of professional experience building and operating distributed systems.

Strong production experience with large-scale machine learning systems.

Hands-on experience with distributed training frameworks such as FSDP, DeepSpeed, Megatron, or comparable technologies.

Deep understanding of model parallelism, including:

Data parallelism

Tensor parallelism

Pipeline parallelism

Expert-level Python programming skills.

Production experience with Python concurrency, error handling, retry mechanisms, and clean software architecture.

Strong understanding of networking and distributed systems fundamentals.

Experience with peer-to-peer systems, gRPC, routing, NAT traversal, and distributed coordination.

Experience optimizing GPU workloads and memory management.

Strong understanding of large-scale compute performance and resource optimization.

Ability to troubleshoot complex distributed systems and performance issues.

Technical Skills

Python

Distributed Systems

Machine Learning Infrastructure

FSDP

DeepSpeed

Megatron

Model Parallelism

GPU Optimization

Memory Management

P2P Networking

gRPC

NAT Traversal

Distributed Coordination

Checkpointing

State Synchronization

Fault Tolerance

Performance Optimization

What We Offer

Equity-heavy compensation with meaningful ownership opportunities.

Competitive base salary for senior engineering positions.

Visa sponsorship available for exceptional candidates.

Remote-first working environment.

Optional access to the company's Melbourne hub.

Opportunity to work with an experienced team of ML researchers and distributed systems engineers.

Exposure to cutting-edge distributed machine learning research and infrastructure.

Opportunity to contribute to an ambitious approach to decentralized AI development.

Work Environment

This is a remote-first engineering position focused on highly technical machine learning infrastructure. Engineers will work on challenging problems involving distributed training, decentralized networking, GPU optimization, and fault-tolerant systems.

The role is particularly suited to engineers who enjoy working on complex infrastructure problems and developing new approaches to large-scale machine learning.

Equal Opportunity

Pluralis Research is committed to building a diverse and inclusive team. Employment decisions are based on qualifications, skills, experience, and organizational needs, without discrimination based on legally protected characteristics.

Responsibilities

  • Design and implement large-scale distributed machine learning training systems.
  • Optimize distributed training for heterogeneous hardware operating in low-bandwidth and high-latency environments.
  • Develop model-parallel training strategies, including data, tensor, and pipeline parallelism.
  • Implement custom sharding techniques to reduce communication overhead.
  • Optimize GPU utilization, memory efficiency, and overall compute performance.
  • Develop reliable checkpointing and state synchronization mechanisms.
  • Build recovery systems for long-running and fault-prone training workloads.
  • Develop monitoring and metrics systems to track training progress, model quality, and infrastructure bottlenecks.

Qualifications

  • 5+ years of professional experience building and operating distributed systems.
  • Strong production experience with large-scale machine learning systems.
  • Hands-on experience with distributed training frameworks such as FSDP, DeepSpeed, Megatron, or comparable technologies.
  • Deep understanding of model parallelism, including data, tensor, and pipeline parallelism.
  • Expert-level Python programming skills.
  • Production experience with Python concurrency, error handling, retry mechanisms, and clean software architecture.
  • Strong understanding of networking and distributed systems fundamentals.
  • Experience with peer-to-peer systems, gRPC, routing, NAT traversal, and distributed coordination.

Benefits

  • Equity-heavy compensation with meaningful ownership opportunities.
  • Competitive base salary for senior engineering positions.
  • Visa sponsorship available for exceptional candidates.
  • Remote-first working environment.
  • Optional access to the company's Melbourne hub.
  • Opportunity to work with an experienced team of ML researchers and distributed systems engineers.
  • Exposure to cutting-edge distributed machine learning research and infrastructure.
  • Opportunity to contribute to an ambitious approach to decentralized AI development.

Skills mentioned

PythonDistributed SystemsMachine LearningDeep LearningPerformance OptimizationgRPCCUDAConcurrencySystems EngineeringTCP/IP

About Pluralis Research

Torentify Jobs is a modern job platform helping candidates find verified job opportunities quickly and easily. From remote jobs to full-time roles, we connect talent with the right companies. Our goal is to simplify job search and help recruiters hire faster with quality candidates.

Technology2-10 employeesIndore, Madhya Pradesh