Principal Distributed Systems Optimization Engineer
About the role
Job Description
We are seeking a Principal Distributed Systems Optimization Engineer to design, analyze, and optimize large-scale, high-throughput distributed systems. This role focuses on improving system performance, reliability, and cost efficiency across complex service-oriented architectures operating at significant scale.
You will work closely with backend engineers, data infrastructure teams, and platform reliability groups to identify bottlenecks, redesign critical paths, and implement robust, fault-tolerant solutions. The ideal candidate has deep expertise in distributed systems theory and hands-on experience tuning production systems under heavy load.
Key Responsibilities
Analyze system performance using metrics, tracing, and profiling tools to identify latency and throughput bottlenecks
Design and implement optimizations across microservices, data pipelines, and storage layers
Improve system reliability through fault isolation, graceful degradation, and resilience patterns
Lead architectural reviews and propose scalable design improvements
Collaborate with cross-functional teams to drive performance best practices
Build tooling and frameworks for performance benchmarking and capacity planning
Mentor engineers on distributed systems concepts and performance engineering
Required Qualifications
8+ years of experience in backend or distributed systems engineering
Strong understanding of distributed systems concepts (consensus, partitioning, replication, consistency models)
Proficiency in one or more programming languages such as Java, Scala, Go, or Python
Experience with large-scale data systems (e.g., streaming platforms, distributed databases)
Deep knowledge of performance tuning, profiling, and observability tools
Preferred Qualifications
Experience with systems like Kafka, Spark, Flink, or similar technologies
Familiarity with cloud platforms and container orchestration (e.g., Kubernetes)
Background in high-QPS, low-latency systems
Experience designing systems with strict SLAs and SLOs
What You’ll Work On
Optimizing services handling millions of requests per second
Reducing tail latency and improving system predictability
Enhancing observability and debugging capabilities in complex environments
Driving initiatives to reduce infrastructure cost while maintaining performance
Responsibilities
- Analyze system performance using metrics, tracing, and profiling tools to identify latency and throughput bottlenecks
- Design and implement optimizations across microservices, data pipelines, and storage layers
- Improve system reliability through fault isolation, graceful degradation, and resilience patterns
- Lead architectural reviews and propose scalable design improvements
- Collaborate with cross-functional teams to drive performance best practices
- Build tooling and frameworks for performance benchmarking and capacity planning
- Mentor engineers on distributed systems concepts and performance engineering
Qualifications
- 8+ years of experience in backend or distributed systems engineering
- Strong understanding of distributed systems concepts (consensus, partitioning, replication, consistency models)
- Proficiency in one or more programming languages such as Java, Scala, Go, or Python
- Experience with large-scale data systems (e.g., streaming platforms, distributed databases)
- Deep knowledge of performance tuning, profiling, and observability tools
Skills mentioned
About Partner Engineering Test Company
Partner Engineering Test Company