Production Support Lead

Optum
United StatesFull-timePosted Aug 29, 2026

About the role

About The Company

Optum is a leading global organization committed to transforming healthcare through innovative solutions and technology-driven care delivery. Our mission is to help millions of people live healthier lives by connecting them with the right care, pharmacy benefits, data, and resources. We foster a culture of inclusion, collaboration, and continuous growth, offering our employees comprehensive benefits, career development opportunities, and a supportive environment. Join us to make a meaningful impact on communities worldwide as we advance health optimization on a global scale. At Optum, your work directly contributes to improving health outcomes and shaping the future of healthcare.

About The Role

We are seeking a highly skilled and experienced Production Support Lead to oversee and manage critical customer-facing applications and services. In this role, you will be responsible for leading production support activities, coordinating application releases, managing incident response efforts, and ensuring operational excellence across cloud and on-premise environments. You will collaborate closely with engineering, DevOps, security, and external vendors to maintain system stability, security, and performance. Your expertise will drive continuous improvement initiatives focused on automation, reliability, and operational efficiency. This position offers an exciting opportunity to work in a dynamic environment, leveraging cutting-edge technologies such as Kubernetes, Kafka, and cloud platforms like GCP, AWS, or Azure.

Qualifications

Bachelor's degree in Computer Science, Information Technology, or related field, or 6+ years of software engineering experience

8+ years of experience in object-oriented programming with Java

5+ years of experience as a lead in production support for critical customer-facing applications and services

3+ years of experience leading or coordinating P1 and P2 incident response efforts, including war room facilitation and stakeholder communication

3+ years of experience working with public cloud platforms such as AWS, Azure, or GCP

2+ years of experience with container technologies like Docker and Kubernetes

1+ year of experience with automation and scripting tools such as Python and Bash

Responsibilities

Lead production support activities for critical customer-facing applications and services, ensuring high availability and performance

Coordinate application release activities, including deployment planning, implementation, and post-release validation

Manage change requests throughout their lifecycle, ensuring proper approval and implementation

Lead or coordinate P1 and P2 incident response efforts, including war room facilitation and stakeholder communication

Drive Root Cause Analysis (RCA) activities and ensure corrective actions are implemented effectively

Present incident reviews and service improvement recommendations to leadership teams

Plan, coordinate, and execute Disaster Recovery (DR) exercises across staging and production environments

Partner with Engineering and DevOps teams to identify and remediate security vulnerabilities and compliance issues

Coordinate monthly security scanning and remediation activities with QA and external partners

Support operational governance activities such as service reviews, reliability initiatives, and Product Lifecycle Management (PLM) reviews

Maintain Configuration Management Database (CMDB) and Configuration Item (CI) documentation

Manage operational forecasts, capacity planning, and health service reporting to ensure optimal performance

Collaborate with third-party vendors and external partners to resolve service requests and production issues

Review upcoming infrastructure and application changes to ensure operational readiness and risk mitigation

Support cloud platform operations, including Kubernetes, Kafka, GCP, and HCP Console technologies

Coordinate onboarding of Kafka topics, event streams, and related data integrations

Drive continuous improvement initiatives focused on automation, reliability, and operational efficiency

Benefits

Comprehensive health insurance plans

Incentive and recognition programs

Equity stock purchase options

401(k) retirement plan contributions

Career development and training opportunities

Supportive and inclusive work environment

Equal Opportunity

UnitedHealth Group is an Equal Employment Opportunity employer. We consider all qualified applicants without regard to race, national origin, religion, age, gender, sexual orientation, gender identity, disability, veteran status, or any other characteristic protected by law. We are committed to creating a diverse and inclusive workplace and adhere to all applicable laws and regulations. Additionally, we are a drug-free workplace, and candidates are required to pass a drug test prior to employment. Pursuant to the San Francisco Fair Chance Ordinance, we will consider qualified applicants with arrest and conviction records.

Responsibilities

  • Lead production support activities for critical customer-facing applications and services, ensuring high availability and performance
  • Coordinate application release activities, including deployment planning, implementation, and post-release validation
  • Manage change requests throughout their lifecycle, ensuring proper approval and implementation
  • Lead or coordinate P1 and P2 incident response efforts, including war room facilitation and stakeholder communication
  • Drive Root Cause Analysis (RCA) activities and ensure corrective actions are implemented effectively
  • Present incident reviews and service improvement recommendations to leadership teams
  • Plan, coordinate, and execute Disaster Recovery (DR) exercises across staging and production environments
  • Partner with Engineering and DevOps teams to identify and remediate security vulnerabilities and compliance issues

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, or related field, or 6+ years of software engineering experience
  • 8+ years of experience in object-oriented programming with Java
  • 5+ years of experience as a lead in production support for critical customer-facing applications and services
  • 3+ years of experience leading or coordinating P1 and P2 incident response efforts
  • 3+ years of experience working with public cloud platforms such as AWS, Azure, or GCP
  • 2+ years of experience with container technologies like Docker and Kubernetes
  • 1+ year of experience with automation and scripting tools such as Python and Bash

Benefits

  • Comprehensive health insurance plans
  • Incentive and recognition programs
  • Equity stock purchase options
  • 401(k) retirement plan contributions
  • Career development and training opportunities
  • Supportive and inclusive work environment

Skills mentioned

JavaGoogle CloudKubernetesDockerApache KafkaPythonAutomationIncident ResponsePerformance OptimizationDevOps

About Optum

Sundayy makes job discovery feel less chaotic and more intentional. Instead of jumping between platforms, repeating the same steps, and getting lost in the process, everything is brought into one place so you can focus on what actually matters. No clutter, no confusion, just a clearer way to move forward.

Technology11-50 employeesNew York