Production Support Lead
About the role
About The Company
Optum is a leading global organization committed to transforming healthcare through innovative solutions and technology-driven care delivery. Our mission is to help millions of people live healthier lives by connecting them with the right care, pharmacy benefits, data, and resources. We foster a culture of inclusion, collaboration, and continuous growth, offering our employees comprehensive benefits, career development opportunities, and a supportive environment. Join us to make a meaningful impact on communities worldwide as we advance health optimization on a global scale. At Optum, your work directly contributes to improving health outcomes and shaping the future of healthcare.
About The Role
We are seeking a highly skilled and experienced Production Support Lead to oversee and manage critical customer-facing applications and services. In this role, you will be responsible for leading production support activities, coordinating application releases, managing incident response efforts, and ensuring operational excellence across cloud and on-premise environments. You will collaborate closely with engineering, DevOps, security, and external vendors to maintain system stability, security, and performance. Your expertise will drive continuous improvement initiatives focused on automation, reliability, and operational efficiency. This position offers an exciting opportunity to work in a dynamic environment, leveraging cutting-edge technologies such as Kubernetes, Kafka, and cloud platforms like GCP, AWS, or Azure.
Qualifications
Bachelor's degree in Computer Science, Information Technology, or related field, or 6+ years of software engineering experience
8+ years of experience in object-oriented programming with Java
5+ years of experience as a lead in production support for critical customer-facing applications and services
3+ years of experience leading or coordinating P1 and P2 incident response efforts, including war room facilitation and stakeholder communication
3+ years of experience working with public cloud platforms such as AWS, Azure, or GCP
2+ years of experience with container technologies like Docker and Kubernetes
1+ year of experience with automation and scripting tools such as Python and Bash
Responsibilities
Lead production support activities for critical customer-facing applications and services, ensuring high availability and performance
Coordinate application release activities, including deployment planning, implementation, and post-release validation
Manage change requests throughout their lifecycle, ensuring proper approval and implementation
Lead or coordinate P1 and P2 incident response efforts, including war room facilitation and stakeholder communication
Drive Root Cause Analysis (RCA) activities and ensure corrective actions are implemented effectively
Present incident reviews and service improvement recommendations to leadership teams
Plan, coordinate, and execute Disaster Recovery (DR) exercises across staging and production environments
Partner with Engineering and DevOps teams to identify and remediate security vulnerabilities and compliance issues
Coordinate monthly security scanning and remediation activities with QA and external partners
Support operational governance activities such as service reviews, reliability initiatives, and Product Lifecycle Management (PLM) reviews
Maintain Configuration Management Database (CMDB) and Configuration Item (CI) documentation
Manage operational forecasts, capacity planning, and health service reporting to ensure optimal performance
Collaborate with third-party vendors and external partners to resolve service requests and production issues
Review upcoming infrastructure and application changes to ensure operational readiness and risk mitigation
Support cloud platform operations, including Kubernetes, Kafka, GCP, and HCP Console technologies
Coordinate onboarding of Kafka topics, event streams, and related data integrations
Drive continuous improvement initiatives focused on automation, reliability, and operational efficiency
Benefits
Comprehensive health insurance plans
Incentive and recognition programs
Equity stock purchase options
401(k) retirement plan contributions
Career development and training opportunities
Supportive and inclusive work environment
Equal Opportunity
UnitedHealth Group is an Equal Employment Opportunity employer. We consider all qualified applicants without regard to race, national origin, religion, age, gender, sexual orientation, gender identity, disability, veteran status, or any other characteristic protected by law. We are committed to creating a diverse and inclusive workplace and adhere to all applicable laws and regulations. Additionally, we are a drug-free workplace, and candidates are required to pass a drug test prior to employment. Pursuant to the San Francisco Fair Chance Ordinance, we will consider qualified applicants with arrest and conviction records.
Responsibilities
- Lead production support activities for critical customer-facing applications and services, ensuring high availability and performance
- Coordinate application release activities, including deployment planning, implementation, and post-release validation
- Manage change requests throughout their lifecycle, ensuring proper approval and implementation
- Lead or coordinate P1 and P2 incident response efforts, including war room facilitation and stakeholder communication
- Drive Root Cause Analysis (RCA) activities and ensure corrective actions are implemented effectively
- Present incident reviews and service improvement recommendations to leadership teams
- Plan, coordinate, and execute Disaster Recovery (DR) exercises across staging and production environments
- Partner with Engineering and DevOps teams to identify and remediate security vulnerabilities and compliance issues
Qualifications
- Bachelor's degree in Computer Science, Information Technology, or related field, or 6+ years of software engineering experience
- 8+ years of experience in object-oriented programming with Java
- 5+ years of experience as a lead in production support for critical customer-facing applications and services
- 3+ years of experience leading or coordinating P1 and P2 incident response efforts
- 3+ years of experience working with public cloud platforms such as AWS, Azure, or GCP
- 2+ years of experience with container technologies like Docker and Kubernetes
- 1+ year of experience with automation and scripting tools such as Python and Bash
Benefits
- Comprehensive health insurance plans
- Incentive and recognition programs
- Equity stock purchase options
- 401(k) retirement plan contributions
- Career development and training opportunities
- Supportive and inclusive work environment
Skills mentioned
About Optum
Sundayy makes job discovery feel less chaotic and more intentional. Instead of jumping between platforms, repeating the same steps, and getting lost in the process, everything is brought into one place so you can focus on what actually matters. No clutter, no confusion, just a clearer way to move forward.