Senior Software Engineer, Site Reliability & Security
About the role
Who We Are: WellSaid Labs
WellSaid is the leading AI voiceover studio for enterprise and professional use. Using carefully sourced voice talent and our proprietary AI model, WellSaid provides ultra-realistic voices that the world’s biggest brands trust to engage listeners. We build AI responsibly and ethically.
How You’ll Contribute:
As Site Reliability Engineer at WellSaid, you'll run our production platform and be working on large-scale projects that improve the availability and performance of our core services and data pipelines.
In your day-to-day, you will:
Build and run our monitoring, tracing and alerting infrastructure
Drive platform security initiatives including preventative efforts to ensure system reliability
Lead our response and recovery efforts from active incidents which include performing root cause analysis
Run our platform and services with high availability and redundancy front and center
Respond to alerts on issues in our systems & be part of our on call response team.
Improve our deployment process with a focus on making all code changes fast, simple, and safe
Deliver a stable and scalable product platform to enable engineering teams to delivery product in a fast and efficient manner
Identify novel ways to handle load and scale resource-intensive applications
Help our engineers write reliable code while retaining our ability to ship fast
What We’re Looking For
To thrive in this role, you ideally have a strong understanding of computer science concepts including algorithms, data structures, and systems design. You have experience with cloud environments and container technology, including AWS or GCP, Kubernetes, and Docker. You ideally also have some combination of the following:
5+ years with a modern programming language (Golang, Typescript, Python, etc...)
Strong understanding of Infrastructure as Code (Terraform, Tofu, Pulumi)
Experience with GitOps/continuous delivery tooling (ArgoCD, Spacelift, Terraform Cloud)
Experience building and troubleshooting Kubernetes environments
The ability to debug and solve issues in complex production environments
An understanding of how to profile applications and databases, to identify and debug performance issues
Fluency working in a UNIX shell to analyze logs, investigate issues handle other common operational tasks
Experience with monitoring tools such as Grafana and Prometheus
To join our team you must also:
be a U.S. Citizen or Permanent Resident
pass a pre-employment background check
What We Offer
WellSaid is proud to support an inclusive work environment that emphasizes each team member’s personal and professional growth. Our team is fully distributed throughout the U.S., and we support flexible schedules - work where and when you work best. You’ll have teammates just a Slack message or video call away if you ever need help solving an exciting challenge, or even if you just have a funny story to tell.
Other perks and benefits:
Competitive salary and stock options
Full medical, dental, and vision insurance
Matching 401(k) plan
Generous vacation policy/paid time off
Parental leave
Learning & development stipend
Home office stipend
What to Expect From Us
We strongly encourage you to apply! If we feel your skills, experience, and values match, we’ll reach out about meeting with the team.
During the interview stage, you can expect:
An introductory interview with the hiring manager (50 minutes); if there’s a match we’ll schedule an interview loop with the team.
A technical screen, either as a live interview or as a take-home assessment
An Interview loop with 3-4 interviews (1 hour each) with the team members you will be potentially working with
All interviews will be remote via Google Meets; we are happy to make accommodations you might need to feel comfortable and set up for success in our process.
WellSaid is honored to be an equal opportunity workplace. We realize that by bringing together teams rich in diverse thoughts and experiences, our people, company, and customers are free to flourish. We are committed to providing equal employment opportunities regardless of race, color, national origin, religion, creed, genetic information, sex (including pregnancy, sexual orientation or gender identity), age, marital status, disability, military or veteran status; or any other protected classifications or characteristics under applicable local laws.
Responsibilities
- Build and run monitoring, tracing and alerting infrastructure
- Drive platform security initiatives
- Lead response and recovery efforts from active incidents
- Run platform and services with high availability
- Improve deployment process for code changes
- Deliver a stable and scalable product platform
- Identify ways to handle load and scale applications
- Help engineers write reliable code
Qualifications
- Strong understanding of computer science concepts
- Experience with cloud environments and container technology
- 5+ years with a modern programming language
- Strong understanding of Infrastructure as Code
- Experience with GitOps/continuous delivery tooling
- Experience building and troubleshooting Kubernetes environments
- Ability to debug and solve issues in production environments
- Fluency in UNIX shell
Benefits
- Competitive salary and stock options
- Full medical, dental, and vision insurance
- Matching 401(k) plan
- Generous vacation policy/paid time off
- Parental leave
- Learning & development stipend
- Home office stipend
Skills mentioned
About WellSaid Labs
WellSaid is the AI voiceover platform built for professional audio production. WellSaid's voice models are developed through consent-based partnerships with pro voice actors, ensuring every voiceover comes with guaranteed commercial rights and no IP liability for our customers. WellSaid is trusted by the world's leading companies, including Microsoft, T-Mobile, Progressive, Humana, and ServiceNow, for training, marketing, product, and creative content.