Lead DevOps (Site Reliability Engineer)
Anza Mortgage InsuranceWilmington, NCOn-site
Salary not published by the employer
Timing
- Posted by employer
- Posting date not provided by source
- Detected by our platform
- Aug 14, 2026, 2:59 PM UTC3 hours ago
- Last confirmed present
- Aug 14, 2026, 2:59 PM UTC
- Detected closed
- —
Description
About AnzaAnza is building a mortgage insurance company from first principles. Most mortgage insurers operate on decades-old technology, manual workflows, organizational structures, and operating assumptions that have evolved incrementally over time. We believe there is a fundamentally better way.Our vision is to build the industry's most technology-enabled mortgage insurer, one where software, automation, AI, analytics, and thoughtful operating design create better experiences for customers while allowing the business to scale efficiently without simply adding people.We are not interested in making legacy processes slightly better. We are interested in building the operating model that mortgage insurance should have had all along.About the roleWe are seeking a highly skilled and motivated Lead Site Reliability Engineer (SRE) to join our team. The Lead SRE will play a critical role in designing, implementing, and maintaining the reliability, scalability, and performance of our cloud-based systems hosted in AWS. You will collaborate closely with software engineers, operations teams, and other stakeholders to enhance system reliability and developer productivity through automation, monitoring, and incident responseWhat you'll doLead a team of SREs consisting of FTEs and contractorsDefine and assign the tasks to SREs, review the PRs and provide the feedbackDesign and implement scalable, reliable, and secure cloud infrastructure in AWS.Develop and maintain monitoring, alerting, and dashboarding solutions to ensure system health and uptime.Automate infrastructure provisioning and configuration management using tools like Terraform and TerragruntImplement CI/CD pipelines to streamline deployments and improve development workflows.Respond to incidents, perform root cause analysis, and implement permanent fixes to prevent recurring issues.Optimize system performance, reliability, and cost-effectiveness in collaboration with engineering teams.Drive infrastructure improvements and advocate for best practices in system design and operations.Establish and manage disaster recovery plans, ensuring system availability during unexpected events.QualificationsMinimum Education BS/BA in Computer Science or equivalent experiencePreferred Education MA in Computer ScienceMinimum SkillsArgo CD and Argo WorkflowsIaC: Terraform and TerragruntKubernetes and and related pluginsGitHub ActionsAWS (EKS, Fargate, Aurora)Security and ComplianceContainerization (Docker)Logging and Monitoring Tools Programming Scripting LanguageDB ManagementVersion Control (Git)Incident ManagementPreferred SkillsExperience with DatadogExperience with CloudflareMortgage Domain KnowledgeAdvanced Security Practices (GuardDuty, Security Hub)Disaster Recovery PlanningExperience with workflow automation tools such as Camunda Preferred Certification(s)AWS Associate, AWS DevOps ProfessionalWhat we offerWe’re committed to creating an environment where our team members can thrive both professionally and personally. We currently offer:Competitive Compensation – Including salary and performance bonuses.Comprehensive Benefits – Health, dental, vision, and mental wellness support.Retirement Savings – 401(k) with company matching.Career advancement opportunities with business growth. Inclusive Culture – A diverse, collaborative, and supportive workplace where every voice is valued.Perks & Extras – Generous PTO, team events, wellness programs, and more.