This job is no longer active. It is kept for history rather than deleted.
Lead DevOps (Site Reliability Engineer)
Anza Mortgage InsuranceMcLean, VAOn-site
Salary not published by the employer
Timing
- Posted by employer
- Posting date not provided by source
- Detected by our platform
- Aug 14, 2026, 2:58 PM UTC16 hours ago
- Last confirmed present
- Aug 14, 2026, 2:58 PM UTC
- Detected closed
- Aug 13, 2026, 5:20 AM UTC
Description
Anza MI is a fintech startup using technology and analytics to drive growth & innovation within the US mortgage market.Anza Mortgage Insurance Corporation is empowering homeownership through credit risk protection.Our MissionOur mission is to empower homeownership by mitigating credit risk through mortgage insurance. We strive to leverage technology and operational excellence to deliver exceptional service to our lender and servicer partners, offer reliable and secure capital support for our insureds, generate leading returns for our shareholders, and foster a fulfilling, rewarding environment for our team members.Technology ForwardWe differentiate ourselves by designing, architecting, and developing our operating systems in-house. This unique approach creates a platform tailored to the needs of today's housing market. It positions us to achieve unparalleled processing efficiency and adaptability to changing industry standards and customer demands. Fundamentally, we are a process optimization company that harnesses technology to revolutionize the mortgage insurance experience and fulfill our mission.Customer CentricWe dedicate ourselves to provide competitive pricing tailored to risk profiles. Our commitment to our partners goes beyond pricing and product offerings. We prioritize a customer-centric approach focused on meeting customer needs and enhancing their experience. This approach is underpinned by transparency throughout the process and prompt issue resolution.Risk Management ExcellenceRisk Management is our foundation. Our risk management framework uses advanced technology, innovative strategies, and modern modeling techniques to manage risks ranging from credit risk to information security risk. Our customers and partners can rely on us to handle their information with the utmost care and responsibility. Our dedication to risk management is designed to reassure our partners about our stability and reliability.About the roleWe are seeking a highly skilled and motivated Lead Site Reliability Engineer (SRE) to join our team. The Lead SRE will play a critical role in designing, implementing, and maintaining the reliability, scalability, and performance of our cloud-based systems hosted in AWS. You will collaborate closely with software engineers, operations teams, and other stakeholders to enhance system reliability and developer productivity through automation, monitoring, and incident responseWhat you'll doLead a team of SREs consisting of FTEs and contractorsDefine and assign the tasks to SREs, review the PRs and provide the feedbackDesign and implement scalable, reliable, and secure cloud infrastructure in AWS.Develop and maintain monitoring, alerting, and dashboarding solutions to ensure system health and uptime.Automate infrastructure provisioning and configuration management using tools like Terraform and TerragruntImplement CI/CD pipelines to streamline deployments and improve development workflows.Respond to incidents, perform root cause analysis, and implement permanent fixes to prevent recurring issues.Optimize system performance, reliability, and cost-effectiveness in collaboration with engineering teams.Drive infrastructure improvements and advocate for best practices in system design and operations.Establish and manage disaster recovery plans, ensuring system availability during unexpected events.QualificationsMinimum Education BS/BA in Computer Science or equivalent experiencePreferred Education MA in Computer ScienceMinimum SkillsArgo CD and Argo WorkflowsIaC: Terraform and TerragruntKubernetes and and related pluginsGitHub ActionsAWS (EKS, Fargate, Aurora)Security and ComplianceContainerization (Docker)Logging and Monitoring Tools Programming Scripting LanguageDB ManagementVersion Control (Git)Incident ManagementPreferred SkillsExperience with DatadogExperience with CloudflareMortgage Domain KnowledgeAdvanced Security Practices (GuardDuty, Security Hub)Disaster Recovery PlanningExperience with workflow automation tools such as Camunda Preferred Certification(s)AWS Associate, AWS DevOps Professional