praxy/jobs
← Back to search

Sr. Site Reliability Engineer

7743

indexed from hirebridge · seen 3d ago · board verified 2026-07-25

https://hbapi.hirebridge.com/careercenter/v2/GetJobListings?cid=7743&language=en&startrow=1&endrow=100

Board last validated live · 2026-07-25

Location
Remote · US-Remote, US
First seen
3d ago
EngineeringSoftware SystemsOperations Execution

About this role

$110,000 - $145,000 / year + Bonus The insurance industry runs on Vertafore. We equip agencies, MGAs, and carriers with the core digital systems, specialized AI, and data-driven foundation to eliminate distribution drag across the insurance lifecycle, spanning sales, servicing, and back-office operations.  Underpinned by unmatched speed and performance power, we are the trusted backbone that’s taking the insurance industry from friction to flow with Distribution Velocity – speed, performance, and trust - to drive growth at scale.  With over 95% of the top agencies and insurers and 50% of industry compliance transactions running through Vertafore, we lead at the intersection of innovation and trust, giving insurance professionals the confidence to transform and win in the AI era.  Our reach is global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India.    We are seeking a Senior Site Reliability Engineer to own the reliability, scalability, performance, and operational integrity of critical production services. This role is accountable for the full-service lifecycle, from design and deployment readiness through production operations, incident response, and continuous improvement. Reliability is a core engineering responsibility, requiring strong software engineering skills and autonomous operation across AWS, hybrid data centers, and customer-hosted environments.  Roles and Responsibilities:  Own production services end to end. Accountable for reliability, availability, scalability, performance, and operational health.  Define and manage SLIs and SLOs, using error budgets to guide delivery decisions.  Influence of service and system design to improve fault tolerance, observability and operational sustainability.  Debug complex production issues across application code, services and infrastructure using software engineering practices.  Perform root cause analysis using logs, metrics, traces, and code-level investigation.  Build automation and self-healing mechanisms to prevent repeat failures.    Execute production changes (patching, certificate management, software releases) with safety, automation, and observability.  Design and operate production observability aligned to service health and customer impact.  Lead and participate in incident response, for high-severity events.  Collaborate with engineering, product, architecture, and operations teams.  Operate with autonomy and sound judgment in reliability decisions.    Qualifications & Requirements  8+ years of hands-on Site Reliability Engineering or reliability-focused engineering experience with end-to-end service ownership.  Proven operation at a senior engineering scope with accountability for reliability outcomes.  Strong software engineering skills in C#, .NET, Java, Python, React, or similar technologies.  Practical experience applying SRE principles (SLIs, SLOs, error budgets).  Hands-on experience with AWS, Kubernetes, CI/CD, infrastructure as code and hybrid environments.  Strong knowledge of Linux and Windows systems, application platforms and relational databases.  Bachelor’s or master’s degree in computer science or equivalent experience.  Participation in an on-call rotation; flexible hours as required. 

Closes fast — jobs here are removed within hours of going off the company's board.

Apply on 7743