
ServiceNow
Senior staff reliability engineer
Company Description It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company...
System Automation Corporation
·24 days agoSystem Automation Corporation
·24 days agoLocation
remote, United States
Salary
$120k – $140k/yr
Commitment
Full Time
Level
Middle (2-4 years)
We're looking for a mid-level Site Reliability Engineer to help build and operate the critical cloud infrastructure behind our platform in Microsoft Azure. You'll define the observability standards (SLOs, SLIs, dashboards, alerting) that tell us whether our systems are healthy, automate away manual toil, and help the team ship changes safely and often through solid CI/CD practices. You'll also share in an on-call rotation, responding to incidents and driving blameless postmortems that make the platform more resilient over time. This role suits someone with hands-on IT Operations or SRE experience, comfort working inside an agile team, and a genuine cloud-native understanding of how to design for reliability, security, and scale.
Reports to: Manager of Platform Operations and Compliance
Location: 100% Remote (must be eligible to work in the U.S.)
Compensation: Base salary, commensurate with experience. Eligible for the company's commission plan and profit sharing.
About the Company System Automation (SA) has built solutions for state and local regulatory agencies for over twenty-five years. Our SaaS platform is trusted by 500 government agencies, spans 5,000 industries, and serves over 20 million licensees. Our mission is to positively disrupt the regulatory software market by incubating new ideas, collaborating with our customers, and building technology with real-world impact. We challenge the status quo, aren't afraid to experiment, and build products that improve the states and cities we live in.
Requirements Reliability & Operations Build, operate, and scale production systems in Microsoft Azure (App Service, Networking, WAF, CosmosDB and related infrastructure) to meet availability and performance targets. Participate in an on-call rotation; triage, respond to, and resolve production incidents, and lead or contribute to blameless postmortems. Define and track SLOs/SLIs and error budgets in partnership with engineering and product teams.
Observability & Monitoring Design and maintain observability and APM tooling (metrics, logs, tracing, dashboards, alerting) so issues are caught before they impact customers. Continuously refine alert thresholds and runbooks to reduce noise and mean time to resolution.
Automation & CI/CD Reduce operational toil through automation — scripting, self-healing systems, and repeatable processes. Build and maintain CI/CD pipelines that let the development team ship safely and frequently. Provision and manage infrastructure as code (Bicep) and follow standard change control and version control practices.
Security & Compliance Ensure application infrastructure meets security and compliance requirements (e.g., SOC 2, GovRAMP) in partnership with the compliance team. Apply security best practices to infrastructure design and change management.
Collaboration & Documentation Partner with the agile development team to translate business requirements into reliable technical solutions. Participate in technical design sessions and produce clear documentation (diagrams, runbooks, architecture notes). Stay current on new Azure capabilities, industry standards, and SRE best practices, and bring recommendations back to the team. Other duties as assigned.
Knowledge, Skills, and Abilities Solid understanding of networking fundamentals, HTTP/S, and observability principles. Ability to evaluate multiple technical approaches and recommend the most effective solution for the context. Strong independent problem-solving skills balanced with effective collaboration in a team environment. Familiarity with software development lifecycle and programming/coding standards. Clear, professional communication, especially under incident pressure.
Qualifications Required 3+ years of experience in an IT Operations, DevOps, or SRE role. Hands-on technical experience with Microsoft Azure in a production environment. Experience with infrastructure as code — Terraform and/or Bicep. Proficiency in at least one scripting/programming language — Python or TypeScript. Experience working with REST and/or GraphQL APIs. Experience defining and tracking KPIs/SLOs for a web-based application. Comfortable participating in an on-call rotation.
Preferred Experience with compliance audits (SOC 2 Type 2, GovRAMP). Familiarity with security frameworks (NIST, ISO 27001). AZ-104 certification, or equivalent Azure networking experience. Experience with Node.js. Experience with low-code platforms (Power Apps, Logic Apps). Familiarity with Scrum/Agile methodology and supporting tools (Confluence, JIRA, Git, Jenkins, Bamboo, TFS). Ability to translate business requirements directly into application/site behavior changes.
What to Expect Hiring Process: Application review -- interviews and technical assessment -- conditional offer Security Screening: Due to the sensitive nature of IT systems and data access, final candidates will receive a conditional employment offer contingent upon successful completion of a drug screening and a fingerprint-based background investigation. We may use technology-assisted tools, including artificial intelligence tools, to assist recruiters and hiring managers in reviewing application materials and identifying qualifications relevant to a position. These tools support, but do not replace human decision-making. All employment decisions are made by qualified human reviewers. Applicants requiring accommodation during the application or selection process may contact us at hr@systemautomation.com. We are an Equal Opportunity Employer and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, protected veteran status, or any other characteristic protected by applicable law.

ServiceNow
Company Description It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company...
MetaRouter
About The Role As a Principal Site Reliability Engineer, you set the reliability strategy for the platform. You will define how we build, deploy, observe, and operate a distributed system that runs...

Rain
About the Company Rain is the global stablecoin payments platform for enterprises, neobanks, platforms, developers, and AI agents. Our technology allows partners to move, store, and use stablecoins...
Bright Vision Technologies
Site Reliability Engineer (SRE) - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United...

Dental Intelligence Inc.
About The Role: We're looking for a Senior Site Reliability Engineer to help us mature and scale the infrastructure behind our multi-cloud SaaS platform. Most of our footprint runs on Microsoft...

Ookla
Site Reliability Engineer The Opportunity We are looking for a highly capable engineer to join our Platform and Site Reliability engineering team. You will be responsible for building, maintaining...

Varo Bank
Varo is an entirely new kind of bank. All digital, mission-driven, FDIC insured and designed for the way our customers live their lives. A bank for all of us. ABOUT THE ROLE Varo’s SRE team is well...