2

24-Mag

·13 days ago

Remote | site reliability engineer (sre & incident management) — $70–$110/hour

Apply now

Location

remote, New York, NY, United States

Salary

$141k – $222k/yr

Commitment

Part Time

Level

Senior (5+ years)

Required skills

IaaS

Job Description

We are sharing a specialised part-time consulting opportunity for experienced Site Reliability Engineering and incident management professionals with strong expertise in production reliability, incident response, operational resilience, root-cause analysis, and service performance. This role focuses on reviewing professional documents, spreadsheets, and presentation materials related to SRE, reliability engineering, and incident management.

Key Responsibilities

Site Reliability Engineering

  • Evaluate work products involving production reliability, service availability, and operational resilience
  • Assess whether recommendations reflect sound SRE principles and realistic production environments
  • Review approaches to reliability, scalability, performance, and service health
  • Identify technically weak assumptions, operational gaps, and impractical recommendations
  • Apply professional judgement grounded in real-world Site Reliability Engineering experience

Incident Management & Response

  • Review incident-response plans, escalation workflows, and operational procedures
  • Assess incident classification, prioritisation, ownership, and coordination
  • Evaluate whether proposed response actions are appropriate for severity and business impact
  • Identify gaps in communication, escalation, containment, or recovery
  • Review incident-management approaches for speed, clarity, and operational effectiveness

Root Cause & Post-Incident Analysis

  • Evaluate root-cause analyses and post-incident reviews
  • Assess whether conclusions are supported by technical and operational evidence
  • Identify shallow causal analysis, unsupported assumptions, or missed contributing factors
  • Review corrective and preventive actions for practicality and effectiveness
  • Evaluate whether lessons learned translate into meaningful reliability improvements

Reliability Metrics & Service Health

  • Review analyses involving availability, latency, reliability, and service-performance metrics
  • Assess use of SLIs, SLOs, error budgets, and related reliability measures where relevant
  • Evaluate whether metrics appropriately reflect service health and user impact
  • Identify inconsistencies between underlying data and reported conclusions
  • Review whether reliability targets and operational recommendations are realistic

Monitoring & Operational Readiness

  • Evaluate monitoring, alerting, observability, and operational-readiness approaches
  • Assess whether alerts are actionable and aligned with meaningful service conditions
  • Review escalation paths, runbooks, and response procedures
  • Identify gaps in detection, diagnosis, or operational preparedness
  • Evaluate whether proposed controls support reliable production operations

Resilience & Failure Management

  • Review scenarios involving outages, degraded performance, capacity constraints, and system failures
  • Assess mitigation, recovery, and resilience strategies
  • Evaluate trade-offs between reliability, performance, complexity, and operational cost
  • Identify single points of failure or poorly addressed dependencies
  • Review recommendations for reducing recurrence and improving service resilience

Documents & Presentation Review

  • Evaluate incident reports, reliability analyses, operational documents, spreadsheets, and slide decks for accuracy and completeness
  • Review stakeholder and executive presentations for clarity and decision usefulness
  • Identify factual, technical, analytical, aesthetic, and formatting issues
  • Assess whether charts, tables, and visuals accurately represent underlying operational information
  • Ensure conclusions and recommendations are clearly supported by evidence

Structured Evaluation & Feedback

  • Assess assigned outputs against domain-specific quality criteria
  • Identify technical, operational, analytical, and presentation weaknesses
  • Distinguish substantive reliability issues from minor editorial concerns
  • Provide clear, structured written feedback explaining identified strengths and weaknesses
  • Apply evaluation standards consistently across different SRE and incident-management work products

Ideal Profile

  • 5+ years of relevant professional experience in Site Reliability Engineering, incident management, reliability engineering, DevOps, production engineering, systems engineering, or a closely related field
  • Strong practical understanding of SRE and production reliability
  • Hands-on experience with incident response, escalation, post-incident review, and root-cause analysis
  • Experience with monitoring, observability, service health, and operational readiness
  • Strong understanding of availability, performance, resilience, and service-level objectives
  • Ability to assess technical recommendations for operational feasibility and reliability impact
  • Highly proficient with Microsoft Office and Google Workspace
  • Advanced proficiency with PowerPoint / Google Slides
  • Strong spreadsheet and analytical skills
  • Native or professional fluency in English
  • Excellent written communication and ability to provide precise, structured feedback
  • Strong attention to technical, operational, analytical, and presentation detail
  • Master's degree or higher from a recognised institution is advantageous

Engagement Details

  • Part-time independent contractor engagement
  • Fully remote
  • Flexible scheduling based on project requirements
  • Compensation: $70–$110/hour
  • Work includes evaluation of incident-management plans, reliability analyses, operational documentation, spreadsheets, and presentation materials
  • Projects may be extended, shortened, or concluded based on project needs and performance
  • Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
  • H1-B and STEM OPT support is unavailable for this engagement

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.

Ready to join the team?

Apply now

Similar Jobs: