We are sharing a specialised part-time consulting opportunity for experienced Site Reliability Engineering and incident management professionals with strong expertise in production reliability, incident response, operational resilience, root-cause analysis, and service performance. This role focuses on reviewing professional documents, spreadsheets, and presentation materials related to SRE, reliability engineering, and incident management.
Key Responsibilities
Site Reliability Engineering
- Evaluate work products involving production reliability, service availability, and operational resilience
- Assess whether recommendations reflect sound SRE principles and realistic production environments
- Review approaches to reliability, scalability, performance, and service health
- Identify technically weak assumptions, operational gaps, and impractical recommendations
- Apply professional judgement grounded in real-world Site Reliability Engineering experience
Incident Management & Response
- Review incident-response plans, escalation workflows, and operational procedures
- Assess incident classification, prioritisation, ownership, and coordination
- Evaluate whether proposed response actions are appropriate for severity and business impact
- Identify gaps in communication, escalation, containment, or recovery
- Review incident-management approaches for speed, clarity, and operational effectiveness
Root Cause & Post-Incident Analysis
- Evaluate root-cause analyses and post-incident reviews
- Assess whether conclusions are supported by technical and operational evidence
- Identify shallow causal analysis, unsupported assumptions, or missed contributing factors
- Review corrective and preventive actions for practicality and effectiveness
- Evaluate whether lessons learned translate into meaningful reliability improvements
Reliability Metrics & Service Health
- Review analyses involving availability, latency, reliability, and service-performance metrics
- Assess use of SLIs, SLOs, error budgets, and related reliability measures where relevant
- Evaluate whether metrics appropriately reflect service health and user impact
- Identify inconsistencies between underlying data and reported conclusions
- Review whether reliability targets and operational recommendations are realistic
Monitoring & Operational Readiness
- Evaluate monitoring, alerting, observability, and operational-readiness approaches
- Assess whether alerts are actionable and aligned with meaningful service conditions
- Review escalation paths, runbooks, and response procedures
- Identify gaps in detection, diagnosis, or operational preparedness
- Evaluate whether proposed controls support reliable production operations
Resilience & Failure Management
- Review scenarios involving outages, degraded performance, capacity constraints, and system failures
- Assess mitigation, recovery, and resilience strategies
- Evaluate trade-offs between reliability, performance, complexity, and operational cost
- Identify single points of failure or poorly addressed dependencies
- Review recommendations for reducing recurrence and improving service resilience
Documents & Presentation Review
- Evaluate incident reports, reliability analyses, operational documents, spreadsheets, and slide decks for accuracy and completeness
- Review stakeholder and executive presentations for clarity and decision usefulness
- Identify factual, technical, analytical, aesthetic, and formatting issues
- Assess whether charts, tables, and visuals accurately represent underlying operational information
- Ensure conclusions and recommendations are clearly supported by evidence
Structured Evaluation & Feedback
- Assess assigned outputs against domain-specific quality criteria
- Identify technical, operational, analytical, and presentation weaknesses
- Distinguish substantive reliability issues from minor editorial concerns
- Provide clear, structured written feedback explaining identified strengths and weaknesses
- Apply evaluation standards consistently across different SRE and incident-management work products
Ideal Profile
- 5+ years of relevant professional experience in Site Reliability Engineering, incident management, reliability engineering, DevOps, production engineering, systems engineering, or a closely related field
- Strong practical understanding of SRE and production reliability
- Hands-on experience with incident response, escalation, post-incident review, and root-cause analysis
- Experience with monitoring, observability, service health, and operational readiness
- Strong understanding of availability, performance, resilience, and service-level objectives
- Ability to assess technical recommendations for operational feasibility and reliability impact
- Highly proficient with Microsoft Office and Google Workspace
- Advanced proficiency with PowerPoint / Google Slides
- Strong spreadsheet and analytical skills
- Native or professional fluency in English
- Excellent written communication and ability to provide precise, structured feedback
- Strong attention to technical, operational, analytical, and presentation detail
- Master's degree or higher from a recognised institution is advantageous
Engagement Details
- Part-time independent contractor engagement
- Fully remote
- Flexible scheduling based on project requirements
- Compensation: $70–$110/hour
- Work includes evaluation of incident-management plans, reliability analyses, operational documentation, spreadsheets, and presentation materials
- Projects may be extended, shortened, or concluded based on project needs and performance
- Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
- H1-B and STEM OPT support is unavailable for this engagement
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.