Nagarro
Nagarro·9 days ago

Associate staff engineer(sre)

Apply now

Location

field, Shanghai, China

Commitment

Full Time

Level

Senior (5+ years)

Required skills

AWSDevOpsKubernetesTerraformCI/CDLinuxPythonObservabilityInfrastructure as CodeProduction SupportIncident ManagementDockerPrometheusGrafanaCloud ComputingAutomation

Job Description

Job Description

Must have Skills: DevOps - AWS (Strong)

Job Description:

Senior Site Reliability Engineer (SRE) Role
We are seeking an experienced Senior Site Reliability Engineer (SRE) to support highly available, business-critical platforms within a global hospitality environment. The role is responsible for ensuring the reliability, availability, performance, scalability, security, and operational readiness of production systems supporting hotel operations, reservation services, digital channels, loyalty platforms, payment services, and other guest-facing applications.

The successful candidate will combine strong expertise in Cloud, DevOps, Kubernetes, Infrastructure as Code, Observability, Automation, and Production Support, and will work closely with Development, Infrastructure, Security, Architecture, and Service Management teams in a global 24x7 operating model.

Key Responsibilities

  • Ensure the availability, reliability, performance, and scalability of critical production services.
  • Define and monitor SLIs, SLOs, SLAs, error budgets, availability, latency, and service health metrics.
  • Act as a senior technical escalation point for major production incidents and P1/P2 issues.
  • Lead troubleshooting, Root Cause Analysis, post-incident reviews, and corrective actions.
  • Reduce operational toil through automation, self-healing, and engineering improvements.
  • Operate and troubleshoot workloads across AWS, Azure, and/or GCP environments.
  • Support enterprise Kubernetes and container platforms, including EKS, AKS, GKE, or OpenShift.
  • Develop and maintain Infrastructure as Code using Terraform, CloudFormation, Bicep, or equivalent technologies.
  • Build and maintain CI/CD pipelines using tools such as Jenkins, GitHub Actions, GitLab CI, Azure DevOps, Argo CD, or Harness.
  • Implement and maintain observability solutions covering metrics, logs, tracing, alerts, dashboards, and APM. Support tools such as Prometheus, Grafana, Dynatrace, Datadog, Splunk, ELK/OpenSearch, or New Relic.
  • Perform performance analysis, capacity planning, load testing, and scalability optimization.
  • Support Disaster Recovery, Business Continuity, failover testing, and RTO/RPO validation.
  • Participate in production release, change, patching, vulnerability remediation, and operational readiness activities.
  • Maintain technical runbooks, operational procedures, monitoring standards, and knowledge documentation.
  • Participate in a global 24x7 on-call / production support model where required.

Hospitality Technology Scope

The role may support platforms including:

  • Property Management Systems (PMS)
  • Central Reservation Systems (CRS)
  • Hotel booking and reservation platforms
  • Loyalty and membership systems
  • Guest-facing websites and mobile applications
  • Payment platforms
  • Property connectivity and integration services
  • API and middleware platforms
  • Revenue management and hotel operational systems

The engineer will help ensure the reliability of critical guest journeys such as hotel search, booking, reservation modification, payment, check-in/check-out, loyalty transactions, and property system integration.

Required Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline.
  • 7+ years of experience in Cloud, DevOps, Infrastructure, Production Engineering, or IT Operations.
  • 3+ years of hands-on experience in an SRE, DevOps, Cloud Operations, or Production Engineering role.
  • Strong hands-on experience with at least one major cloud platform: AWS, Azure, or GCP. Strong Kubernetes, Docker, and container troubleshooting skills.
  • Hands-on experience with Terraform or other Infrastructure as Code technologies.
  • Strong experience with CI/CD, deployment automation, and release management.
  • Strong knowledge of Linux and production troubleshooting.
  • Experience with enterprise monitoring, logging, tracing, and observability platforms.
  • Experience managing major incidents, RCA, problem management, and production stability.
  • Good knowledge of networking concepts including DNS, HTTP/HTTPS, load balancing, firewall, routing, VPN, and CDN.
  • Experience supporting microservices, distributed systems, APIs, databases, and messaging platforms.
  • Scripting or programming experience with Python, Bash, PowerShell, Go, Java, or similar languages.
  • Good understanding of ITIL-based Incident, Problem, Change, and Knowledge Management processes.
  • Strong written and verbal English communication skills.

Preferred Qualifications

  • Experience in hospitality, travel, airline, e-commerce, financial services, or other 24x7 high-availability industries.
  • Experience supporting high-volume transactional or reservation platforms.
  • Experience with PCI DSS, GDPR, ISO 27001, DevSecOps, and vulnerability management.
  • Experience with Kafka, Redis, API gateways, service mesh, or event-driven architectures.
  • Experience with cloud cost optimization / FinOps.
  • Experience with resilience testing, chaos engineering, or automated recovery.
  • Relevant certifications such as AWS, Azure, GCP, CKA, Terraform Associate, or ITIL.

Ready to join the team?

Apply now

Over 200,000 job seekers trust us

Diego Martínez
The AI tools here are outstanding. I found a job within a week of signing up, and the personalized job suggestions made the whole process effortless. I'm beyond impressed!
Feb 16, 2025
Brian Raine
I was skeptical, but this platform exceeded my expectations. The AI-driven resume builder helped me craft a professional resume. The auto-apply feature is a game-changer — saves so many tedious hours. I found a remote position eventually. Highly recommend Global Work AI.
Feb 16, 2025
Tomi Horvat
This site is amazing! The AI tools saved me so much time. I secured two contracts.
Feb 16, 2025

Frequently Asked
Questions

Everything you need to know about
finding your next role
  • How does Global Work AI apply to jobs without spamming?

    Global Work AI only applies to roles that closely match your experience, seniority, and career goals. Each application is personalized and submitted intentionally-one role at a time-so it looks human, relevant, and never spammy.
  • Is Global Work AI safe for my professional reputation?

    Yes. Protecting your professional reputation is built into the product. We avoid mass applications, irrelevant roles, and generic language. Every application is tailored to the role and designed to reflect how a strong candidate would apply manually.
  • Do I need a “perfect” resume to use Global Work AI?

    No. You don’t need a perfect resume to get started. Global Work AI uses your existing resume to build a Candidate Profile, then adapts your resume for each role-highlighting what matters most.
  • What will my resume and cover letter look like before they’re submitted?

    You stay fully in control. Before anything is submitted, you can review how your resume and cover letter are tailored for the role-clear, focused, and aligned with the job description.
  • Will recruiters be able to tell that I used AI to apply?

    No. Applications are designed to look natural and intentional. Because each one is role-specific and personalized, they don’t resemble the generic or automated submissions that recruiters are accustomed to filtering out.
  • What if I’m not satisfied with Global Work AI?

    We want you to feel confident trying Global Work AI. If you’re not satisfied, you can rely on our Money-Back Guarantee. Contact our support team at contact@globalwork.ai , and we’ll assist you.
Associate staff engineer(sre) at Nagarro in Shanghai, CHN