Job Description
Staff Operational Support Engineer L2
37639950 Hourly pay: $70/hr
Worksite: Leading audio, video, and voice technologies company (Atlanta, GA 30308 - Onsite)
W2 Employment, Group Medical, Dental, Vision, Life, Retirement Savings Program, PSL 40 hours/week, 12 Month Assignment
A leading video, audio, and voice technologies company is seeking a Staff Operational Support Engineer L2 to provide operational support for 24/7 live video streaming, advertising, player, and real-time delivery platforms.
Responsibilities:
- Own escalated customer issues from Level 1 Support through resolution, troubleshooting complex production incidents affecting live streams, VOD playback, ad insertion, DRM, and real-time WebRTC services.
- Operate directly in production environments to perform configuration changes, CDN adjustments, mitigations, and emergency changes when required, while providing clear and timely customer-facing communication and leading or contributing to live incident bridges with customers, internal teams, and partners.
- Work with Infrastructure as Code as the primary mechanism for safe, auditable, and repeatable production changes, using Terraform, Helm, Kubernetes manifests, GitOps workflows, CI/CD and deployment pipelines.
- Validate and execute infrastructure and configuration changes through codified workflows and collaborate with Engineering and DevOps to improve deployment reliability and operational safety.
- Improve operational efficiency and incident response, including AI-assisted incident triage and classification, automated runbook execution, AI-based incident pattern detection, intelligent alert correlation and noise reduction, automated or improved incident communications, accelerated troubleshooting workflows, and identification of recurring or systemic issues.
- Drive adoption of automation-first and AI-augmented operational practices.
- Support pre-event operational readiness for critical customer events through runbook checks, monitoring coverage validation, risk identification and mitigation planning, and rehearsed incident-response strategies.
- Respond to critical alerts within defined SLAs for stream health, player errors, and delivery infrastructure.
- Perform and contribute to root cause analyses, document findings and corrective and preventive actions, identify recurring issues and partner with Engineering and Product teams to eliminate them.
- Improve runbooks, operational playbooks, and knowledge bases across player, advertising, live-streaming, and real-time products.
- Support production deployments and defect resolution, provide feedback on observability, tooling gaps, and operational risks, and serve as the operational voice during post-incident reviews.
Qualifications:
- 5+ years of relevant experience in operational, support, or similar customer-facing roles.
- Experience supporting production video streaming platforms, OTT services, and live systems.
- Troubleshooting skills across distributed systems, including APIs, microservices, and cloud infrastructure.
- Familiarity with HLS, DASH, CMAF, WebRTC, DRM, and CDN architectures.
- Experience using monitoring, alerting, and logging tools such as Grafana, Kibana/ELK, Prometheus, and Loki to diagnose live incidents.
- Ability to correlate backend streaming metrics, player telemetry, and CDN signals to diagnose live customer issues end-to-end.
- Comfort performing controlled changes in production environments.
- Working knowledge of incident management and on-call operations.
Shift: 24/7 On-Call Rotation: Includes nights, weekends, and holidays as part of a global support model, ensuring effective handoffs between shifts and regions.