1:00 PM - 2:00 PM
Senior Product Manager Interview
Sarah Jenkins

Accelerant
·10 days agoAccelerant
·10 days agoLocation
remote, United States
Commitment
Full Time
Level
Senior (5+ years)
We're building the financial data platform at Accelerant — the premium, claims, and paid data products that underpin financial processing, reserving analysis, and the monthly close — and it needs to stay fast, resilient, and observable as we scale. You'll drive the reliability and observability strategy across the platform and the enterprise systems it depends on: Velocity, MuleSoft, D365, Snowflake, Fabric, and the streaming and integration layers that move data through it. You are a key decider about what gets measured, how we define reliability, and where engineering needs to invest to keep production healthy.
We need someone who can prove a repeatable, define-to-alert observability pipeline, harden it, and scale it into systems that have never had real SLOs — and build modern, AI-assisted operational tooling that lets a small team punch far above its weight.
Own the reliability roadmap end to end. Prove a repeatable define → emit → ingest → dashboard → alert metric pipeline, set SLOs and error budgets, prioritize the work, and drive execution. You'll partner with engineering on what we monitor, how, and when — indexing on user impact over low-level infrastructure.
Take the financial data platform from functional to enterprise-grade, with a focus on availability, performance, and recoverability. Strengthen deployment paths, straight-through processing, and failover so the monthly close runs faster and cleaner as legacy hops are retired.
Extend instrumentation across the six target systems — Velocity, Red Panda, MuleSoft, Snowflake, Fabric, and AWS (with D365 ledger to follow) — proving both push (OpenTelemetry) and pull (agent) ingestion. Cover service health (latency, error rates, throughput) and business KPIs (match rate, reconciliation completeness, settlement correctness and latency).
Build the on-call, alerting, and blameless postmortem process that keeps reliability high as systems and the team grow. Route alerts Datadog → Incident.io with ServiceNow as the system of record, and set severity standards, escalation norms, and follow-up tracking that actually closes the loop.
Build the tooling that automates routine operations, self-heals common failures, and surfaces signal over noise. Establish data lineage and retention, and validate reliability at scale — 5,000+ transactions before go-live — through auto-remediation, capacity planning, and actionable dashboards.
Design and ship AI agents for incident triage, log analysis, and root-cause investigation (to name a few). Use Cursor as your build environment. Treat the agents as products solving specific problems.
Partner with the AI platform team to deploy your agents on the org's AI fabric. Make them discoverable, governed, and reusable across functions.