What we operate
The retainer covers the full operational surface of a modern cloud platform. Scope is locked at SOW; the bullets below are the menu most engagements draw from.
- 24×7 monitoring and alerting. Synthetic checks, SLI / SLO tracking, alert routing into your on-call tooling (PagerDuty / OpsGenie / Opsgenie). We tune the noise out so on-call actually sleeps.
- IAM and access governance. Least-privilege reviews, role drift detection, JIT elevation workflows, key and credential rotation, audit-log monitoring, SSO and SAML / OIDC posture.
- Patching and vulnerability management. OS patch cadence, container image refresh, dependency vulnerability tracking, CVE triage, emergency patch path.
- Backup, DR, and business continuity. Backup policy enforcement, restore-test automation, cross-region failover validation, RTO / RPO reporting, tabletop facilitation.
- Security posture management. CIS / CSPM scanning, drift remediation, secret rotation, network egress review, public-bucket / public-IP monitoring, GuardDuty / Security Command Center / Defender integration.
- Compliance management. SOC 2, ISO 27001, HIPAA, DPDP, RBI guidelines — control evidence collection, audit-prep coordination, gap remediation. We do not run the audit; we make sure you pass it.
- Incident management. Triage, comms, runbook execution, post-incident reviews. We take a slot in your on-call rotation for incidents in our scope.
- Capacity and cost management. Monthly trend, anomaly detection, commitment tracker maintenance, rightsizing recommendations, FinOps guardrails.
- Change management. IaC stewardship, module updates, drift detection, terraform-state hygiene, approval workflows for high-blast-radius changes.
- Monthly platform review. A four-page memo: cost trend, reliability metrics, security posture, capacity headroom, top three risks worth your time. Delivered to leadership; built for the CTO to forward.
Service tiers (illustrative — actual SOW is scope-quoted)
We do not run a fixed price-list for the retainer because the right shape depends on which of the above are in scope, how many cloud accounts, what SLA you need, and which compliance frameworks apply. Three patterns we engage against:
- Essential. Monitoring, IAM, patching, monthly review. Business-hours response. For platforms in steady state where you want a second pair of eyes and a defensible compliance posture.
- Standard. Adds backup / DR, security posture management, incident participation, capacity management. After-hours SEV-1 response. For platforms with regulated workloads or 99.9% SLO commitments.
- Premium. Adds compliance management, dedicated change management, quarterly DR tabletops, and a higher response SLA. For platforms with 99.95%+ SLO commitments or active audit cycles.
How we differ from a traditional MSP
- No ticket queue. Communication happens in your existing channels (Slack, Teams, the on-call tool). We do not have a portal you log into.
- Senior engineer, not L1. The platform owner on your account is the same person across months. No offshored bench rotation.
- Per-quarter retainer, not per-ticket. Incentives align — we are paid to keep tickets from happening, not to bill against them.
- Genuinely embedded. We attend your platform reviews, your DR exercises, your security audits. We are part of the team, billed monthly.
Service-level commitments
- Response time — agreed in the SOW per tier (typically 30 minutes for SEV-1, 4 hours SEV-2, 1 business day SEV-3 on Standard).
- Availability of platform owner — minimum 8 hours / day overlap with your team time zone.
- Monthly review delivery — by the 5th business day of the following month.
- Compliance evidence collection — within 5 business days of audit request.
How we engage
Onboarding (~2 weeks) is included in the first quarter — credentials, IaC handoff, runbook walkthrough, monitoring access, IAM review. Monthly retainer is invoiced in advance. Quarterly auto-renew; cancel any quarter, no penalty.
Capacity
We deliberately keep our retainer portfolio small. We accept new retainers only when capacity is available — usually one or two per quarter. If we are at capacity, we will say so up front and recommend an alternative.
AIOps Managed — signal hygiene as a retainer sub-scope
For clients already running Datadog Watchdog, Dynatrace Davis, PagerDuty AIOps, or cloud-native anomaly detection, AIOps signal hygiene is available as a named scope within the Platform Management retainer — or as a standalone AIOps Managed engagement at $4–8k/month. Scope: monthly correlation-rule review, ownership-map updates, alert-volume trend reporting, and vendor liaison. The broader retainer absorbs this automatically at Standard tier and above; smaller teams can engage it standalone.
If you have not deployed an AIOps tool yet or your current deployment is not performing as expected, start with the AIOps Readiness Scorecard ($500, 5 business days) before deciding whether to expand into the managed retainer.
FAQ
Application incidents?
Out of scope. We cover the cloud and data platform layer; application code stays with your team. The runbook makes the boundary explicit.
Can you also do execution work during the retainer?
Small changes are in scope. Larger changes (a new pipeline, a new warehouse, a re-architecture) are scoped as a separate Execution SOW.
What clouds do you cover?
AWS, GCP, Azure, OCI. Data platform layer: BigQuery, Snowflake, Databricks, Redshift. Other stacks — let's talk, but probably no.