Skip to content
All services
Cloud Platform Management

Managed services without the ticket queue.

A managed-services retainer with the discipline of an MSP and the judgement of a staff engineer. 24×7 monitoring, IAM governance, patching, backup and DR, security posture, compliance, incident management, capacity and SLO tracking — operated by a named senior engineer who knows your platform.

What we operate

The retainer covers the full operational surface of a modern cloud platform. Scope is locked at SOW; the bullets below are the menu most engagements draw from.

  • 24×7 monitoring and alerting. Synthetic checks, SLI / SLO tracking, alert routing into your on-call tooling (PagerDuty / OpsGenie / Opsgenie). We tune the noise out so on-call actually sleeps.
  • IAM and access governance. Least-privilege reviews, role drift detection, JIT elevation workflows, key and credential rotation, audit-log monitoring, SSO and SAML / OIDC posture.
  • Patching and vulnerability management. OS patch cadence, container image refresh, dependency vulnerability tracking, CVE triage, emergency patch path.
  • Backup, DR, and business continuity. Backup policy enforcement, restore-test automation, cross-region failover validation, RTO / RPO reporting, tabletop facilitation.
  • Security posture management. CIS / CSPM scanning, drift remediation, secret rotation, network egress review, public-bucket / public-IP monitoring, GuardDuty / Security Command Center / Defender integration.
  • Compliance management. SOC 2, ISO 27001, HIPAA, DPDP, RBI guidelines — control evidence collection, audit-prep coordination, gap remediation. We do not run the audit; we make sure you pass it.
  • Incident management. Triage, comms, runbook execution, post-incident reviews. We take a slot in your on-call rotation for incidents in our scope.
  • Capacity and cost management. Monthly trend, anomaly detection, commitment tracker maintenance, rightsizing recommendations, FinOps guardrails.
  • Change management. IaC stewardship, module updates, drift detection, terraform-state hygiene, approval workflows for high-blast-radius changes.
  • Monthly platform review. A four-page memo: cost trend, reliability metrics, security posture, capacity headroom, top three risks worth your time. Delivered to leadership; built for the CTO to forward.

Service tiers (illustrative — actual SOW is scope-quoted)

We do not run a fixed price-list for the retainer because the right shape depends on which of the above are in scope, how many cloud accounts, what SLA you need, and which compliance frameworks apply. Three patterns we engage against:

  • Essential. Monitoring, IAM, patching, monthly review. Business-hours response. For platforms in steady state where you want a second pair of eyes and a defensible compliance posture.
  • Standard. Adds backup / DR, security posture management, incident participation, capacity management. After-hours SEV-1 response. For platforms with regulated workloads or 99.9% SLO commitments.
  • Premium. Adds compliance management, dedicated change management, quarterly DR tabletops, and a higher response SLA. For platforms with 99.95%+ SLO commitments or active audit cycles.

How we differ from a traditional MSP

  • No ticket queue. Communication happens in your existing channels (Slack, Teams, the on-call tool). We do not have a portal you log into.
  • Senior engineer, not L1. The platform owner on your account is the same person across months. No offshored bench rotation.
  • Per-quarter retainer, not per-ticket. Incentives align — we are paid to keep tickets from happening, not to bill against them.
  • Genuinely embedded. We attend your platform reviews, your DR exercises, your security audits. We are part of the team, billed monthly.

Service-level commitments

  • Response time — agreed in the SOW per tier (typically 30 minutes for SEV-1, 4 hours SEV-2, 1 business day SEV-3 on Standard).
  • Availability of platform owner — minimum 8 hours / day overlap with your team time zone.
  • Monthly review delivery — by the 5th business day of the following month.
  • Compliance evidence collection — within 5 business days of audit request.

How we engage

Onboarding (~2 weeks) is included in the first quarter — credentials, IaC handoff, runbook walkthrough, monitoring access, IAM review. Monthly retainer is invoiced in advance. Quarterly auto-renew; cancel any quarter, no penalty.

Capacity

We deliberately keep our retainer portfolio small. We accept new retainers only when capacity is available — usually one or two per quarter. If we are at capacity, we will say so up front and recommend an alternative.

AIOps Managed — signal hygiene as a retainer sub-scope

For clients already running Datadog Watchdog, Dynatrace Davis, PagerDuty AIOps, or cloud-native anomaly detection, AIOps signal hygiene is available as a named scope within the Platform Management retainer — or as a standalone AIOps Managed engagement at $4–8k/month. Scope: monthly correlation-rule review, ownership-map updates, alert-volume trend reporting, and vendor liaison. The broader retainer absorbs this automatically at Standard tier and above; smaller teams can engage it standalone.

If you have not deployed an AIOps tool yet or your current deployment is not performing as expected, start with the AIOps Readiness Scorecard ($500, 5 business days) before deciding whether to expand into the managed retainer.

FAQ

Application incidents?

Out of scope. We cover the cloud and data platform layer; application code stays with your team. The runbook makes the boundary explicit.

Can you also do execution work during the retainer?

Small changes are in scope. Larger changes (a new pipeline, a new warehouse, a re-architecture) are scoped as a separate Execution SOW.

What clouds do you cover?

AWS, GCP, Azure, OCI. Data platform layer: BigQuery, Snowflake, Databricks, Redshift. Other stacks — let's talk, but probably no.

Who this is for

Does any of this sound familiar?

If it does, the next section explains how the engagement model is structured to address each one.

Your MSP sends L1 tickets for every alert. You've spent three months teaching them what's noise and what's real.

A named senior engineer — who knows your platform and its alert patterns — owns the account. There is no offshored L1 tier between you and the person who understands the blast radius of a change.

You're heading into a SOC 2 audit. Nobody has been collecting the control evidence your auditor needs. It's month nine.

Compliance management is a retainer-level deliverable: control evidence collection, audit-prep coordination, gap remediation. We make sure you pass — without you needing to be in the room for every evidence request.

Your IAM posture drifted when three engineers left and two contractors onboarded. Nobody ran a privilege audit since Q1.

IAM and access governance is a core retainer scope: least-privilege reviews, role drift detection, JIT elevation workflows, key rotation, SSO posture. Monthly, not quarterly — because drift happens between reviews.

You have 99.9% SLO commitments. Your on-call rotation is four people who are also building features. One more alert and they quit.

The retainer includes a named engineer in your on-call rotation for incidents in scope, with agreed SEV-1 response time in the SOW. The incentive is aligned — a retainer paid to keep incidents from happening, not to bill against them.

Your leadership asks for a monthly platform health update. Engineering gives them a Slack dump. The board wants something defensible.

A four-page monthly platform review memo — cost trend, reliability metrics, security posture, capacity headroom, top three risks — delivered by the 5th business day of each month. Built for a CTO to forward to a board without edits.

The Migration Engine

One discipline, every engagement.

Every migration runs the same six-stage pipeline: Inventory, Plan, Convert, Validate, Reconcile, Report. Human sign-off gates enforce the stage boundaries that matter.

STAGE 1 INVENTORY Read-only sweep STAGE 2 PLAN Wave + dependency map HUMAN GATE STAGE 3 CONVERT Rule library + handlers STAGE 4 VALIDATE Checksums + diffs STAGE 5 RECONCILE Daily diffs, parallel run STAGE 6 REPORT Co-signed cutover doc Scope sign-off required before conversion begins
Read how the engine works stage by stage
Why us

The rest of the market vs. what we do differently.

Every promise in this table lives in the contract, not the pitch deck.

Dimension Traditional T&M SI Replatform
Pricing model Time & materials — final cost unknown at project start Fixed-fee against a written Appendix A — no surprises
Staffing Senior pitched, junior delivered — bench economics drive the swap The engineer on the first call is the engineer in the repo
Validation Verbal sign-off or sampling — "it looks right" Deterministic row counts, checksums, query diffs — signed artefact you keep
Scope creep Absorbed into T&M — change orders often verbal, billed later Anything out-of-scope is a written Change Order before work begins
Timeline Months of ramp, discovery, re-discovery, re-scoping Fixed Discovery Sprint produces inventory + wave plan before you commit to execution
How to engage

How to start a platform management retainer.

A quick cost read, a pre-retainer assessment to scope the SOW, then ongoing management. Every engagement starts with an explicit scope and a named engineer.

01 · Free

FinOps Quick-Check

Free
2 minutes · no email

If cost governance is part of your platform management concern, start here — a quick waste-range estimate before scoping a retainer that includes FinOps guardrails.

Run the Quick-Check
02 · Cost review

Cloud Cost X-Ray

Free
through Q3 2026 · 90-min live session

Useful pre-retainer to understand your current spend shape and whether cost management should be a primary or secondary focus in the retainer scope. After Q3 2026 reverts to $100.

Send me your bill
03 · Onboarding sprint

Platform Assessment

Fixed-fee
scoped per engagement

A 2-week platform assessment: current state mapping, monitoring gaps, IAM posture, compliance gaps, incident-response runbooks. Produces the retainer SOW scope and baseline posture document.

Start with a platform assessment
04 · Retainer

Platform Management Retainer

Contact
scope-quoted · quarterly auto-renew

Essential (monitoring + IAM + patching), Standard (adds DR + security posture + incident participation), or Premium (adds compliance management + higher SLA). Named senior engineer, no ticket portal.

Talk to us
Founder-led delivery

The people doing the work.

Yash Maheshwari
Founder · Replatform

Cloud and data platform engineer with 6+ years across migrations, lakehouse architecture, FinOps, and managed platform operations on AWS, GCP, Azure, Snowflake, Databricks, BigQuery, and Redshift. The engineer on your first call is the engineer in your repo — no bench hand-offs, no junior substitutions.

Not ready to talk?

Download the Migration Readiness Checklist

If a migration or major platform change is in your roadmap alongside ongoing management, this checklist covers what to have in order first — runbooks, IAM posture, monitoring baselines, and compliance constraints.

Get the checklist