Operations
How to staff a data platform without an org redesign
Central bottleneck or embedded sprawl is a false choice. The topology that works at 50 to 500 engineers, and the one hire that decides it.
The Cloud Practice4 min readOperations
Data platform staffing swings between two failure modes. The central team that owns everything becomes a ticket queue with a six-week latency, and every domain team's roadmap quietly becomes the platform team's backlog. The fully embedded model fixes the queue and dissolves the platform: five teams, five ingestion patterns, five half-built orchestration wrappers, no one on call for the whole.

The topology that works in the 50-to-500-engineer range, consistent with the platform-as-product thinking in Team Topologies, is thin-platform, embedded-analytics:
A small platform team owns the paved road. Ingestion patterns, the orchestrator, the warehouse or lakehouse, contracts tooling, cost visibility. Its customers are internal engineers; its product is "boring by default". It says no to bespoke pipelines and yes to making the paved road cover one more case per quarter.
Analytics engineers sit in the domains. They own models, metrics and marts for their business area, on the platform's rails. Domain knowledge stays where the questions are; the platform team never becomes the bottleneck for a marketing question.
The interfaces are contracts, not meetings. Boundary contracts and a shared catalog do the coordination that liaison roles pretend to do.
sizing the thin platform
Thin has a number. In the 50-to-500 range we see healthy platform teams at three to six engineers, and the ratio matters less than the trend: a platform team that grows faster than its consumer base is absorbing work the domains should own, usually because saying yes to bespoke requests feels helpful in the meeting where it happens. The counter-force is a public paved-road definition. One page, versioned, listing what the platform supports and, more importantly, the current answer to "what if my case is not covered": you build it on our primitives, you own it, and if three teams build the same thing we adopt it onto the road. That last clause is how the road grows without the team bloating, and it converts the domains from ticket filers into contributors with a promotion path for their tooling.
the migration between topologies
Most organizations we assess are not choosing a topology; they are escaping one. From the central-queue starting point, the move is to stop accepting model-level work first, publish the paved road, and transfer mart ownership domain by domain, each transfer paired with an embedded analytics hire or a renamed existing analyst. From the fully-embedded starting point, the move is the reverse: inventory the five ingestion patterns, pick the one with the fewest 2 a.m. surprises, and consolidate behind it one pipeline at a time. Both migrations take two to four quarters and both fail the same way, by announcing an org change before the paved road exists for anyone to move onto. Roads first, then reorganization. The reverse order produces a memo and a mess.
how to tell it is working
Three measurements, taken quarterly, tell you whether the topology holds. Queue latency on platform requests: the paved road exists to make most requests unnecessary, so the queue should shrink even as the company grows. Divergence count: the number of pipelines running off the road, which trends down when the road is good and up when the platform team has been saying yes in meetings again. And the on-call weather report: pages per week for the platform team, because a thin platform with healthy boundaries sleeps, and a platform team that does not sleep is absorbing operational debt the domains created. None of these requires a survey or a maturity model. All three come out of systems you already run.
Two edge cases we are asked about in most engagements. Below fifty engineers, do not build a platform team at all; one senior data engineer with strong paved-road instincts and a part-time analyst per domain covers everything, and the topology conversation is premature by two funding rounds. Above five hundred, the thin platform usually splits into two: an infrastructure layer (compute, storage, catalogs) and an enablement layer (tooling, contracts, paved-road advocacy), because the customer conversations and the capacity planning stop fitting in one team's head. Both splits are fine. What does not survive scale is the liaison committee that gets proposed instead, the "data council" whose members represent teams without authority to bind them. We have never seen one produce anything but minutes.
The one hire that decides it: the first platform engineer must be someone who treats internal teams as customers. A brilliant infrastructure engineer who despises stakeholder conversations will rebuild the ticket queue with better tooling. We look for evidence of shipped internal products, not badge-collection, and in assessments this single role explains more platform outcomes than any technology choice on the diagram.