Platform Engineering: Building an Internal Developer Platform Teams Actually Use
Most internal platforms fail not for technical reasons but because they were mandated instead of made useful. Here is how to build one developers choose.
Somewhere between the tenth and fiftieth service, every engineering organisation discovers the same thing: the constraint on shipping software is no longer writing it. It is the eleven steps between a finished feature and production, six of which involve waiting for a different team.
Platform engineering is the response. Done well, it turns infrastructure from a ticket queue into a product. Done badly, it produces an expensive internal tool that teams route around while telling the platform team it is going well.
The problem platforms are meant to solve
Without a platform, cognitive load is distributed to every application team. Each one independently decides how to build container images, wire CI, configure secrets, emit metrics, define alerts, handle database migrations and pass a security review. The results are inconsistent by definition, and the cost is paid three times: once in duplicated effort, once in operational variance, and once in the review meetings that exist to manage the variance.
The platform's job is to absorb that decision-making into a supported default, so application teams spend their attention on the domain problem their employer is actually paying them to solve.
Golden paths, not golden cages
The central design concept is the golden path: an opinionated, fully supported route to production for a common case. A new HTTP service should be creatable from a template that arrives with CI, container build, deployment pipeline, secrets management, structured logging, metrics, tracing, health checks, a security scan and an ownership record already configured.
Two rules make this work rather than fail:
- The path must be genuinely faster. If the golden path takes three days and the workaround takes two hours, you have built bureaucracy with better branding.
- Deviation must be permitted, at a cost. Teams with legitimate reasons to differ should be able to, while accepting the operational burden themselves. Mandates produce compliance theatre; better defaults produce adoption.
The four capability layers
| Layer | What it provides | Failure if missing |
|---|---|---|
| Infrastructure orchestration | Provisioning environments, databases, queues by declaration | Ticket queues and bespoke environments |
| Delivery pipeline | Build, test, scan, deploy, progressive release | Every team maintains its own snowflake CI |
| Observability by default | Metrics, logs, traces, SLOs pre-wired | Incidents diagnosed by guesswork |
| Developer portal | Catalogue, ownership, docs, templates, self-service actions | Nobody knows what exists or who owns it |
Build these in that order. A portal over a platform that cannot provision anything is a very tidy list of things you still have to raise tickets for.
Platform as a product, not a project
This is the shift most organisations get wrong. A platform team funded as a project delivers a launch and disperses. A platform team run as a product has:
- Named users — the application teams — whose satisfaction is measured, not assumed.
- A roadmap driven by demand, informed by where teams currently waste time.
- Documentation treated as an interface, because an undocumented capability does not exist.
- Support commitments for the paths it declares golden.
- Deprecation discipline, with migration paths rather than announcements.
The single best diagnostic question for any platform team: can you name the last three things you removed because users found them unhelpful? Teams that cannot answer are building for themselves.
Anti-patterns that kill platforms
- The mandate. Adoption enforced by executive memo produces malicious compliance and hidden shadow infrastructure.
- The abstraction that leaks under pressure. If a developer must understand your abstraction and the underlying system to debug an incident, you have added a layer without removing a burden.
- Portal-first delivery. A beautiful catalogue over no self-service capability is documentation with a build pipeline.
- Rebadging the old ops team. New name, same ticket queue, worse morale.
- Ignoring the local development experience. If getting the service running on a laptop still takes a day, nothing else you built will be forgiven.
Metrics that reveal the truth
Measure outcomes for application teams, not activity for the platform team.
- Time to first deploy for a brand-new service — the clearest single indicator.
- Change lead time and deployment frequency across teams on and off golden paths.
- Change failure rate and recovery time, to prove speed did not cost stability.
- Golden path adoption percentage, tracked voluntarily rather than mandated.
- Tickets eliminated — the work the platform absorbed that used to require a human.
- Developer sentiment, surveyed regularly and acted on visibly.
A staged build sequence
Stage 1 (0–3 months): find the worst bottleneck by interviewing teams, not by architecture review. It is usually environment provisioning or the deployment approval chain. Solve exactly that, for two willing teams.
Stage 2 (3–6 months): productise it into one golden path with a template, and onboard three or four more teams. Write the documentation as though the reader is competent but new.
Stage 3 (6–12 months): add observability and security scanning by default, then introduce a service catalogue with clear ownership.
Stage 4 (12+ months): add self-service actions, cost visibility per team, and progressive delivery capability. Begin deprecating the bespoke pipelines the platform has made redundant.
The measure of a successful platform is not how sophisticated it is. It is whether a new engineer can ship a small change to production safely in their first week — and whether the teams using it would object if you took it away.
Frequently Asked Questions
What is platform engineering?
Building an internal product — usually a platform with self-service capabilities and an internal developer portal — that application teams consume to ship software without raising tickets for infrastructure.
How is it different from DevOps?
DevOps is a culture and set of practices; platform engineering is the organisational structure that makes those practices sustainable at scale by productising them rather than expecting every team to reinvent them.
Do we need an internal developer portal?
Once you have a service catalogue nobody can hold in their head, yes. Below roughly twenty services, good templates and documentation usually suffice; beyond that, discoverability becomes the bottleneck.
What is a golden path?
A supported, opinionated route to production for a common case — a templated service with CI, observability, security scanning and deployment already wired. Teams may deviate, but then they own the consequences.
How do we measure platform success?
Time to first deploy for a new service, change lead time, change failure rate, percentage of teams on golden paths, and the volume of tickets the platform eliminated. Adoption is the real signal, and it must be earned rather than mandated.