# AI Platform Operating Model

Source: John Bradshaw — https://bradshaw.cloud
URL: https://bradshaw.cloud/strategy/ai-platform-operating-model/
Type: insight
Tags: Platform Engineering, AI, Infrastructure

> AI does not replace platform engineering. It expands the platform’s job from delivery pipelines to model access, data context, evaluation and agent governance.

## The AI platform service catalogue

Common services should remove repeated setup without taking product ownership away from delivery teams.

### Make AI capabilities services, not exceptions

A useful platform exposes a small set of repeatable AI services: model access, data context, agent runtime, evaluation, serving, observability and policy. Teams consume a supported path instead of building a private stack around each pilot.

**Business take:** The catalogue gives each capability an owner, a service level and a default control. It also tells teams where the platform ends and their product responsibility starts.

### The control plane is more than an API gateway

A model gateway can provide credentials and routing. The AI control plane adds evaluated model changes, tool registration, data permissions, cost attribution and policy checks. It governs the whole runtime path.

**Business take:** A central policy point replaces scattered provider keys, untracked model changes and after-the-fact cost analysis with controls that work before the request runs.

### Extend the platform instead of creating a new empire

An AI silo centralises specialist skills but can become a bottleneck. An enablement overlay helps teams start but often leaves shared controls unfinished. A modular extension lets the internal platform own common services while product teams keep delivery ownership.

**Business take:** The practical boundary is undifferentiated AI heavy lifting: policy, cost, shared runtimes and common evaluation. Keep domain prompts, user journeys and business outcomes with product teams.

### AI needs an acceptance test at runtime

Code either passes a test or fails it. AI quality is probabilistic, so teams need representative task sets, release thresholds, trace sampling and a rollback path when a model or prompt change moves the result.

**Business take:** A platform that can report latency but not task quality only monitors half the system. Evaluation belongs beside deployment, not in a separate governance presentation.

## Platform maturity for AI

The operating model changes from pilot sprawl to a platform product with measurable adoption.

### AI platform maturity

A practical path for teams extending an internal developer platform

- Pilot sprawl — Teams buy models and build integrations independently — Duplicate risk
- Shared credentials — A common gateway without consistent delivery paths — Partial control
- Common AI services — Models, data and evaluation are available by default — Repeatable
- Paved AI paths — Self-service paths include cost, quality and policy controls — Governed speed
- Platform product — The roadmap follows adoption and measurable outcomes — Compounding

_Do not centralise domain decisions. Centralise the services that every team would otherwise rebuild with different controls._

## Platform decisions

### Which AI capabilities should be common services?

**Start with the undifferentiated heavy lifting.**

Model credentials, routing, evaluation scaffolding, common tool registration, observability and budget policy benefit from standardisation. Keep product prompts, workflow design and customer experience with delivery teams.

### Should the organisation form a separate AI platform team?

**Extend the existing platform where possible.**

A separate team can bootstrap expertise, but it often creates a second delivery path. The stronger pattern adds specialist capabilities to the platform products teams already use.

### Where should AI quality be checked?

**In the release path and in production traces.**

Test representative tasks before deployment, monitor task quality after release and keep a rollback path. Latency and availability remain necessary, but they do not prove the answer is useful.

### How should platform success be measured?

**Measure adoption and outcomes, not services provisioned.**

Track time to a safe first deployment, use of supported paths, cost attribution coverage, evaluation coverage and the number of teams that can work without a platform ticket.

## Frequently asked questions

### What is AI platform engineering?

AI platform engineering applies platform-product thinking to AI systems. It gives product teams supported paths for models, data context, agents, evaluation, deployment, observability, cost controls and policy enforcement.

### How is AI platform engineering different from MLOps?

MLOps focuses on the lifecycle of trained models. AI platform engineering covers the broader production estate: model access, application and agent runtime, tool and data controls, evaluation, cost attribution and developer self-service.

### Should every company build a separate AI platform?

No. Most organisations should extend their internal developer platform with modular AI services. A separate platform is justified only when a distinct operating environment, specialised workload or regulatory boundary demands it.

### What should an AI platform provide first?

Start with governed model access, cost attribution, a basic evaluation path and traceability. These services stop early pilots from becoming a collection of credentials and unrepeatable integrations.

## Related

- [Platform operating model](https://bradshaw.cloud/strategy/platform-operating-model/) — the platform-as-product foundation.
- [Developers as customers](https://bradshaw.cloud/strategy/developers-as-customers/) — measure whether the paved road is worth using.
- [FinOps for AI](https://bradshaw.cloud/strategy/finops-for-ai/) — attribution and policy before the invoice arrives.
- [AI Grid](https://bradshaw.cloud/strategy/ai-grid/) — the runtime and placement decisions underneath serving.
