TRM

Case study: platform engineering and IT

Explaining why platform engineering (DevOps, SRE, Data, QA and so on) and IT matter is hard, and not only for non-technical people. Each team also works in a very different space, which makes it hard to compare the value of their initiatives. This exact problem is what led me to create TRM.

Everything these teams create falls into time, risk or money. That made it easy to show progress every week, month and quarter. Our quarterly reviews looked something like the charts below. None of them is actual data.

Time

For time, we focused on where we could save the company the most. We showed time saved as a share of one full-time employee (FTE), to give it context. Hours sound impressive, with over two thousand working hours a year per FTE, but they don’t tell the company what it got. We also broke the savings down by the teams and departments that received them.

Illustrative bar chart of FTE time saved by quarter, rising steadily from Q1 to Q4, where it passes two full-time employees' worth. Not actual data.

This one shows more than two FTEs’ worth of time saved by the end of the year.

Risk

We measured each risk by impact and likelihood. That gave every risk a score we could compare. The trick is showing where you are and where you came from.

This is my favorite visualization: a radar chart. The goal with risk is to reduce it as far as you can, but it will never be eliminated. On a radar chart the area shrinks as the company’s exposure shrinks.

Illustrative radar chart of risk across infrastructure, compliance, backup, access, availability and monitoring, with a start, current and future line. Each line encloses a smaller area than the one before. Not actual data.

The chart shows at a glance where the company started, where it stands now and where it’s headed. At each point you can see where you’re doing well and where the opportunities are. Like time, it can be broken down further, by vendor for example.

Money

Money is the most straightforward of the three. Cutting vendor spend was our focus, so it’s where we spent the most time. I wish I could share our real tracking. Here is roughly what it looked like.

Illustrative bar chart of cumulative vendor savings by month, January to December, against a flat target line. Savings pass the target in March and keep climbing to about three times the target by December. Not actual data.

The teams

Each team has its own focus. How does each one line up with TRM? In each triangle, the thick yellow shape is the value the team creates and the thin pink shape is its price.

Every score here will vary from company to company and year to year. The bigger a team’s customer base, the more value it can provide.

DevOps (+43%)

TRM triangle for a DevOps team: the yellow value shape reaches far toward risk and partway toward time and money, while the pink price shape reaches far toward money.

DevOps teams focus on CI/CD, deployments, production availability and security. How many people did I just confuse?

Value. DevOps serves three customers: product engineering, risk management and finance. For product engineering, it provides a structured, consistent environment where teams can test code and release it to production without dealing with the complexity of making that happen. That is time. For risk management, DevOps (or DevSecOps) makes sure environments, production above all, follow security and compliance best practice. And at SaaS companies especially, DevOps often holds the keys to the most expensive vendor spend. That makes it finance’s best friend or worst enemy.

Price. Like every team, DevOps costs money. DevOps engineers are paid above average, and they often take longer to onboard. That is time.

SRE (+38%)

TRM triangle for an SRE team: the yellow value shape reaches far toward time and risk, while the pink price shape reaches far toward money.

SRE teams focus on production monitoring, observability, alerting and availability. Clear?

Value. At a SaaS company, SRE has three main customers: risk management, product engineering and customer support. For risk management, SRE minimizes the impact of production incidents, by making them less frequent and handling them quickly. For product engineering, it makes it easy to tell when there is an incident (and when there isn’t), and provides the tools and training to resolve it quickly. That is time. It also provides processes so incidents don’t repeat. Customer support gains time too: it can see the scope of an incoming problem. Just this user? People on the east coast? Everyone?

Price. Like any team, there are people costs. There is also the cost of the observability tools.

QE/QA (+29%)

TRM triangle for a QA team: the yellow value shape reaches all the way to the risk corner and partway toward time, while the pink price shape reaches furthest toward time.

QA and QE teams make sure unit, manual, integration, regression, smoke and load testing happens before a release. Obviously?

Value. QA’s main customers are risk management and product engineering. For risk management, QA reduces the chance that a change to production causes an incident, through all those kinds of testing. For product engineering, the sooner a bug is found, the exponentially less time it takes to fix, so they can focus on delivering their own value.

Price. Time is the biggest expense: testing slows development down significantly. There is a people cost too, but QE engineers usually cost less than other engineers.

Platform (+32%)

TRM triangle for a platform team: the yellow value shape reaches almost to the time corner, while the pink price shape reaches far toward money.

I haven’t seen this team anywhere else, but we had a team focused on the cross-functional parts of software development. Clear as mud?

Value. Snagajob’s platform team’s main customer is product engineering. It cuts the time spent developing new products and features.

Price. Mostly people. Its tooling costs much less than the other platform engineering groups’ do.

Data (+61%)

TRM triangle for a data team: the yellow value shape reaches almost to the time corner and well toward risk, while the pink price shape reaches far toward money.

At SaaS companies, data is gold. Transactional data stores, message buses, data warehouses and lakes, ML and AI, data visualization, reporting and backups power every part of the company.

Value. Data’s main customer is every employee, particularly in SaaS. Accurate, secure, timely access to data cuts down on communication across the company, gives direction, and creates a competitive advantage.

Price. Data engineers cost more than most engineers. The tools to store, retrieve and back up data are expensive too.