Automation · Snagajob

Freeing up people with automation

Automating the data pipeline into Snowflake, and removing a deployment step that cost more than it saved.

The problem

Engineering, not just AI, can automate tasks and free people to do more valuable work. Two places at Snagajob needed it: getting product data into analytics, and deploying code.

Data. Snagajob’s strength is its data, and that data runs on Kafka. Kafka is useless for ad-hoc analytics, so we turned Kafka topics into Snowflake tables for real-time analytics. Getting there took several teams in sequence:

  • A product team asks for a topic.
  • The data team creates it and checks it follows best practice. The product team can’t start until this is done.
  • The product team ships code that writes to the topic.
  • The data team writes code to ingest the raw data.
  • The data team writes code to flatten it into easy-to-use tables.
  • Every new field on the topic repeats all of the above except the first step.

New features went live, and weeks passed before anyone could analyse them. No data was lost, however long it took.

Deployment. Deploying new code is risky. Problems usually come from change: in code, infrastructure or load. We had built a blue/green pipeline that ramped traffic slowly, to limit how much traffic saw a change at once.

How I thought about it

The data process was built on purpose, to make data follow patterns and to make up for past data-quality problems. But it cost far too much time and far too many people for what it protected.

The deployment ramp traded time for lower risk. That trade can make sense where bad deployments are common. So I checked how often ramping had actually saved us. In the last month: zero. The last quarter: zero. The last year: zero. We weren’t reducing the risk, and we were paying for it in time on every deployment and in backed-up deployment queues.

What I did

Data, as it works now:

  • A product team starts building straight away, with no hand-off.
  • The data team is told about the new topic and reviews it for anomalies, on its own schedule.
  • Code and infrastructure we built ourselves ingest every new topic into Snowflake, raw and flattened. New fields are flattened automatically too.

Deployment: here the fix was to remove automation. We took out the ramp.

Results

In 2023, automation cut operational headcount needs by 15%.

  • Data. Product teams build faster. The data team’s operational cost fell sharply, and the oversight is still there. The company sees how a new feature is doing straight away. Snowflake costs went way down, and AWS went up a little. There is one set of code to support instead of two per topic, and monitoring went from crude to full-featured, for free, on our internal libraries instead of a third-party tool that only logged.
  • Deployment. Forty-five minutes back to engineers on every deployment. It also cut tech debt (legacy infrastructure) and cost, because we ran two sets of infrastructure for less time.

What I’d do today with AI

The deployment story is the one I’d keep front of mind: I only found the waste by asking how often the safeguard had actually fired. A harness makes that question routine, because every retro records what a check caught and memory keeps the count. I’d have agents write the ingestion code, but I’d want the verify step to prove each new topic lands in Snowflake before anyone calls it done.