Skip to content

Top opportunities, overall

1-tech-capability

Skills silos across data ops, applications and domain knowledge

Symptoms

The same individuals always pick up the same work. Pipeline changes wait for one engineer; when they are on leave that work stops rather than moving to whoever is free. Requirements and troubleshooting route to named people the same way. The group cannot cover itself, and the person who holds the knowledge cannot take leave without bracing for what they will come back to.

Signals & measures

  • Highest-voted item in the workshop: 16 of 52 dots across 9 participants
  • SME self-rating 1.7/3 — the lowest of the five needs rated
  • Ranked #1 in the spend stack at $340k, so this is a funded, live gap rather than a neglected one
  • "If Engineer C is on leave the pipeline work stops" — workshop, 2026-07-16

2-delivery-metrics

Lead time is dominated by waiting, not by working

Symptoms

A change is finished and then sits. Engineers describe the queue rather than the build as the thing they wait on. Delivery health itself is rebuilt by hand for each steering meeting, so performance is argued about rather than read. Nobody shares one picture they trust, so the room spends its calm on reconciling spreadsheets.

Signals & measures

  • Median lead time 19h, of which 14h is review wait — 74% of elapsed time (platform dashboard, 90-day median, 2026-08-01)
  • Three approvers required on every change regardless of size; "3 approvers for a one-line change" recorded in the S44 retro
  • "Review queue" raised in three of the last four retros as a standing item with no owner
  • SME self-rating 1.2/3; no shared dashboard, three private spreadsheets identified
  • Assumes the dashboard reports change-failure rate as a fraction (0.12 entered as 12%) and MTTR in minutes (45 entered as 0.75h). Neither unit is labelled; both to be confirmed with the platform team

1-people-purpose

No slack capacity for upskilling, incident response or design work

Symptoms

Every hour is committed before the week starts, so anything unplanned is paid for out of the next planned thing. Improvement items are carried forward rather than dropped, which reads as intent without capacity. The week starts already behind, so there is no slack left for learning, incidents or design — only the next planned item.

Signals & measures

  • SME self-rating 1.9/3; 5 of 52 workshop dots
  • "Reduce review turnaround" carried as a standing retro item for four sprints with no owner and no action
  • No protected improvement time identified in the last two quarters
  • Leadership reads this as headcount ("we need 4 more heads"). We have not tested that. On the evidence so far, 74% of lead time is review wait rather than work, which four more engineers would add to rather than relieve
  • What would settle it: throughput per engineer before and after the approval policy changes

3-incident-mgmt

On-call carries knowledge that was never written down

Symptoms

Restoring service depends on who answers the page rather than on a runbook. On-call is something people dread, because the outcome hangs on who happens to be holding the knowledge that night.

Signals & measures

  • 9 of 52 workshop dots; SME self-rating 1.5/3
  • 6 Sev1/2 incidents a month across 3 teams (platform dashboard, 2026-08-01)
  • "Nobody wants to be on call for quota-keeper" — workshop, 2026-07-16
  • The internal findings doc (2026-08-05) records incident response as "fine — we have a rota". The workshop and the incident rate disagree with that, and the gap is worth raising directly

Internal platform friction

Teams route around the internal platform rather than through it

Symptoms

The platform is used where it is mandatory and avoided where it is not. Teams do not trust it enough to bet a delivery date on it. This heading is deliberately not a need id: the workshop named the theme before we had a need to hang it on, so it is held here, unranked, until it is mapped.

Signals & measures

  • 8 of 52 workshop dots; ranked #2 in the spend stack at $180k
  • "We use the platform where we have to and route around it where we can" — workshop, 2026-07-16
  • A prior vendor assessment (2026-06) named toolchain modernisation as the highest-leverage investment, but showed no measurements and held a tooling partnership. We have not confirmed that reading, and the retro evidence points at approval policy rather than tooling
  • What would settle it: adoption figures per platform service, held against the review-queue data above

tbc

Test data availability — raised in the workshop, not yet placed against a need

Symptoms

Named on the wall with 3 of 52 dots. Not discussed long enough to establish whether this is an environments problem, a data-governance one, or a test-design one, and it is not yet clear which need it belongs to.