← projects
The Before · Experiment Results

Eleven+ columns, no recommendation.

experimentation platform · internal The experimentation platform's dense metrics table: rows of statistics with no summary or recommendation
fig. i · the metrics view · gpt image 2 mock

what the table couldn't say

  1. 01

    a wall of numbers

    Readers face eleven+ columns of statistics and still have to make the ship call themselves.

  2. 02

    one view, four audiences

    The detail suits data analysts. Engineers, product managers, and S&O kept asking where their experiment stood.

  3. 03

    health checks need context

    The platform flags exposure and imbalance issues. Non-analysts struggled to judge which ones could change the decision.

← projects

The Solution

I proposed two projects.

Scorecard gives teams an answer. Conclude guides the next steps.

Part 01 · Scorecard

Scorecard

One consolidated results view that recommends an action.

  1. Overview, variants, and success + guardrail metrics in one place
  2. A recommendation up front: wait, ship, revert, review, or diagnose
  3. Health checks explained in plain language
Part 02 · Conclude Workflow

Conclude

A guided workflow to conclude an experiment.

  1. Six steps: variants, checks, guardrails, metrics, learnings, ship
  2. Every field pre-fills from the analysis, so nothing is re-typed
  3. The platform saves the final results and ships the winning variant
← projects
Goals & Team

Two goals, a team of four.

the goals
01

fewer status asks

Cut #ask-experimentation status asks by about a quarter.

02

conclude in minutes

Bring concluding an experiment from four hours down to minutes.

the team

Me

led the projects · front end · leadership syncs

Back-end engineer

conclude apis · data model

Data scientist · Product manager

decision rules · requirements

Stakeholders

analytics leadership · engineers & pms across orgs

← projects
The Scorecard

The same metrics, one recommendation.

experimentation platform · scorecard The Scorecard view: a slim ship-treatment recommendation strip, a treatment variant tile with passing success and guardrail pills, and short success and guardrail metric tables
fig. iii · the scorecard · gpt image 2 mock
← projects
The Scorecard · Under the Hood

The full analysis, one recommendation.

~6 parallel queries cached + memoized backend apis GetHealthCheck mht guardrail check GetAnalysisResults primary + guardrail GetDimensionResults dimensional slices GetExperimentAnalysis report + conclude state front-end interpretation layer Significance p < 0.05 + direction check ? Per-variant status gains / losses / neutral ? Priority-ordered decision engine ? One recommendation per analysis Ship Revert Review Wait Diagnose
← projects
The Impact

How it went.

the goals
01

fewer status asks

~40% fewer

The goal was 25%. Teams found the experiment status in Scorecard before opening a thread.

02

conclude in minutes

10 minutes

Owners used six guided steps instead of a doc, form, and spreadsheet.

extra wins

Demoed at the analytics offsite

cited in the platform's h1 accomplishments

Base of the leadership dashboard

v2 is driving broader adoption

Business-impact math

exploring one shared formula next