Measuring UX at Scale

Designing a company-wide, in-product UX measurement system for Brightspace

Quantitative Strategic Research

At a glance

  • Situation: D2L had strong pre-release research but no shared way to see the experience after release — where it was improving, where friction persisted, or which user groups struggled most.

  • My role: Senior UX Researcher and project lead — owned the strategy for what we measured, how we prompted, how the data was interpreted, and how teams across the business used it.

  • Method: A three-layer in-product measurement system (UMUX Lite + Experience Metrics + Workflow Metrics), built cross-functionally with design, legal, customer experience, engineering/vendors, and BI.

  • Scale (FY26 Q2): 108k+ instructor and 370k+ learner prompt impressions across 500+ client organizations; 80% / 82% completion.

  • Biggest insight: Learners are already in a good place — the gap is instructors. A persistent, role-driven ~13-point UMUX gap, worst among K-12 instructors, made closing the instructor gap the highest-impact priority.


The problem

Before this, D2L shipped changes with no consistent way to see how the live experience was doing — hard to benchmark UX over time, prioritize follow-up, or show UX impact.

My goal was to design an in-product measurement system that could generate trustworthy quantitative signal at scale — strong enough to benchmark experience over time, specific enough to diagnose where problems lived, and careful enough to protect the user experience, privacy, and client trust. That made it as much a systems-and-alignment problem as a measurement one.



Building a system teams could trust

Before launching anything, I worked a set of strategic and operational decisions through partners across the company — and each partnership existed to make the system credible before it went live:

  • Design (principal designers): shaped the measurement approach and where prompting would be least disruptive.

  • Legal: guardrails around anonymity, PII, and age-to-consent.

  • Customer experience / customer-facing leaders: where clients might fear over-surveying, and what opt-out controls would build trust.

  • Senior developers + vendor partners: what was technically possible, how to control prompt rates reliably and cost-consciously.

  • BI: making sure incoming data could be analyzed usefully rather than becoming overwhelming.

This wasn't something I could launch alone in Qualtrics; it required real roadmap and development investment, which is why cross-functional buy-in was the work.



1. What to measure — three connected layers

A single metric couldn't answer every product question, so we designed three layers, each solving a different problem:

  • UMUX Lite — the top-line benchmark. Two 7-point items (easy to use; meets my needs), lightweight enough to scale, established enough to trust, and interpretable against usability benchmarks. A credible baseline to track over time.

  • Experience Metrics — the diagnostic layer. Built on the experience areas the org already used (Administration, Assessment, Data & Analytics, Finding Learning, Monitoring & Results, Teaching & Learning), so the system matched how teams understood the product — 10 instructor and 5 learner prompts, each tied to a major area.

  • Workflow Metrics — the targeted layer. Time-bound, role-specific prompts tied to a launch, for pre/post-release measurement, without overloading the benchmark.

Each layer answers a different question: UMUX Lite = overall benchmark; Experience Metrics = which product areas shape it; Workflow Metrics = the impact of a specific change.


2. Who, where, when, and how often to prompt

The harder problem was collecting feedback without negatively impacting the product experience. The guardrails:

  • Who: started with instructors and learners (the two largest, most experience-defining groups) for statistical usefulness and a controlled rollout; added administrators as the system matured. Learners under 13 were never prompted — excluding K–12 learners entirely.

  • Where: on tool landing pages, never mid-workflow — a firm decision not to interrupt task completion.

  • How often: no more than once per year, low prompt rates, and user/client opt-out — direct answers to customer-facing concerns about over-surveying.


We piloted internally, then phased the rollout (UMUX Lite for instructors → learners → additional experience prompts), which built organizational trust and gave us room to resolve vendor and accessibility issues before scaling.


3. What the data revealed

Engagement was strong enough to trust the signal. 108k+ instructor and 370k+ learner prompt impressions across 500+ client organizations, with 80% completion.

A persistent role gap. The most consistent signal across the year was a gap between instructors and learners. Learners rated Brightspace "very good" (UMUX 79); instructors "marginal/okay" (66) — a ~13-point gap that was both statistically and practically significant, and stable quarter over quarter. On Top-2-Box (the share rating 6–7 of 7) the gap was even wider: learners ~65% vs. instructors under 50%.

I moved past averages — reading rating distributions, not just means — and broke UMUX down across segment, region, client size, and product usage:

  • Client size and product usage: no meaningful effect.

  • Segment: K–12 instructors scored lowest.

  • Region (the strongest signal): LATAM instructors rated markedly higher than North American instructors — statistically and practically significant — which I flagged for future strategic research into what drives the stronger LATAM experience and whether it could help close the North American gap.



Experience Metrics. At the task level, instructors were strongest on communicating with and evaluating learners, and weakest on data & analytics and centrally managing content. Because the low scores were broad rather than concentrated, the instructor problem was systemic — not one bad tool.

Throughout, I framed the program's limit plainly: it shows what is happening and at what scale — not why on its own — so it's read alongside qualitative research.



  1. Outcomes and impact

The program gave D2L something it hadn't had: a consistent, company-visible way to measure and compare UX across Brightspace over time — benchmark scores, role-based comparison, and a clear view of where the experience was strongest and weakest. It also made UX impact visible: instead of relying only on pre-release findings, teams could now see where usability was improving, flat, or in need of follow-up.

The organizational impact came from reframing, not just reporting. The key story wasn't "the numbers moved" — it was "the instructor experience is consistently weaker, and that gap is strategically important." The value wasn't "we have a dashboard" — it was a trustworthy signal that guides research and product decisions. That synthesis is what turned the program into a strategic input rather than a reporting mechanism.


Reflection

The biggest lesson: measurement is only useful when it's designed to support decisions — not to collect more feedback for its own sake, but to create a signal teams can trust, interpret responsibly, and use alongside other evidence.

If I continued, I'd make the path from signal to action more explicit and repeatable: define clearer trigger points for when a pattern moves from "monitor" to "investigate," build a structured drill-down from broad benchmarks into experience areas and segments, and connect those signals more systematically to workflow-level measurement and targeted qualitative research — so a result like the instructor–learner gap leads directly to a concrete next step.

Ready to build something amazing together?

I'd love to connect with you!

Ready to build something amazing together?

I'd love to connect with you!

Ready to build something amazing together?

I'd love to connect with you!