Topic 20 of 20 · Part V : Advanced Decision Intelligence

Business experimentation

Experimentation applies the same control pattern to the product itself, and becomes the outer feedback loop around the whole system.

The control pattern

Data asset

Customer behaviour under controlled product changes

Knobs

Audience, Intervention, Timing, Randomisation unit, Success metric, Guardrail metric

Outcome

Product improvements supported by evidence

Measurement

See how to measure this below

What it is

Business experimentation applies the same control pattern to the product itself. The data asset is customer behaviour under a controlled change: a new summary format, alert policy, onboarding flow, explanation panel or ranking presentation.

Audience, timing, channel, randomisation unit, exposure duration and success metric are the knobs. A well-designed experiment estimates whether the change caused an improvement, instead of relying on anecdotes or before-and-after comparisons that may be confounded by market conditions.

Successful experiments identify behaviours worth operationalising, and their clean examples can later support prompt updates or distillation. In this way, product experimentation becomes the outer feedback loop around the entire DataKnobs system.

The knobs in detail

Each row is one adjustable property of the data asset, and what moving it tends to do.

KnobWhat you adjustLikely effect
AudienceWho is exposed to the changeDetermines what the result generalises to
InterventionExactly what changedMust be isolated to be attributable
TimingWhen the test runsMarket conditions confound short windows
Randomisation unitUser, session or accountPrevents contamination between arms
Success metricWhat improvement meansThe most consequential choice in the design
Guardrail metricWhat must not get worseDetects harmful optimisation

Applied: Stocks Assistant

For Stocks Assistant, experiments might compare concise and detailed earnings summaries, test whether visible citations increase trust, evaluate different explanations of the Health Score, or measure whether personalised watchlist alerts improve return visits. Primary metrics should reflect user value : successful research completion, source inspection, comprehension, retention. Guardrails should detect increased unsupported confidence, excessive alerts, slower responses, reduced source usage or a rise in impulsive trading behaviour.

How to measure it

Evidence that the knob produced the intended behaviour, rather than shifting the problem elsewhere.

  • Effect on the pre-declared primary metric
  • Guardrail metrics, checked every time
  • Results across meaningful segments, without fishing for significance
  • Logged model, prompt, data and scoring versions used during the test

Common mistakes

Revenue or engagement alone as the definition of success for a financial-information product.

An unlogged AI change during the test window, which contaminates the intervention.