Core idea
Every change follows the same control pattern
A data asset is the thing you change. Knobs are its adjustable properties. The outcome is what should change when a knob moves, and the measurement is the evidence that it did : rather than the problem moving somewhere else.
Data asset
The thing you change : training examples, retrieved knowledge, test cases, user context, feedback, source records.
Knobs
The adjustable properties of that asset : quality, freshness, weighting, segmentation, permissions, thresholds, retention.
Outcome
The measurable behaviour that should change when the knob moves. Stated before the change, not after.
Measurement
The evidence that the change helped rather than shifting the problem somewhere else.
“Data is where the leverage is. The properties of a data asset are the knobs. Metrics show whether moving a knob produced the intended behaviour.”
Why it matters
What the framework buys you
A shared language
Product, data, AI, risk and engineering argue about the same object : a named data asset with named knobs : instead of trading opinions about “the model”.
Traceable results
Every user-facing answer connects back to sources, transformations, model versions, policies and metrics.
Cheaper improvement
Most quality problems are data-asset problems. Turning a knob is faster, cheaper and more reversible than changing the model.
Earned trust
Users see not just a score, but which inputs drove it, how current they are, what assumptions applied, and how confident the system should be.
The guide
Twenty topics, five parts
Each topic is one data asset with its own knobs, outcome and measurement. Open any of them for the detail, the worked Stocks Assistant application, and the metrics that prove the knob moved what you intended.
Shaping Model Behaviour
What a model learns, and how reliably that behaviour transfers to production.
Topic 1
Fine-tuning
Fine-tuning turns a general model into a specialist by repeatedly showing it the inputs and outputs that define the desired behaviour.
- Data asset
- training examples
- Knobs
- example quality, volume, task mix, difficulty, weighting, curriculum order
- Outcome
- specialised and repeatable model behaviour
Read the topic →
Topic 2
Retrieval-Augmented Generation
RAG changes model behaviour by changing the evidence placed in the model's context at answer time, rather than changing the weights.
- Data asset
- retrieved knowledge
- Knobs
- source hierarchy, chunk size and overlap, metadata, freshness rules, top-k and reranking, citation threshold
- Outcome
- grounded answers tied to current evidence
Read the topic →
Topic 3
Prompt optimisation
Prompt optimisation treats instructions, demonstrations and output formats as a configurable data layer : the fastest and most reversible control available.
- Data asset
- instructions and context examples
- Knobs
- example selection, ordering, diversity, output schema, tool instructions, response constraints
- Outcome
- more consistent responses without retraining
Read the topic →
Topic 4
Synthetic data
Synthetic data fills gaps in training and evaluation : but only when it is designed from a coverage plan rather than produced as a large undifferentiated batch.
- Data asset
- generated training and test examples
- Knobs
- generator model, constraints, coverage plan, rarity, perturbation level, human review rate
- Outcome
- coverage of important cases that real data does not supply reliably
Read the topic →
Grounding, Evaluation and Orchestration
What evidence the system uses, how quality is measured, and how work is routed.
Topic 5
Evaluation
Evaluation data determines what the organisation can see, and therefore what it can improve.
- Data asset
- test datasets and scoring rules
- Knobs
- scenario mix, difficulty, time split, segment coverage, scoring method, pass thresholds
- Outcome
- reliable measurement of system quality
Read the topic →
Topic 6
Hallucination control
Hallucination control is the disciplined management of what the model is allowed to claim : and requiring any citation is not enough.
- Data asset
- grounding evidence and claim verification
- Knobs
- source quality, required citations, retrieval threshold, claim materiality, abstention policy, verifier strictness
- Outcome
- fewer unsupported or misleading claims
Read the topic →
Topic 7
Model routing
Routing is not merely cost optimisation : it improves quality by sending each request to the component best suited to it.
- Data asset
- request characteristics and runtime signals
- Knobs
- complexity, domain, risk, confidence, latency and cost budgets, tool requirements
- Outcome
- the right model and workflow for each request
Read the topic →
Topic 8
Agent systems
An agent's behaviour depends on the data describing its environment : tools, observations, permissions and workflow state become knobs for autonomy and risk tolerance.
- Data asset
- tool, state and workflow data
- Knobs
- available tools, permissions, memory, planning depth, stop conditions, action budget
- Outcome
- better multi-step decisions with controlled execution
Read the topic →
Personalisation and Responsible AI
Adapting to the user while keeping safety, fairness and review boundaries explicit.
Topic 9
Personalisation
Good personalisation does not change objective facts. It changes prioritisation, explanation and presentation.
- Data asset
- user-context data
- Knobs
- relevance, recency, consent, retention, risk profile and horizon, portfolio exposure
- Outcome
- more useful responses for the individual user
Read the topic →
Topic 10
Safety and guardrails
Guardrails should be layered: a prompt alone is too fragile, while a rigid rule set alone cannot interpret every context.
- Data asset
- risk examples, policies and runtime checks
- Knobs
- risk categories, coverage, thresholds, adversarial cases, permissions, escalation rules
- Outcome
- safer outputs and actions
Read the topic →
Topic 11
Bias and fairness
Bias in a financial AI system often begins with uneven data representation : abundance mistaken for quality.
- Data asset
- representation in training, retrieval and evaluation data
- Knobs
- group balance, market coverage, labels, peer definitions, counterfactual examples, weighting
- Outcome
- more equitable and less systematically distorted behaviour
Read the topic →
Topic 12
Active learning
Active learning chooses which production examples deserve human review instead of labelling data at random.
- Data asset
- selected production examples for human labelling
- Knobs
- uncertainty threshold, disagreement, sampling method, business impact, novelty, review priority
- Outcome
- faster improvement with fewer labelled examples
Read the topic →
Data and Learning Operations
Keeping sources, feedback, privacy and monitoring healthy after deployment.
Topic 13
Feedback learning
Feedback signals are not equally trustworthy : a click may indicate curiosity rather than agreement.
- Data asset
- user reactions and expert corrections
- Knobs
- feedback type, reviewer reliability, weighting, recency, context, aggregation
- Outcome
- behaviour that improves with real-world use
Read the topic →
Topic 14
Data quality
A sophisticated model cannot reliably repair a pipeline that confuses millions with billions or maps a ticker to the wrong legal entity.
- Data asset
- source records and canonical data models
- Knobs
- completeness, validity, consistency, identity resolution, units and adjustments, lineage and deduplication
- Outcome
- dependable downstream analysis
Read the topic →
Topic 15
Privacy
The first and strongest privacy knob is collection: data that is not necessary should not be gathered.
- Data asset
- sensitive user and portfolio information
- Knobs
- collection, consent, redaction, retention, access, deletion and auditability
- Outcome
- useful personalisation with lower privacy risk
Read the topic →
Topic 16
Data drift
Financial markets make drift normal rather than exceptional : and an alert should not automatically trigger retraining.
- Data asset
- changing production inputs and relationships
- Knobs
- monitoring window, segment, baseline, alert threshold, seasonality, retraining trigger
- Outcome
- early detection of changing conditions
Read the topic →
Advanced Decision Intelligence
Extending the framework to complex documents, auditable scores, causal questions and product decisions.
Topic 17
Multimodal AI
Many financial facts lose meaning when reduced to plain text. Treating layout as data preserves the relationships that make a number interpretable.
- Data asset
- text, tables, charts, images and document layout
- Knobs
- resolution, ocr engine, page selection, layout preservation, modality balance, cross-modal validation
- Outcome
- better understanding of complex financial documents
Read the topic →
Topic 18
Explainability
An explanation should be faithful to the actual system rather than a persuasive story generated after the fact.
- Data asset
- evidence, metadata and model trace data
- Knobs
- provenance depth, factor visibility, lineage, citation granularity, confidence reporting, version disclosure
- Outcome
- auditable decisions and understandable scores
Read the topic →
Topic 19
Causal inference
Even when causal certainty is impossible, the framework forces the product to state which assumptions connect the evidence to the conclusion.
- Data asset
- observational and experimental data
- Knobs
- treatment, control, outcome window, confounders, matching, identification assumptions
- Outcome
- better distinction between correlation and cause
Read the topic →
Topic 20
Business experimentation
Experimentation applies the same control pattern to the product itself, and becomes the outer feedback loop around the whole system.
- Data asset
- customer behaviour under controlled product changes
- Knobs
- audience, intervention, timing, randomisation unit, success metric, guardrail metric
- Outcome
- product improvements supported by evidence
Read the topic →
Method
How to turn a knob without breaking something else
Name the data asset
Which data asset are you actually changing? If the answer is “the model”, keep looking.
State the outcome first
Write the behaviour that should change, and by how much, before you touch anything.
Move one knob
Change a single property. Two at once and you cannot attribute the result.
Measure by slice
Score the segments that matter operationally, not one headline average.
Check for displacement
Confirm the problem moved off the board rather than into another segment.
The rule that keeps this honest: a change ships only if it improves the intended metric without crossing a guardrail elsewhere. That single sentence turns evaluation from a report into a control.
Operating model
Twenty topics, one loop
Not twenty initiatives : four stages that feed each other, continuously. What you learn in the last stage changes what enters in the first.
Distillation comes last: it compresses stable, well-evaluated behaviour into efficient student models : never behaviour that has yet to be validated.
Sequencing
What to build first
The order matters more than the ambition.
Foundations
Canonical source metadata and point-in-time evaluation
Nothing above this layer is trustworthy without it
Evidence
Claim-level citations and confidence-based escalation
Now the system can say what it does not know
Efficiency
Distil high-volume extraction, classification and summary tasks
Compress only what is already stable and measured
Product knobs
Expose score, personalisation, risk and time-horizon controls to users
The framework becomes visible to the people it serves
Only then: automate more complex agent workflows, or treat ranking changes as proven signals. Sequencing this the other way produces confident systems built on unverified ground.
Artifacts
The minimum you must be able to show
If a control layer has no artifact, it is an intention rather than a control. A good test: pick any number on the screen and ask which artifact proves it.
| Control layer | Minimum production artifact |
|---|---|
| Sources | Authority ranking, timestamps, fiscal-period metadata and lineage |
| Models | Versioned prompts, student/teacher routing, schemas and confidence |
| Scores | Factor definitions, peer groups, weights and formula versions |
| Risk | Claim verification, guardrails, permissions and escalation rules |
| Learning | Frozen benchmarks, production review queue, drift alerts and a change log |