The control pattern
Data asset
User reactions and expert corrections
Knobs
Feedback type, Reviewer reliability, Weighting, Recency, Context, Aggregation
Outcome
Behaviour that improves with real-world use
Measurement
See how to measure this below
What it is
Feedback learning uses reactions to model outputs as data. Explicit signals include ratings, corrections and written comments; implicit signals include whether a user opens a source, edits a watchlist, asks a follow-up, or abandons a workflow.
These signals are not equally trustworthy. Feedback type, reviewer expertise, context, freshness and weighting are knobs that determine how much influence each observation receives. A click may indicate curiosity rather than agreement, and high engagement does not necessarily mean high-quality financial guidance.
Feedback should be linked to the exact model, prompt, data snapshot and evidence bundle that produced the answer, so the issue can be reproduced.
The knobs in detail
Each row is one adjustable property of the data asset, and what moving it tends to do.
| Knob | What you adjust | Likely effect |
|---|---|---|
| Feedback type | Explicit rating vs. implicit behaviour | Different signals mean different things |
| Reviewer reliability | Expert vs. general user | Weights factual corrections appropriately |
| Weighting | Influence per observation | Stops loud minorities dominating |
| Recency | How current the signal is | Keeps stale complaints from steering the roadmap |
| Context | What produced the answer | Makes issues reproducible |
| Aggregation | How signals combine | Detects campaigns and duplication |
Applied: Stocks Assistant
Stocks Assistant can ask targeted questions instead of relying only on a generic thumbs-up control. Users might flag an incorrect number, a stale source, a missing risk, a confusing explanation, or an irrelevant recommendation. Each category routes to a different fix. Domain-expert corrections should carry more weight for factual labels, while broad user feedback is valuable for clarity and workflow design.
How to measure it
Evidence that the knob produced the intended behaviour, rather than shifting the problem elsewhere.
- Whether feedback-driven changes improve a held-out benchmark
- Long-term trust metrics rather than short-term clicks
- Recurrence of the specific issue that was reported
- Rate of duplicated or abusive feedback detected
Common mistakes
Training directly on raw feedback without deduplication, abuse detection and verification.
Optimising for trading frequency or emotional engagement rather than informed decisions.