Change one operating dimension: such as retrieval depth: without rewriting unrelated parts of the system.
How orthogonal knobs help build reliable agentic harnesses
AI agents can search, plan, invoke tools, execute code, modify files, call APIs, delegate work and operate across long-running tasks. The engineering challenge is no longer simply calling an LLM: it is controlling a system whose behavior emerges from models, prompts, tools, context, memory, data, policy and runtime decisions.
Models, prompts, retrieval, memory, tools, planning, runtime, governance and more.
Assembles context, tools, state, policies, execution environments and stopping conditions.
Traces runs, measures outcomes, compares configurations and creates evidence for governance.
The central idea
A harness gives agents infrastructure. Knobs make that infrastructure controllable.
A production agent can fail because of the model, prompt, retrieval, tool design, memory, planning, routing, runtime limits, policy or even the evaluation criterion. Orthogonal knobs turn those influences into explicit control dimensions that can be configured, observed and tested.
Record the full harness configuration with every run so telemetry can be grouped by knob value and version.
Compare configurations systematically using A/B tests, factorial designs, champion/challenger or sensitivity analysis.
Real agent systems contain interactions. The engineering objective is to isolate dimensions enough that their individual and interaction effects can be measured. A knob should be independently configurable, observable and testable wherever possible.
From harness to control system
Three planes turn agent engineering into a systematic discipline
The harness is most useful when it is separated from the controls that configure it and the measurement system that evaluates it.
Orthogonal Knobs
Control plane. Defines the dimensions that can vary: model, prompt, context, tools, planning, memory, runtime, governance and evaluation.
Agentic Harness
Execution plane. Reads the chosen operating point, assembles the agent and applies context, tools, policies, state and execution boundaries.
Evaluation
Measurement plane. Records outcomes, supports comparison, surfaces regressions and creates evidence for optimization and governance.
A monolithic agent configuration makes root-cause analysis difficult because many influences move together. A knob-based architecture makes the system easier to debug because each run has a clear, versioned operating point.
Knob registry
Treat each major agent dimension as a first-class control
The source article identifies a broad set of knob families. Select one to see how the family affects agent behavior and what could be versioned in the harness.
Model knob
Controls the core reasoning capability available to the agent without forcing changes to retrieval, tools, policies or evaluation.
| Knob family | Example knobs | What it controls |
|---|---|---|
| Model | Provider/model, model size, reasoning tier | Core reasoning capability |
| Prompt | System prompt, task prompt, instructions | Agent behavior |
| Context | Context window, document set, ordering | Information available to the agent |
| Retrieval | Top-K, similarity threshold, reranker | Knowledge selection |
| Memory | Enabled, horizon, summarization strategy | Persistent state |
| Tools | Enabled tools, descriptions, permissions | Agent capabilities |
| Planning | Planner type, decomposition strategy | How work is organized |
| Reasoning | Depth, reflection, retry strategy | How much computation is applied |
| Routing | Specialist selection, escalation thresholds | Which agent handles work |
| Multi-agent | Number of agents, roles, delegation pattern | Collaboration architecture |
| Runtime | Max steps, timeout, token budget | Resource boundaries |
| Sampling | Temperature, seed, stochastic settings | Output variability |
| Governance | Approvals, blocked actions, permissions | Operational control |
| Evaluation | Judge model, rubric, success metrics | How success is measured |
Execution plane
The harness becomes a knob execution engine
Instead of hardcoding behavior throughout application code, the harness reads a configuration that describes the desired operating point, assembles the components and runs the agent.
- Change retrieval top-K without changing the prompt or tool definitions.
- Swap planners while holding model, retrieval and evaluation constant.
- Apply a stricter governance profile without redesigning the agent.
- Version the complete operating point so the run can be compared or reproduced.
model: model-A prompt: prompt-v17 retrieval: strategy: hybrid-v4 top_k: 8 memory: rolling-summary tools: [search, database, calculator] planner: planner-v3 max_steps: 12 policy: financial-high-risk evaluation: quality-rubric-v5
Debugging
Turn traces into causal clues
Agent failures are often stochastic and interaction-driven. Orthogonal knobs make observability more useful because every trace can be associated with the exact operating configuration.
Model = model-A Prompt = v17 Retriever = hybrid-v4 Top-K = 8 Memory = enabled Planner = planner-v2 Max steps = 10 Tool set = finance-v3 Policy = regulated-finance-v2
- Group failures by model, prompt, planner, tool set or policy version.
- Measure interaction effects rather than assuming a single root cause.
- Distinguish behavioral regressions from infrastructure failures.
- Use reproducible configurations to retest suspected causes.
Evaluation → experimentation
Ask which knobs move the outcome: not just which agent version wins
Once the harness exposes knobs and evaluation metrics, agent development can move from informal prompt tweaking to controlled experimentation.
| Experiment | Model | Prompt | Retrieval | Planner | Task success |
|---|---|---|---|---|---|
| A | M1 | P1 | R1 | Plan1 | 82% |
| B | M1 | P2 | R1 | Plan1 | 86% |
| C | M1 | P2 | R2 | Plan1 | 89% |
| D | M1 | P2 | R2 | Plan2 | 91% |
Illustrative experiment from the supplied article. The numbers demonstrate the method, not a vendor benchmark.
A/B testing
Change one knob at a time when you want a clean comparison between a baseline and a candidate.
Factorial experiments
Vary multiple knobs in a structured design to estimate both main effects and interaction effects.
Champion / challenger
Keep a proven harness configuration in production while continuously testing candidate configurations.
Sensitivity analysis
Identify which knobs have the greatest effect on quality, latency, reliability, cost or safety.
Model ≠ system
Separate model capability from system capability
Agent performance is a product of the full system: model × prompt × context × data × tools × planning × memory × governance × runtime.
A larger model may help: but a better tool interface, better retrieval strategy or better routing decision may produce a larger improvement. Orthogonal knobs let teams measure these effects independently instead of treating every agent problem as a reason to buy a larger model.
Model upgrade: 85% → 88% accuracy. Tool-interface improvement: 85% → 93%. The exact values are illustrative; the architectural point is that system-level knobs may matter more than model size for a specific workload.
Risk-adaptive governance
Risk should configure the harness
Different agents can use similar models and tools while receiving radically different autonomy. Governance belongs in the knob system rather than being buried in application code.
Marketing content agent
Writes draft content. human_approval = optional
Customer service agent
May issue account credits. approval_required_above = $100
Financial operations agent
Can move money. human_approval = every_transaction
Do not hardcode risk into one opaque application flow. Represent approvals, permissions, blocked actions, spending limits and escalation thresholds as versioned governance knobs tied to agent, task and risk level.
Reproducibility
Know what configuration caused the behavior
Without configuration discipline, “the agent worked better last week” is nearly impossible to investigate. A knob-based harness gives every execution a complete versioned configuration.
- Version prompts, models, tools, policies and evaluators.
- Give each complete operating point a harness configuration ID.
- Record that ID with every execution trace and result.
- Compare and rerun configurations against the same evaluation dataset.
config_id: H-284 model: model-A@2026-08 prompt: support-v17 tools: service-toolset-v3 retrieval: hybrid-v4 policy: customer-credit-v2 evaluator: task-quality-v5
Evolving harnesses
Make every architectural assumption replaceable
Models improve quickly. Capabilities that once required elaborate orchestration may later be handled directly by a stronger model. If planning, memory and orchestration are separate knobs, the harness can evolve without a rewrite.
Example: simplify planning
planner: complex-tree-planner
planner: native-model-planning
Compare on the same evaluation set
- Quality and task completion
- Latency and token usage
- Reliability and recovery
- Cost per successful task
- Governance and policy compliance
If the simpler operating point performs equally well, the old planner can be retired. Knobs give teams a disciplined way to simplify as models improve.
Long-running agents
The longer the task, the more valuable the knob model becomes
Coding, research and business-process agents may work across many files, searches, context windows or hours of execution. State management and continuation strategies become part of the harness: and therefore part of the knob registry.
Checkpoint frequency, continuation policy, artifact persistence and recovery behavior.
Context compression, reset strategy, memory horizon and summarization approach.
Retry count, maximum runtime, maximum cost and human-escalation thresholds.
Harness as a knob
Even the orchestration architecture can be experimental
Teams do not need to argue abstractly about whether multi-agent systems are “better.” Treat the harness pattern itself as a controlled dimension and measure it.
Closed-loop agent engineering
Orthogonal knobs create an agent control plane
The larger idea is not “add configuration options.” It is to create a control plane where the organization knows what can vary, how the agent runs, what happened, how well it worked and which configurations are allowed.
Implementation blueprint
A practical way to introduce knob-based harnesses
The following implementation path synthesizes the architecture described in the supplied article into a build sequence for engineering teams.
Create a knob registry
Define the dimensions that can vary, their allowed values, owners and expected effects.
Externalize configuration
Move model, prompt, retrieval, tool, runtime and policy choices out of hardcoded orchestration logic.
Version everything
Version prompts, tools, policies, planners, evaluators and complete harness operating points.
Trace every run
Record the configuration ID, tool path, runtime events, outputs and evaluation results.
Define evaluators
Measure quality, correctness, safety, latency, cost, task success and policy compliance.
Run controlled experiments
Use A/B, factorial or champion/challenger designs to learn which knobs actually matter.
Govern allowed operating points
Attach risk tiers, approvals, permissions and blocked actions to configurations: not just code paths.
Continuously simplify
As models improve, retest whether planners, memories, retries or multi-agent patterns are still necessary.
Frequently asked questions
Orthogonal knobs and agentic harnesses
From prompt engineering to agent engineering
A production agent is not just a prompt wrapped around a model. It is a system of models, context, tools, memory, orchestration, policies, environments and evaluation. Orthogonal knobs provide the abstraction needed to control that complexity.
The objective is not to restrict intelligence: it is to make increasingly powerful intelligence controllable.
Content is grounded in the supplied “How Orthogonal Knobs Help Build Reliable Agentic Harnesses” article. The implementation blueprint is a structured synthesis of the article’s architecture and recommendations.