Orthogonal Knobs + Agentic Harnesses
Agent engineering architecture

How orthogonal knobs help build reliable agentic harnesses

AI agents can search, plan, invoke tools, execute code, modify files, call APIs, delegate work and operate across long-running tasks. The engineering challenge is no longer simply calling an LLM: it is controlling a system whose behavior emerges from models, prompts, tools, context, memory, data, policy and runtime decisions.

Control planeWHAT CAN VARY
Orthogonal Knobs

Models, prompts, retrieval, memory, tools, planning, runtime, governance and more.

Execution planeHOW IT RUNS
Agentic Harness

Assembles context, tools, state, policies, execution environments and stopping conditions.

Measurement planeWHAT HAPPENED
Observability + Evaluation

Traces runs, measures outcomes, compares configurations and creates evidence for governance.

The central idea

A harness gives agents infrastructure. Knobs make that infrastructure controllable.

A production agent can fail because of the model, prompt, retrieval, tool design, memory, planning, routing, runtime limits, policy or even the evaluation criterion. Orthogonal knobs turn those influences into explicit control dimensions that can be configured, observed and tested.

K
Configurable

Change one operating dimension: such as retrieval depth: without rewriting unrelated parts of the system.

O
Observable

Record the full harness configuration with every run so telemetry can be grouped by knob value and version.

E
Experimentable

Compare configurations systematically using A/B tests, factorial designs, champion/challenger or sensitivity analysis.

Orthogonal does not mean perfectly independent.

Real agent systems contain interactions. The engineering objective is to isolate dimensions enough that their individual and interaction effects can be measured. A knob should be independently configurable, observable and testable wherever possible.

From harness to control system

Three planes turn agent engineering into a systematic discipline

The harness is most useful when it is separated from the controls that configure it and the measurement system that evaluates it.

Agent outcome = f(K₁, K₂, K₃ … Kₙ)

Orthogonal Knobs

Control plane. Defines the dimensions that can vary: model, prompt, context, tools, planning, memory, runtime, governance and evaluation.

Agentic Harness

Execution plane. Reads the chosen operating point, assembles the agent and applies context, tools, policies, state and execution boundaries.

Evaluation

Measurement plane. Records outcomes, supports comparison, surfaces regressions and creates evidence for optimization and governance.

Why this matters:

A monolithic agent configuration makes root-cause analysis difficult because many influences move together. A knob-based architecture makes the system easier to debug because each run has a clear, versioned operating point.

Knob registry

Treat each major agent dimension as a first-class control

The source article identifies a broad set of knob families. Select one to see how the family affects agent behavior and what could be versioned in the harness.

Model knob

Controls the core reasoning capability available to the agent without forcing changes to retrieval, tools, policies or evaluation.

Knob familyExample knobsWhat it controls
ModelProvider/model, model size, reasoning tierCore reasoning capability
PromptSystem prompt, task prompt, instructionsAgent behavior
ContextContext window, document set, orderingInformation available to the agent
RetrievalTop-K, similarity threshold, rerankerKnowledge selection
MemoryEnabled, horizon, summarization strategyPersistent state
ToolsEnabled tools, descriptions, permissionsAgent capabilities
PlanningPlanner type, decomposition strategyHow work is organized
ReasoningDepth, reflection, retry strategyHow much computation is applied
RoutingSpecialist selection, escalation thresholdsWhich agent handles work
Multi-agentNumber of agents, roles, delegation patternCollaboration architecture
RuntimeMax steps, timeout, token budgetResource boundaries
SamplingTemperature, seed, stochastic settingsOutput variability
GovernanceApprovals, blocked actions, permissionsOperational control
EvaluationJudge model, rubric, success metricsHow success is measured

Execution plane

The harness becomes a knob execution engine

Instead of hardcoding behavior throughout application code, the harness reads a configuration that describes the desired operating point, assembles the components and runs the agent.

  • Change retrieval top-K without changing the prompt or tool definitions.
  • Swap planners while holding model, retrieval and evaluation constant.
  • Apply a stricter governance profile without redesigning the agent.
  • Version the complete operating point so the run can be compared or reproduced.
harness-config.yaml
model: model-A
prompt: prompt-v17
retrieval:
  strategy: hybrid-v4
  top_k: 8
memory: rolling-summary
tools: [search, database, calculator]
planner: planner-v3
max_steps: 12
policy: financial-high-risk
evaluation: quality-rubric-v5

Debugging

Turn traces into causal clues

Agent failures are often stochastic and interaction-driven. Orthogonal knobs make observability more useful because every trace can be associated with the exact operating configuration.

Agent Run #48312trace snapshot
Model       = model-A
Prompt      = v17
Retriever   = hybrid-v4
Top-K       = 8
Memory      = enabled
Planner     = planner-v2
Max steps   = 10
Tool set    = finance-v3
Policy      = regulated-finance-v2
Example diagnostic insight: planner-v2 shows a higher failure rate when retrieval top-K is large. The point is not the specific percentage; the point is that telemetry can be grouped by knob values to reveal interaction patterns that a generic “agent failed” log cannot.
  • Group failures by model, prompt, planner, tool set or policy version.
  • Measure interaction effects rather than assuming a single root cause.
  • Distinguish behavioral regressions from infrastructure failures.
  • Use reproducible configurations to retest suspected causes.

Evaluation → experimentation

Ask which knobs move the outcome: not just which agent version wins

Once the harness exposes knobs and evaluation metrics, agent development can move from informal prompt tweaking to controlled experimentation.

ExperimentModelPromptRetrievalPlannerTask success
AM1P1R1Plan182%
BM1P2R1Plan186%
CM1P2R2Plan189%
DM1P2R2Plan291%

Illustrative experiment from the supplied article. The numbers demonstrate the method, not a vendor benchmark.

A/B testing

Change one knob at a time when you want a clean comparison between a baseline and a candidate.

Factorial experiments

Vary multiple knobs in a structured design to estimate both main effects and interaction effects.

Champion / challenger

Keep a proven harness configuration in production while continuously testing candidate configurations.

Sensitivity analysis

Identify which knobs have the greatest effect on quality, latency, reliability, cost or safety.

Model ≠ system

Separate model capability from system capability

Agent performance is a product of the full system: model × prompt × context × data × tools × planning × memory × governance × runtime.

A larger model may help: but a better tool interface, better retrieval strategy or better routing decision may produce a larger improvement. Orthogonal knobs let teams measure these effects independently instead of treating every agent problem as a reason to buy a larger model.

Illustrative comparison from the source article

Model upgrade: 85% → 88% accuracy. Tool-interface improvement: 85% → 93%. The exact values are illustrative; the architectural point is that system-level knobs may matter more than model size for a specific workload.

Risk-adaptive governance

Risk should configure the harness

Different agents can use similar models and tools while receiving radically different autonomy. Governance belongs in the knob system rather than being buried in application code.

Lower consequence

Marketing content agent

Writes draft content. human_approval = optional

Moderate consequence

Customer service agent

May issue account credits. approval_required_above = $100

High consequence

Financial operations agent

Can move money. human_approval = every_transaction

Architecture principle:

Do not hardcode risk into one opaque application flow. Represent approvals, permissions, blocked actions, spending limits and escalation thresholds as versioned governance knobs tied to agent, task and risk level.

Reproducibility

Know what configuration caused the behavior

Without configuration discipline, “the agent worked better last week” is nearly impossible to investigate. A knob-based harness gives every execution a complete versioned configuration.

  • Version prompts, models, tools, policies and evaluators.
  • Give each complete operating point a harness configuration ID.
  • Record that ID with every execution trace and result.
  • Compare and rerun configurations against the same evaluation dataset.
Harness Configuration H-284reproducible run
config_id: H-284
model: model-A@2026-08
prompt: support-v17
tools: service-toolset-v3
retrieval: hybrid-v4
policy: customer-credit-v2
evaluator: task-quality-v5

Evolving harnesses

Make every architectural assumption replaceable

Models improve quickly. Capabilities that once required elaborate orchestration may later be handled directly by a stronger model. If planning, memory and orchestration are separate knobs, the harness can evolve without a rewrite.

Example: simplify planning

beforecomplex harness
planner: complex-tree-planner
candidatesimpler harness
planner: native-model-planning

Compare on the same evaluation set

  • Quality and task completion
  • Latency and token usage
  • Reliability and recovery
  • Cost per successful task
  • Governance and policy compliance

If the simpler operating point performs equally well, the old planner can be retired. Knobs give teams a disciplined way to simplify as models improve.

Long-running agents

The longer the task, the more valuable the knob model becomes

Coding, research and business-process agents may work across many files, searches, context windows or hours of execution. State management and continuation strategies become part of the harness: and therefore part of the knob registry.

State & checkpoint knobs

Checkpoint frequency, continuation policy, artifact persistence and recovery behavior.

Context knobs

Context compression, reset strategy, memory horizon and summarization approach.

Runtime knobs

Retry count, maximum runtime, maximum cost and human-escalation thresholds.

Harness as a knob

Even the orchestration architecture can be experimental

Teams do not need to argue abstractly about whether multi-agent systems are “better.” Treat the harness pattern itself as a controlled dimension and measure it.

Single agentOne agent owns the task end to end.
Planner → WorkerPlanning is separated from task execution.
Orchestrator → WorkersA coordinator decomposes and dispatches work.
Generator → Evaluator → GeneratorGeneration and critique form an iterative loop.
Router → SpecialistsTasks are assigned to specialized agents based on routing logic.

Closed-loop agent engineering

Orthogonal knobs create an agent control plane

The larger idea is not “add configuration options.” It is to create a control plane where the organization knows what can vary, how the agent runs, what happened, how well it worked and which configurations are allowed.

Configure
↓
Execute
↓
Observe
↓
Evaluate
↓
Compare
↓
Govern
↓
Optimize

Implementation blueprint

A practical way to introduce knob-based harnesses

The following implementation path synthesizes the architecture described in the supplied article into a build sequence for engineering teams.

01

Create a knob registry

Define the dimensions that can vary, their allowed values, owners and expected effects.

02

Externalize configuration

Move model, prompt, retrieval, tool, runtime and policy choices out of hardcoded orchestration logic.

03

Version everything

Version prompts, tools, policies, planners, evaluators and complete harness operating points.

04

Trace every run

Record the configuration ID, tool path, runtime events, outputs and evaluation results.

05

Define evaluators

Measure quality, correctness, safety, latency, cost, task success and policy compliance.

06

Run controlled experiments

Use A/B, factorial or champion/challenger designs to learn which knobs actually matter.

07

Govern allowed operating points

Attach risk tiers, approvals, permissions and blocked actions to configurations: not just code paths.

08

Continuously simplify

As models improve, retest whether planners, memories, retries or multi-agent patterns are still necessary.

Frequently asked questions

Orthogonal knobs and agentic harnesses

An agentic harness is the infrastructure surrounding an AI agent that manages context, tools, environments, state, delegation, execution observation and termination. The model supplies intelligence; the harness supplies the operating structure around that intelligence.
An orthogonal knob is a configurable dimension designed so that it can be changed and evaluated largely independently from other dimensions. Model selection, retrieval depth, tool availability and maximum agent steps are examples.
A configuration file stores settings. A knob system treats settings as experimental and governance dimensions. Values are versioned, traced, compared, evaluated and potentially optimized.
Yes. Once knobs and metrics are formally defined, experimentation systems can search across configurations using A/B testing, factorial experiments, bandit methods, Bayesian optimization or other optimization techniques.
No. Some interactions are unavoidable. Model choice may interact with prompts, tools or context strategy. Orthogonality is primarily a design objective: isolate dimensions sufficiently that individual and interaction effects can be measured.
Governance policies become explicit controls. Permissions, human approvals, tool restrictions, spending limits, data access and escalation rules can be attached to agents, tasks and risk levels and recorded with each run.

From prompt engineering to agent engineering

A production agent is not just a prompt wrapped around a model. It is a system of models, context, tools, memory, orchestration, policies, environments and evaluation. Orthogonal knobs provide the abstraction needed to control that complexity.

Orthogonal Knobsdefine the controls.
Agentic Harnessapplies the controls.
Observabilityrecords what happened.
Evaluationmeasures the outcome.
Experimentationdiscovers what works.
Governancedetermines what is allowed.

The objective is not to restrict intelligence: it is to make increasingly powerful intelligence controllable.

Content is grounded in the supplied “How Orthogonal Knobs Help Build Reliable Agentic Harnesses” article. The implementation blueprint is a structured synthesis of the article’s architecture and recommendations.