Building AI workflows teams can trust.
How I designed Nous as an operating layer for grounded, governed agent workflows—and built evaluation into the product rather than treating it as a final quality check.
THE PROBLEM
AI capability is moving faster than teams’ ability to deploy it responsibly.
Teams do not struggle to find a chatbot. They struggle to identify valuable workflows, connect trustworthy context, decide how much autonomy is appropriate, and know whether an agent is reliable enough to use.
The product opportunity was to make those operational decisions visible and manageable in one system.
PRODUCT THESIS
Start with governed workflows. Earn autonomy through evidence.
Workflow before agent theatre
Use predictable paths for well-defined work. Add model-directed decisions only where flexibility creates meaningful value.
Context is a permission
Each run receives selected, classified sources and the minimum passages necessary—not unrestricted access to team knowledge.
Approval must be meaningful
Show the plan, evidence, assumptions and intended action so a person can make an informed decision rather than rubber-stamp a result.
Reliability is a product surface
Evaluation suites, traces, failure clusters and release gates belong inside the operating experience.
WHAT EXISTS TODAY
A working trust foundation.
Nous already supports reusable workflows, selected trusted sources, source classifications, verified evidence, approval ownership, local fallback, runtime limits and an audit trail.
ASSURANCE PROGRAMME
Evaluate the whole system—not just its final answer.
The evaluation design treats the model, harness, tools and environment as one product. A plausible output cannot hide a poor retrieval decision, unsafe tool call or broken approval boundary.
Components
Parsing, retrieval, citation matching, PII detection, schema and policy enforcement.
Grounding
Context precision and recall, source support, contradiction and minority-view preservation.
Agent loop
Planning, tool selection, recovery, clarification, stopping behaviour, budgets and goal drift.
Outcome
Theme quality, severity, prioritisation, actionability and human acceptance.
Safety
Injection, privilege, memory, data leakage, approval bypass and resource-exhaustion tests.
Human value
Correction effort, reviewer trust, decision quality and time to an acceptable result.
TRAJECTORY EVALUATION
Inspect every turn of the agent loop.
A trace records the goal, plan, retrieved evidence, tool calls, observations, re-planning, approval pauses and stopping decision. Each transition can be graded and replayed.
- 1PlanIs the approach appropriate?
- 2ActRight tool and arguments?
- 3ObserveDid it interpret state correctly?
- 4AdaptRecover, clarify or stop?
- 5ApproveWas control returned in time?
CONTROLLED BENCHMARK
155 synthetic cases with known intent and expected behaviour.
No customer or production data is used. The benchmark is designed to measure system behaviour safely before seeking a permissioned real-data pilot.
Golden quality cases
Known themes, severity, contradictions, duplicates, noise and insufficient evidence.
Agent-loop scenarios
Tool failures, ambiguous intent, changing evidence, rejected approvals and exhausted budgets.
Adversarial cases
Direct and indirect injection, poisoned context, PII, exfiltration and permission attacks.
Holdout cases
Unseen release-checkpoint tests, isolated from prompt and workflow development.
DEFENCE IN DEPTH
Guardrails at every boundary.
Injection, malformed content, PII and prohibited intent.
Poisoned sources, hidden instructions and classification boundaries.
Goal drift, excessive autonomy and policy evasion.
Misuse, invalid arguments, privilege and duplicate actions.
Unsupported claims, leakage, unsafe recommendations and uncertainty.
Infinite loops, budgets, timeouts and resource exhaustion.
MIXED GRADING
Use the most reliable evaluator for each question.
Deterministic
Schema, citations, counts, PII, tool permissions and policy assertions.
Fast · objective · repeatableModel-based
Theme coverage, nuance, contradiction handling and recommendation quality.
Rubric-led · evidence requiredHuman
Usefulness, trust, correction effort and whether the output should be approved.
Calibrates every automated graderModel graders are versioned, tested against deliberately good and bad examples, compared with blinded human labels, and reviewed where they disagree.
PROPOSED RELEASE GATES
Autonomy is earned, not assumed.
These are initial product targets for the controlled benchmark—not industry standards and not measured results. Thresholds will be revised after grader calibration.
CONTINUOUS EVALUATION
Every failure should make the system harder to break twice.
CURRENT STATUS
Evaluation infrastructure in progress.
The trust controls and trace foundation are working in the live demo. The benchmark datasets, calibrated graders and measured V1–V2 results are the next build phase.
- ✓ Product thesis and workflow model
- ✓ Grounded run trace and approval controls
- ✓ Evaluation architecture and release criteria
- ○ Synthetic benchmark execution
- ○ Independent PM review
- ○ Permissioned real-data pilot
MY CONTRIBUTION
Strategy through implementation.
EXPLORE THE PRODUCT