AI PRODUCT CASE STUDY2026 · PERSONAL PRODUCT EXPLORATION

Building AI workflows teams can trust.

How I designed Nous as an operating layer for grounded, governed agent workflows—and built evaluation into the product rather than treating it as a final quality check.

TRUSTED CONTEXTEvidence
→
AGENT LOOPPlan · Act · Observe
→
CONTROL POINTHuman approval
→
LEARNING LOOPEvaluate · Improve

THE PROBLEM

AI capability is moving faster than teams’ ability to deploy it responsibly.

Teams do not struggle to find a chatbot. They struggle to identify valuable workflows, connect trustworthy context, decide how much autonomy is appropriate, and know whether an agent is reliable enough to use.

The product opportunity was to make those operational decisions visible and manageable in one system.

PRODUCT THESIS

Start with governed workflows. Earn autonomy through evidence.

01

Workflow before agent theatre

Use predictable paths for well-defined work. Add model-directed decisions only where flexibility creates meaningful value.

02

Context is a permission

Each run receives selected, classified sources and the minimum passages necessary—not unrestricted access to team knowledge.

03

Approval must be meaningful

Show the plan, evidence, assumptions and intended action so a person can make an informed decision rather than rubber-stamp a result.

04

Reliability is a product surface

Evaluation suites, traces, failure clusters and release gates belong inside the operating experience.

WHAT EXISTS TODAY

A working trust foundation.

Nous already supports reusable workflows, selected trusted sources, source classifications, verified evidence, approval ownership, local fallback, runtime limits and an audit trail.

AI CONTROL CENTEREvidence engine active
Human approvalAlways on
Sensitive-data blockingAlways on
External actionsNone permitted
Source accessSelected only
Every run is attributable and reviewable.✓

ASSURANCE PROGRAMME

Evaluate the whole system—not just its final answer.

The evaluation design treats the model, harness, tools and environment as one product. A plausible output cannot hide a poor retrieval decision, unsafe tool call or broken approval boundary.

01

Components

Parsing, retrieval, citation matching, PII detection, schema and policy enforcement.

02

Grounding

Context precision and recall, source support, contradiction and minority-view preservation.

03

Agent loop

Planning, tool selection, recovery, clarification, stopping behaviour, budgets and goal drift.

04

Outcome

Theme quality, severity, prioritisation, actionability and human acceptance.

05

Safety

Injection, privilege, memory, data leakage, approval bypass and resource-exhaustion tests.

06

Human value

Correction effort, reviewer trust, decision quality and time to an acceptable result.

TRAJECTORY EVALUATION

Inspect every turn of the agent loop.

A trace records the goal, plan, retrieved evidence, tool calls, observations, re-planning, approval pauses and stopping decision. Each transition can be graded and replayed.

  1. 1PlanIs the approach appropriate?
  2. 2ActRight tool and arguments?
  3. 3ObserveDid it interpret state correctly?
  4. 4AdaptRecover, clarify or stop?
  5. 5ApproveWas control returned in time?

CONTROLLED BENCHMARK

155 synthetic cases with known intent and expected behaviour.

No customer or production data is used. The benchmark is designed to measure system behaviour safely before seeking a permissioned real-data pilot.

60

Golden quality cases

Known themes, severity, contradictions, duplicates, noise and insufficient evidence.

30

Agent-loop scenarios

Tool failures, ambiguous intent, changing evidence, rejected approvals and exhausted budgets.

40

Adversarial cases

Direct and indirect injection, poisoned context, PII, exfiltration and permission attacks.

25

Holdout cases

Unseen release-checkpoint tests, isolated from prompt and workflow development.

DEFENCE IN DEPTH

Guardrails at every boundary.

INPUT

Injection, malformed content, PII and prohibited intent.

CONTEXT

Poisoned sources, hidden instructions and classification boundaries.

PLAN

Goal drift, excessive autonomy and policy evasion.

TOOLS

Misuse, invalid arguments, privilege and duplicate actions.

OUTPUT

Unsupported claims, leakage, unsafe recommendations and uncertainty.

RUNTIME

Infinite loops, budgets, timeouts and resource exhaustion.

MIXED GRADING

Use the most reliable evaluator for each question.

Deterministic

Schema, citations, counts, PII, tool permissions and policy assertions.

Fast · objective · repeatable

Model-based

Theme coverage, nuance, contradiction handling and recommendation quality.

Rubric-led · evidence required

Human

Usefulness, trust, correction effort and whether the output should be approved.

Calibrates every automated grader

Model graders are versioned, tested against deliberately good and bad examples, compared with blinded human labels, and reviewed where they disagree.

PROPOSED RELEASE GATES

Autonomy is earned, not assumed.

These are initial product targets for the controlled benchmark—not industry standards and not measured results. Thresholds will be revised after grader calibration.

Citation correctness≥ 98%
Unsupported claims≤ 2%
Theme recall≥ 85%
Critical policy tests100%
Approval bypass0
PII leakage0
Loop completion≥ 95%
Tool recovery≥ 90%
Human acceptance≥ 80%

CONTINUOUS EVALUATION

Every failure should make the system harder to break twice.

Production trace→Automatic checks→Failure cluster→Human review→Regression case→Canary release

CURRENT STATUS

Evaluation infrastructure in progress.

The trust controls and trace foundation are working in the live demo. The benchmark datasets, calibrated graders and measured V1–V2 results are the next build phase.

  • ✓ Product thesis and workflow model
  • ✓ Grounded run trace and approval controls
  • ✓ Evaluation architecture and release criteria
  • ○ Synthetic benchmark execution
  • ○ Independent PM review
  • ○ Permissioned real-data pilot

MY CONTRIBUTION

Strategy through implementation.

Product strategyAI system designEvaluation designUX and prototypingGuardrails and governanceHands-on implementation

EXPLORE THE PRODUCT

See the operating model in action.

Open Nous ↗All case studies →