Evaluation method

How we evaluate coaching over time.

The Coaching Arc Lab gives Nexus Labs a stable set of fictional lives against which product behavior can be examined, compared, and revised. Its purpose is to expose failures that only become visible across a developing coaching relationship.

The unit of evaluation is a changing life

Each case begins with a deeply specified person: his daily life, social environment, important relationships, history, goals, patterns, speech, and the information he is not ready to share. The intended direction of the case is defined, but the individual sessions are not scripted in advance.

The current corpus contains six fictional cases and 51 fixed voice sessions. The longest case spans 182 days of simulated time. More than 200 alternative session transcripts are retained from the process used to construct the final case histories.

Cases advance one session at a time

For each session, the person's private situation and likely behavior are specified. Multiple conversation candidates are generated under the same case conditions. A human reviewer selects the session that best preserves the person, the coaching relationship, and the useful ambiguity of a real conversation.

Once selected, that session becomes fixed case evidence. Later sessions must respect what was said, what remained private, which real-world actions occurred, and what the coach could reasonably know at the time. The history is not rewritten to make a later product version look better.

What the lab is used to evaluate

The fixed cases can be supplied to the systems that support coaching. The first major use was evaluating the structured continuity system that carries relevant context from one session into the next. The same case states can also support evaluation of post-session summaries, chat, outreach, and other components.

A component receives the history and context it would receive in the product. Its output is then reviewed against the person's established life and the coaching method. This allows the evaluation to ask a concrete question: would this output make the next coaching interaction more accurate and useful?

Evaluation criteria

Fidelity to the person

The system must preserve what was actually said, distinguish fact from interpretation, and avoid inventing motives, diagnoses, or events.

Longitudinal continuity

Important people, commitments, unfinished work, changing priorities, and earlier disclosures must remain available when they matter.

Coaching judgment

The output should support the right work for the current barrier and avoid generic advice, premature action, false certainty, or automatic agreement.

Proportion and restraint

A minor event should not rewrite the entire understanding of the person. One successful action should not erase a longstanding pattern.

Coaching boundaries

Practice must remain distinct from real-world completion, uncertainty must remain visible, and clinical claims must not be manufactured from ordinary life material.

How a change is reviewed

Evaluation begins with an observable failure. Examples include a missed disclosure, invented certainty, a stale formulation, an inappropriate Challenge, or a summary that overstates progress. The proposed revision is then run on the same input that exposed the problem.

The comparison does not end with the target case. The revised component is also run against other relevant case states, especially cases where the new instruction could create the opposite error. Individual outputs are critiqued, results are reviewed across the set, and the change is accepted, revised, rejected, or placed under watch.

Generated evidence and human judgment

The people and conversations in the Arc Lab are synthetic. Models play both the coach and the fictional user during transcript generation, and several plausible alternatives are retained. Human judgment determines which session becomes part of the fixed case and how product outputs are ultimately accepted.

This arrangement makes controlled comparison possible, but it also has a known limitation: generated users, generated coaches, and model critics can share assumptions. We account for that by preserving alternatives, reviewing complete histories, making acceptance decisions explicitly, and testing live voice behavior through a separate process.

Live voice requires a different test

The Arc Lab is suited to questions about understanding, continuity, pacing, intervention choice, and boundaries across time. It cannot reproduce speech recognition, interruption, latency, turn-taking, silence, audio quality, or transitions into deeper reasoning.

Those behaviors are evaluated in live sessions using the actual product configuration. Findings from live testing can lead to changes in the voice instructions, provider integration, turn control, or the way product systems appear inside the conversation.

Scope of the method

The Coaching Arc Lab is an internal product-evaluation system. It shows how a specified component behaves against controlled fictional evidence and whether a revision improves that behavior under the conditions examined.

It is not a clinical trial, a study of customer outcomes, or a guarantee that every live conversation will meet the standard. Those questions require separate forms of evidence.