After Agent 1 was rebuilt from a fiction generator into a real classifier, the critical test was whether the full pipeline could reason correctly over genuinely real, structured input — not whether it could produce a plausible-looking demo.
Client: Crossroads Infrastructure (Construction, $175K retainer, flagged at-risk in the seed portfolio).
Three real events were present in sentinel_raw_events at ready status: a calendar no-show, a manual note from a chance forum encounter with the client contact, and a manual note logging a phone follow-up attempt.
Agent 1 classified all three — correctly identifying each as negative and concerning, with rationales citing the specific evidence ("no-show with no reschedule attempt," "evasive language re: reviewing advisory relationships").
Agent 2 scored the account 0/100, state at_risk. Zero delta from the prior run — consistent, since the account's history was already trending 10 → 0 → 0.
Agent 3 ran (correctly routed, since at_risk requires diagnosis) and classified the root cause as relationship breakdown, citing the specific dated signals rather than inventing a narrative.
Agent 4 drafted two actions: a re-engagement email and call preparation notes, both correctly prefixed with the review-before-sending disclaimer.
Total pipeline time: approximately 33 seconds. state["error"] was None throughout.
This run is the dividing line in the project between a demo that looks intelligent and a system that actually is. Every prior version of this same client, scored against fabricated LLM-generated signals, could produce a similar-looking output — the difference is that this score, this diagnosis, and this drafted email are reasoning about three things that genuinely happened, expressed by a real account manager, resolvable back to source.