What the Design Earned, What It Couldn't Reach, and Where It Goes Next


Property Value
๐Ÿ“… Project OmniRisk ยท Financial Workflow Dashboard
๐Ÿ”– Phase Outcome
๐ŸŽฏ Purpose Honest reflection on what worked, what didn't, and what comes next

Outcome at a Glance

OmniRisk validated the trust thesis at the interaction layer. Three operator personas navigated the sequential workflow without instruction. The AI Recommendation modal earned trust ratings of 4.3/5 in post-task survey. Manual Override read as a first-class action and produced the MiFID II audit trail the hedge-fund analyst usually builds manually in Excel. The two-speed design (badges for triage, modal for depth) worked for both the middle office operator and the front office trader.

Three findings did not hold on first contact and were addressed โ€” one iteration shipped within the project window (the phase stepper), two findings were scoped for next iteration with designed responses (confidence band legend, "See AI reasoning" relabel). The design's interaction layer behaved as the research and synthesis predicted.

The case study also surfaced what testing could not reach โ€” behavior under live volatile market conditions, compliance audit adequacy from a regulatory reviewer, longitudinal adoption patterns, and cross-workflow generalization. These are the boundaries of what a conceptual case study can claim. They are documented honestly, not hidden.


What Worked

Sequential CTAs read as orchestration, not as constraint. All three operators navigated the Fetch FCM โ†’ Run Analysis โ†’ View Insights sequence without instruction. The hedge-fund analyst explicitly preferred the enforced order, citing a past incident where out-of-sequence data processing caused a major reconciliation break at his previous fund. Decision 02 (Enforce the Sequence) held because it reflected real workflow dependencies โ€” not because it was imposed by the design.

The two-speed pattern served both personas without compromise. Compact badges in the Currency Trade Table gave the FX trader the triage speed her work demands. The AI Recommendation modal gave the fintech-side analyst the reasoning depth her trust threshold requires. The same design served both โ€” without forcing either into the other's view. Decision 01 (Confidence in the Modal) and in-table highlighting (Core feature) composed into a pattern that worked for diverging user needs.

The override compliance flow worked as a designed experience, not as paperwork. The hedge-fund analyst completed the Manual Override flow during his session โ€” override notice, required reason, three risk acknowledgments, audit trail block with MiFID II / Dodd-Frank specifics. He read the compliance block carefully. His post-task feedback specifically noted that the override flow generates the audit trail he otherwise builds manually in Excel later. Decision 03 (Override Sits Beside Accept) โ€” the most sophisticated decision in the case study โ€” held both perceptually and functionally.

Trust signals composed. The trust thesis predicted that explainability, controllability, and auditability would compose into AI adoption. The sessions confirmed this at the interaction layer: the fintech-side analyst verified before accepting even at high confidence; the FX trader paused at medium confidence to verify against external data; the hedge-fund analyst overrode when he had information the AI didn't. All three behaviors are exactly what the trust thesis predicted skeptical operators would do with a design that meets the conditions.


What Didn't Work

Three earlier-round findings did not hold on first contact. The intermediate state hesitation, the confidence score calibration drift, and the "More Details" underuse were real failures of the initial design to communicate what it needed to communicate. One was resolved with the phase stepper iteration; the other two are scoped for next iteration with designed responses. The case study does not claim everything worked โ€” and the gap between hypothesis and observation is where most of the design learning lives.

The testing depth was scoped to what the project window could support. Three sessions at 30-45 minutes each, focused on confirming interaction-layer assumptions, is appropriate methodology for a conceptual case study at this stage. It is not the methodology that would be appropriate for a production validation program before launch. The honest position is that the testing answered the questions it was designed to answer and did not attempt to answer questions it was not equipped to handle.

The audit trail's regulatory adequacy was not validated by a compliance reviewer. The Manual Override modal's compliance documentation flow is designed against MiFID II and Dodd-Frank requirements based on operations experience, but no compliance officer reviewed it for audit-review adequacy. This is a production-testing concern, not a research gap to hide.