THE TRIBUNAL
Last Tuesday, an AI agent released a $312,000 wire. Its own retrieval had flagged it 74% fraud. Nobody validated the reasoning - because your team has no playbook for agents that plan, act, and self-correct. This is that playbook.
Your validators can validate a credit model.
They have no playbook for an agent.
SR 11-7 - spelled out, "S-R eleven dash seven" - is the Fed's model risk guidance. It assumes a model you can inspect: fixed inputs, fixed logic, a documented output. An agent breaks every assumption. It plans. It calls tools. It changes its mind. And it writes its own justification after it acts.
Four questions a tribunal asks.
Every one maps to an SR 11-7 pillar.
Adversarial validation isn't hostility for its own sake. It's a disciplined interrogation along four axes. Tap each to see what the examiner is really testing.
In re: the $312,000 wire.
You are the examiner. Rule on each step.
Here is the fraud-triage agent's full decision trace. Read each step and rule - reasoning holds or reasoning breaks - before you reveal the tribunal's finding. This is the skill: reading a trace adversarially.
The tribunal's findings.
Verdict: FAIL.
Three of your five rulings should have been "breaks." Here is why - each finding mapped to its SR 11-7 pillar, with severity and the remedy that closes it.
Detection isn't the product.
Closure is.
A FAIL verdict is a finding, not an outcome. The tribunal's value is the loop: apply the controls, re-convene, and rule on residual risk only. Apply each remedy and watch the verdict journey.
Checkpoint one.
Reading the trace.
Two questions on what the examination just taught. Pick an answer - the reasoning comes either way.
Now build your own tribunal.
In re: the credit-limit agent.
A different agent, a different trace: it approved an $8,000 → $25,000 limit increase. Assemble the adversarial exam - pick the challenges that will actually test it. Choose well; two of these are traps that a real examiner would reject.
Who validates the validator?
Audit the examiner's own reasoning.
An AI examiner is a model too - and effective challenge applies to it. Here are three of the tribunal's own rulings on the credit case. One over-flagged a sound step. One missed a real problem. One holds. Judge the judge.
Checkpoint two.
The discipline.
The last two questions of the program.
You've learned to run the tribunal.
Now here's the one that runs at scale.
Six of seven modules complete. You can brief a model, countersign its output, keep the record, operate the agent, record the flight - and now, put it on trial.
TRIBUNAL is the instrument that runs this at scale: it cross-examines an agent's decision trace like a hostile examiner, generates SR 11-7 evidence with a hash-chained audit trail, and documents the remediation loop from FAIL to PASS. Find out if you'd pass the bank exam - before the bank runs it.