
There is a predictable moment in these conversations. The demonstration has gone well, and then somebody whose job is independent review asks a question that sounds almost too basic to have been held until the end. From there the meeting either becomes a real project or a follow-up nobody schedules. We have sat through it often enough to be direct about the question and why it is fair.
The question that stalls the conversation
The first question model-risk teams in banking ask about generative AI is deceptively simple: can you explain how the model produced a specific output?
Not how the model works in general. That output — the summary it produced, the recommendation it surfaced on that file, on that date — and the path that led to it.
For a traditional credit model or a fraud detection system, that question has a settled answer. You show feature weights, decision trees, or logistic regression coefficients. The explanation is a property of the model itself, and it can be handed over. Somebody can read it without you in the room, disagree with it on its own terms, and file it where it still makes sense in two years.
Generative AI does not come with that artifact ready-made, which is where most conversations stall. What you can offer instead sits in a different category: retrieval citations tying the output back to identifiable sources, logging of the prompt and context assembled at inference time, structured outputs whose fields trace to specific inputs, reasoning traces the system produced as it worked. That is a real answer, often a good one. But it is not the same kind of object as a coefficient, and pretending otherwise costs you the room faster than admitting the gap.
Why the traditional answer cannot simply be transferred
The old answer travelled well because in a linear model the explanation was not an interpretation of the model — it was the model, and nothing had to be reconstructed afterwards.
The person asking is not being obstructive. They are thinking about the day they must defend one specific decision to somebody with no interest in how the technology works. On that day, “the system performs well in aggregate” is not a defence of a single case. Aggregate performance and per-case justification are different claims, which is why the question is asked at the level of one output rather than the system.
The questions that follow the first one
Generative AI does not get a new set of questions. It gets the existing ones from a well-established discipline, asked harder, because the usual answers no longer land.
What the model was exposed to
Data lineage. Where the training or grounding data came from, what sits in the retrieval corpus today, whether anything is in there that should not be, and what changed last month. A traditional model had a documented feature set drawn from known systems of record. With a generative system the boundary is fuzzier, and answering honestly requires inventory work most teams have not done.
Whether the same input produces the same output
Reproducibility and version control. If someone re-runs the case in six months, do they get the same answer, and if not, can you say why? That means naming which model version, which prompt version, which retrieval index and which configuration produced the original. Traditional models were versioned as a matter of course; generative systems change constantly, sometimes by a vendor-side upgrade nobody logged.
How you would notice it getting worse
Monitoring and drift. Credit and fraud models have outcome data attached: the loan defaults or it does not, the transaction was fraudulent or not. Much generative output has no clean label, which makes the monitored quantity a hard design question. Structured evaluation sets, sampled human review, escalation and correction rates tracked over time — those are the answers that hold up. “Users would tell us if it got bad” is not a control.
Who can overrule it, and whether they actually do
Human-in-the-loop design and override authority. Not whether a human is nominally in the loop, but who holds the authority to reject an output, what they see when they use it, and whether there is evidence they use it at all. If review volume is high enough that approval has become reflexive, the control exists on paper and not in practice — which is what a reviewer is probing for.
What happens when the dependency is somebody else’s
Vendor and third-party dependency. In most cases the model is not yours. It can change without notice, its training data may not be visible to you, and your contractual position on notification and audit rights may be thinner than your risk documentation assumes. What happens if the capability is withdrawn, and how much collapses with it, is fair to ask before the workflow is load-bearing.
They are not rejecting it — they are grilling it
Worth being clear about why: model-risk officers are not rejecting the technology. They are grilling it. That is the job, and it is the right job. The discipline exists because models have been trusted past their evidence before, and the people asking are usually the ones who cleaned that up.
We would go further: these questions improve the system regardless of who asks them. If you cannot say what changed in your retrieval corpus last month, that is an operational gap, not a compliance formality. These are questions operators should be able to answer before deployment rather than during.
What preparing properly looks like
The useful takeaway is about preparation. Walk in knowing what you are offering in place of coefficients — name the substitute rather than letting the room find the absence — and be ready to defend it at the level of one specific output, not the system in general. Bring a single case end to end: the input, the context, the output, the trace, the review step, the person who signed it.
Then say plainly what your evidence does not cover. Conceding the limit is what makes the rest credible. It also tends to reveal that the honest fix is narrower scope, not better argument — a system that drafts for human sign-off answers these questions far more easily than one deciding on its own.
A capability story, however good, does not answer the question that was actually asked.
Where this work actually lives
This work is delivered through Interactive Intel, Alton Worldwide’s agentic-AI practice — a small, Miami-based team that designs, builds and runs production AI agents in environments where somebody independent will eventually ask how an output came to be. Every engagement is led by Paul personally and scoped around a measurable outcome.
The entry point is deliberately small: a fixed-price AI Opportunity Scan — one workflow, two weeks, the payback math in writing before committing to anything larger. If the workflow cannot be evidenced to the standard your reviewers hold it to, we would rather tell you in week two than let you find out in front of them.
Related reading: AI for medical practices · AI consulting for small business