Workplace change requests at a major bank
Designing where an AI agent belongs in a governance process, and what it should not be trusted to do
Client details removed to respect confidentiality.
The same question, with and without the agent. The version on the right is what the rest of this case study is about.
Snapshot
Sector: Banking
My role: Lead designer, end to end
Platform: ServiceNow, a requester portal and an agent workspace
Status: Built as a proof of concept, on its way to production
A leader submits a request to make a workplace change: closing an office, restructuring a team, changing shift patterns. Reviewers from HR, risk, legal and corporate affairs check it before it goes ahead. Requests were coming back weeks later with feedback nobody could act on. The bank asked us to put AI agents into the process to speed it up. I designed a form that shows leaders what reviewers check for while they're still writing, and a reviewer's screen that makes the specific comment the easy one to send.
The problem
A leader does this once or twice in their career. The reviewers do it every day.
So the leader answers a fifteen question form with no way of knowing what a good answer looks like, guesses, and finds out weeks later that it wasn't enough. What comes back is a set of comments they have to interpret. Some are vague. Some assume knowledge of a policy the leader has never had to think about. Either way they're guessing again, and the next return is another three weeks.
The delay was never in the writing. It was in the round trips. So the agent does one job throughout: it reads an answer and works out what's missing, against what reviewers check for. Everything below is a different place that same job appears, and the design questions are mostly about how much authority it gets in each one.
The checks are visible while the answer is being written
The standard reviewers hold in their heads was never written down anywhere the leader could see.
What reviewers look for, listed under the answer and scored as it's written. Three of the five are covered here. The two that aren't are what would otherwise come back weeks later as a request for more detail.
A point not yet covered, and the same point once the answer addresses it. The tick is earned by what's written, not by pressing anything.
Every point present in the answer, and none of it scored by the person writing it.
The form opens with three questions that decide what else gets asked, so a leader only sees what applies to their change. The rail alongside runs the whole time: what needs attention, who will review this, and what it is likely to take.
Who is judging this answer and on what. The agent shows the lens rather than passing judgement itself.
Two ways to get help writing an answer
Someone facing a blank box and someone who has already written a paragraph need different things.
Four short questions produce a draft that already covers what reviewers ask for. Drafting stays disabled below two answers, since one answer makes a sentence too thin to be worth offering.
The draft can replace what's in the box or be added to it. They are different acts, so they are different buttons.
For someone who has already written something, their sentence stays in grey and every suggestion sits alongside it, accepted or rejected on its own.
It refuses to invent an answer from nothing, and points to the tab that can help instead.
A head start is offered, not assumed
Collapsed by default. The questions are the front door; describing the change in a sentence is the shortcut.
Every extracted fact names the words it came from. A fact a leader can't check isn't worth offering.
Findings name the reviewer who would raise them
The form reads across its own answers. A hundred or more people affected but no people impacts ticked. A site move described in the prose with the location box left blank.
Blockers above queries, each with a named owner and the action that resolves it.
The empty state says what was checked rather than just confirming everything is fine, and a resolved finding stays visible.
Being wrong is a designed path
An agent that can't be contradicted is one people quietly stop trusting.
Three ways out: take the fix, park it, or say it's wrong. Dismissing and contradicting were doing the same job and meant different things, so they were separated.
What the check actually read, in the words the leader saw when they answered it.
A contradicted claim doesn't disappear. The count clears, but the claim drops into its own group, struck through, so the disagreement stays on the record.
The prediction is visible before anything is sent
What the request is likely to take, derived from answers already given rather than asking for the same facts twice.
Always visible, and it names the reason rather than counting problems. It replaces a Submit button that told you nothing until it was too late.
The governance panel run early. The verdict leads, the issues group beneath it, and everything the agent can fix is one action. These are the same checks that would otherwise arrive three weeks later.
A returned request is its own designed state
The reviewer's words are quoted, then decoded into specific gaps, attributed to the agent rather than the reviewer. The old answer sits beside the new one.
Needs tick off individually, so progress is visible instead of binary.
When an ask is genuinely unclear, the leader can ask the reviewer back. It says plainly that asking pauses the clock and drafting does not.
Writing a specific comment has to be easier than writing a vague one
Reviewers don't write vague comments out of laziness. Writing out exactly what you need takes about ten minutes. A short comment that gestures at the problem takes seconds, and both send the request back. With eleven assessments open and one overdue, the short one wins.
So the same check runs in the reviewer's comment box, and the asks arrive already written and already selected.
Nothing appears until a field is chosen. The agent has no opinion about a whole record, only about a specific answer.
Three options, not two. Answering it yourself removes a round trip rather than improving one: if the HR business partner already holds the answer, sending it back so someone else can find it helps nobody.
The comment rewrites as the options change, and stays editable. What's being sent is always visible.
Raising an ask moves a point to amber, never to green. An ask is not an answer, and green is only for points genuinely closed.
What the leader sees. On the left, three vague cards to work out. On the right, the same screen after the reviewer used theirs: specific asks, and the points the reviewer answered herself marked as round trips that never happened.
Outcome
The bank asked for agents to speed things up, and the first prototype had a chatbot on the form. We dropped it once it was clear the delay wasn't in the writing, it was in the round trips.
What replaced it does one job in several places, and holds to a few rules: it never overwrites what someone wrote, it shows what it read, it can be told it's wrong, and it never says an answer is good, only what reviewers check for. The screen with no AI on it was kept as a control the whole way through, which is how we could tell what the restructure earned on its own.
Reflection
The brief named a technology. The work was finding where it belonged.
Adding an agent to the writing would have looked like progress and changed nothing, because the time was going somewhere else. Most of the value came from moving a check that already existed to a point where it was still cheap to act on.
An agent's limits are part of its design.
Deciding what it won't do, invent an answer, overwrite a sentence, or mark something as good, mattered more than what it can do. Those limits are why a reviewer would put their name to what it produced.