
Analysts rewrite the same rationale for every exception
Prepared for Dana Reyes, September 2026. Inside: the problem and the goal, then the 2-week sprint we recommend for finding the first piece of work and gathering what the AI needs to do it. Every number in it is one your team gave us.
Executive summary
The problem
Price exceptions should be approved the same day, with reasoning a manager can defend to finance.
Dana listed 3 problems. The one to start on: analysts rewrite the same rationale for every exception. Today it runs about 42 times a week and takes 35 minutes per run.
The goal
| Done means | An exception packet a manager can approve in twenty minutes, with the reasoning already written and sourced. |
| Measured on | Time savedAccuracy matching a human reviewer |
The proposed solution
An AI assistant, working inside Word, where reviewers already edit the packet, takes over 4 of the 9 steps in this workflow, starting with the rationale paragraph. Priya Sharma checks every output before it moves on, so the judgment and the sign-off stay with your team.
What this document does
It gets your team started. The next pages walk you through the first steps on this workflow, with the people you already have. If you want help, from day one or once the sprint is done, the last page says how.
The plan on one page
Price exceptions should be approved the same day, with reasoning a manager can defend to finance.
Done means: An exception packet a manager can approve in twenty minutes, with the reasoning already written and sourced.
Start here
We would build "Analyst writes the rationale paragraph" first, in Copilot. It takes 2 hours of the 15.5 hours a full pass takes today, it follows the same rules every run, and it comes around again and again. That is the kind of step you can build in two weeks and judge in an afternoon.
The math
| 35 min per run × 42 runs a week | 25 h a week |
| 65% of that is routine work AI can take | 16 h a week |
| × 46 working weeks a year | 733 h a year |
| ÷ 7.5 hours a working day | 98 working days |
These figures are estimates built from your answers. We would welcome the chance to run a full discovery program with your team, map each of these in detail, and make sure they hold up.
Today
What a first build needs
This week
- Sit with the workflow owner through one exception end to end and time each step. Write the number next to "35 minutes per run" from the form. That is your baseline.
- Ask the decision owner what makes a packet easy to approve, what sends one back, and how often. The rework rate is your accuracy baseline.
- Cut the first piece down until the owner can show a real draft by Friday: one exception type, one contract clause family, and the output is the paragraph only, not the packet.
What to measure
You chose time saved, accuracy matching a human reviewer. Write each of them down in week 1, and again at day 30.
- Time saved. Hours per run, before and after. Our projection is 16 h a week, or about 98 working days a year.
- Accuracy matching a human reviewer. The share of outputs the reviewer accepts as they are. The target is nine in ten.
The workflow, drawn
Reading the map
- 9 steps, 15.5 hours a full pass by the step times given.
- 4 of them marked for AI. They are built one at a time, in the order they run.
- The first piece, "Analyst writes the rationale paragraph", is 2 hours of every run.
- Dana Reyes keeps the decision. Priya Sharma keeps the review.
Where you are, and what's next
Where your team is right now
Northline has 120 AI licenses across Copilot and ChatGPT, 40 of them paid, and usage is low. We see this pattern often: licenses bought, usage low. The fix is to pick the work before you pick the tool, then measure it honestly. You have already done the first half by choosing this workflow.
The tooling you need
Nothing new. The work happens in Word, Excel, Outlook and Salesforce, and you want the AI to sit inside Word, where reviewers already edit the packet. Copilot in Word is that tool, and Northline already has the licenses. Everything in the sprint below can be done with it and a shared folder.
We recommend starting with a 2-week sprint
Two weeks keeps it concrete, and each week is real work. Adjust the pace to your team.
What you will have at the end
- A one-page description of the rationale step, with the time it takes today.
- The workflow owner named, and a reviewer for the first results.
- A folder with the packet template and three approved packets to learn from.
- A prompt the owner can run in Copilot on the next real exception.
- A short list of the things the AI cannot reach yet, and who can fix that.
Week 1 · Identify and Refine
The aim this week is not to build anything. It is to see the work as it really happens and cut it down to one piece one person can own.
Who to meet, and why
Questions to ask them
New questions, beyond the form, to get closer to how the work really runs.
- What would you need to see before you trusted a rationale you did not write?
- Which of the 9 steps waits on a person, and which waits on a decision?
- When you write the rationale, what do you look at first, and what do you copy from last time?
- Show me a rationale that was approved first time, and one that came back. What was the difference?
Cut the first piece down until you can show it by Friday
The first piece is the rationale paragraph. On its own it is 2 hours of a 7.5 hour run. Cut it further: one exception type, one contract clause family, and the output is the paragraph only, not the packet. If the owner cannot show a real draft by Friday, the piece is probably still too big.
Week 2 · Feed and Standardize
Gather what a new analyst would need on day one, and hand it to the AI.
The five things the AI needs, for your case
Turn it into a working prompt
A first version to start from. Making it reliable on real cases is the part we help with.
Standardize the output: same sections, same order
Every packet has the same five sections, in this order, so the AI fills them the same way every time: account and contract reference, the clause the exception rests on, rationale paragraph, recommended action, reviewer sign-off line.
The first 30 days
You gave yourself two weeks to a working version. That is tight, so run weeks 1 and 2 together and start building on the first three finished examples while the rest are still being collected.
- Sit with the workflow owner through one exception end to end and time each step. Write the number next to "35 minutes per run" from the form. That is your baseline.
- Ask the decision owner what makes a packet easy to approve, what sends one back, and how often. The rework rate is your accuracy baseline.
- Cut the first piece down until the owner can show a real draft by Friday: one exception type, one contract clause family, and the output is the paragraph only, not the packet.
- One page on the rationale step, with the time it takes today
- A baseline in hours, written down
- One owner and one reviewer named
- Clean up the example you have, then collect nine more finished packets from the last quarter. Mark the best three as the standard, and seal three more as the exam the AI never sees.
- Write the checklist: the five things every packet has to contain, one line each, and save it in SharePoint, in the Revenue Ops packet templates.
- In Copilot, build one prompt that does "Analyst writes the rationale paragraph" and nothing else. Give it the checklist and the three gold examples as context.
- Run it on the three sealed cases. Priya Sharma scores each output against the human version: accept as is, accept with edits, or reject.
- Access: "The billing system only gives us a nightly PDF, and the contract archive is scanned paper…" For a test, a nightly export into a folder the AI can read is enough. Get that running now rather than in week 3.
- 3 gold examples and 3 sealed cases in one folder the team can open
- A checklist the reviewer has signed
- Accept rate on the sealed cases, written down
- Kick it off the way the work really starts (a person kicks it off) and land the output where the finished packet lives today.
- Priya Sharma reviews every output before it moves on. Nobody else's work changes yet.
- Keep a log of the first 20 runs: how long each took, whether it was accepted or sent back, and why.
- 20 real runs logged
- Accept rate at or above week 2
- Minutes of review per run, written down
- Compare the log against the baseline on two numbers: time per run, and the accept rate.
- Make one of four calls. Go, and hand over "Analyst finds the contract and the discount clause" the same way. Narrow, and shrink to the simplest variant before running again. Fix foundations, because the standard or the access was the problem rather than the AI. Or not yet, and park it with the reason written down.
- Write one paragraph with the two numbers in it and send it up.
- A before number and an after number
- One of the four calls, with the reason
- Leadership has read it
One workflow, done by AI, at an accuracy the reviewer would accept from a colleague: nine out of ten outputs accepted as they are, checked against sealed cases the AI never saw. Everything else in this plan is there to get one workflow over that line.
That last stretch of accuracy is what we do. If you find yourself stuck below the line, bring us the workflow and we will show you how deep it goes: the sources, the standard, the examples, and the review.
The roadblocks you will hit
The first 70 percent comes easily. The last 25 is where projects die.
The demo looks great, and then the real cases arrive. The reviewer sends back three in ten, and after the third bad one the team goes back to doing it by hand without telling anyone.
Only judge it on the sealed cases, never on a demo. Keep a log of every send-back with its cause, then fix the most common cause and run it again. In our experience the culprit is usually a rule that lives in someone’s head, or an example nobody wrote down. It is rarely the model.
You have been stuck between 80 and 95 percent for two weeks. That stretch is most of what we do.
The AI cannot read the one system that matters.
You already named it: "The billing system only gives us a nightly PDF, and the contract archive is scanned paper…" What tends to happen next is that someone pastes into the prompt by hand, and the test ends up measuring their patience instead of the AI.
For a test, a nightly export into a folder is all you need, so ask IT for that rather than an integration. If the documents are scans, run them through OCR once before anything else.
There is no export at all, or IT tells you the integration is a quarter away. We connect to most of these already.
Nobody agrees what good looks like.
The reviewer says the output is wrong. The builder says it matches the example. Both are right, because nobody ever agreed on the example in the first place.
Three gold examples and a one-page checklist, signed off by the person who actually reviews the work. An hour spent cleaning up one great example does more than a page of prompt instructions.
The reviewers cannot agree on what belongs on the checklist. Writing that standard down is the first thing we do on site.
The licenses exist. The habit does not.
The people on the free tier get worse answers and tell everyone the tool is not very good, and any paid seats sit with a handful of power users.
Put the first piece inside Copilot, where the work already happens, so nobody has to open a new window. Name one champion on the team, and hold a 30-minute open office hour every week for the first month. Most of the questions will take ten seconds to answer.
Usage has been flat for a month. Adoption is a training and habit problem, and that is what SuperHumans is built for.
Why most first projects stop here
How this worked for others
Thousands of reviewer hours a month were going into long technical reports, and known error classes were still slipping through.
A multi-agent reviewer, scored against sealed past reports, with every finding traced back to its source line.
- Over 90% of expert issue classes re-found
- Issues the experts had missed, caught
- About $2 a report
The sales team was spending 30 to 45 minutes researching each account by hand before any outreach.
Multi-agent account research inside Salesforce, with a person approving every enriched record before it is used.
- Under 2 minutes per account
- Thousands of accounts a week
- $1M+ in qualified pipeline
Copilot licenses across 600-plus staff, with usage concentrated in a handful of people.
Training built from their own lease, tenant and market workflows, a champion in each department, and adoption tracked month by month.
- +36 points on the Microsoft adoption score
- 20%+ time savings reported within weeks
- Self-sustaining champion network
Full-length case studies, with the method and the numbers behind each one: aiexperts.com/case-studies
The road after the first piece
Get the team using the AI it already pays for, on its own work, every week.
Map the next three workflows and rank them by hours spent and how fixed the rules are.
Run the 30 days against sealed cases and make the decision in writing.
Hand the next step on the map to the AI, behind the same review gate.
Your answers, and the words we used
Your details
- Name
- Dana Reyes
- dana.reyes@northline.com
- Phone
- +1 212 555 0148
- Company
- Northline Capital
- Role
- VP or department head
- Company size
- 250–999
- Team this touches
- 10–49
Current stack
- Tools in use
- Microsoft Copilot, ChatGPT, Meeting recaps in Teams
- AI licenses
- 120
- Tier
- 40 paid, the rest on the free tier
- Driving adoption of a paid tool?
- Yes, and usage is low
- On-prem or own cloud?
- No, everything is vendor-hosted
- Where the work happens
- Word documents, Excel spreadsheets, Outlook or email, A CRM or ERP
- Where the AI should sit
- Inside Word, where reviewers already edit the packet
The problem
- Decision or output to improve
- Price exceptions should be approved the same day, with reasoning a manager can defend to finance.
- Problems that may be low-hanging fruit
- Analysts rewrite the same rationale for every exception Contract terms are looked up by hand across four systems The approval packet is reformatted before every review
- Who owns the decision
- Dana Reyes, Revenue Operations
- What done looks like
- An exception packet a manager can approve in twenty minutes, with the reasoning already written and sourced.
- How success is measured
- Time saved, Accuracy matching a human reviewer
Priority
- Top priority workflow
- Analysts rewrite the same rationale for every exception
- Ranked
- 1. Analysts rewrite the same rationale for every exception 2. Contract terms are looked up by hand across four systems 3. The approval packet is reformatted before every review
- Trusted for this output
- Priya Sharma
Breakdown
- The workflow, step by step
- 1. Analyst exports the exception list from the billing report (20 min) 2. Analyst filters out exceptions already approved last cycle (30 min) 3. Analyst opens each account in the CRM (30 min) 4. Analyst finds the contract and the discount clause (2 hrs) 5. Analyst checks the clause against current policy (1 hr) 6. Analyst writes the rationale paragraph (2 hrs) 7. Analyst assembles the packet in Word (1 hr) 8. Analyst emails the packet to the manager (10 min) 9. Manager reviews, approves or escalates (1 day)
- Step to build first
- Analyst writes the rationale paragraph
- Time to a working version
- Two weeks
Output
- Final form
- A document, SharePoint or a shared drive
- Example available?
- Yes, but it needs cleaning up
- Where the example lives
- SharePoint, in the Revenue Ops packet templates
- Has to contain, every time
- Account and contract reference The clause the exception rests on Rationale paragraph Recommended action Reviewer sign-off line
- Produced with AI before?
- Not yet
Access
- Hard for AI to reach
- The billing system only gives us a nightly PDF, and the contract archive is scanned paper before 2019.
- Tools flagged
- Salesforce, SharePoint, A legacy on-prem system, The monthly billing report
Workflow
- Steps handed to AI
- Analyst finds the contract and the discount clause; Analyst checks the clause against current policy; Analyst writes the rationale paragraph; Analyst assembles the packet in Word
- Named on the map
- Priya Sharma (step 1), Priya Sharma (step 2), Dana Reyes (step 9)
- Time per run today
- 35 minutes per run
- Runs a week
- 42
- Who reviews the output
- A named individual
- What starts it
- A person kicks it off
- Who has to change how they work
- A single team
Glossary
Context. What you would hand a new hire before asking them to do the job: the policy, the records, three good examples. AI without it guesses.
First piece. One step out of the workflow, small enough to build and judge on its own.
Baseline. Today's number, written down before anything changes.
Gold example. A finished output everyone agrees is good. The AI is judged against it.
Sealed case. A real past case the AI never sees during setup. It is the exam.
Send-back. An output the reviewer rejects. Counting them, by cause, is how accuracy is measured.
Handoff. The point where the AI stops and a person takes over. Marked on the map with a dashed box.
Human in the loop. AI does a step, a person checks it, then the work moves on.
SuperHumans. Our training. It gets a team using AI on their own work.
SuperTools. Our engineering. Custom AI built for one company, scoped from a first piece that already proved itself.
Longer write-ups of most of these are on the AI Experts blog at aiexperts.com/blog.
Keep going with your own AI
How to use these
- Open Claude, ChatGPT or Copilot and pick the most capable model available (the reasoning or "thinking" option if there is one).
- Attach this whole PDF. The answers on the previous page are what the prompt works from.
- Copy one prompt below exactly as written and paste it in. Answer its questions when it asks; it is meant to interview you.
- Bring the result to the person who owns the decision. None of this replaces the day-30 decision.
You are helping me run week 2 of the attached plan. Using the answers in the appendix, draft the one-page checklist for the output described under "Has to contain, every time". Then ask me, one question at a time, what a reviewer looks for that is not on the list yet, until you have a standard I would sign. Finish with the checklist as a numbered list I can paste into a document.
You are helping me run week 2 of the attached plan. The step to build first is named under "Step to build first" in the appendix. Interview me, one question at a time, about the inputs that step needs, the rules it follows, and what a good output looks like. Then write a reusable prompt I can run in my AI tool for that one step, with a place to paste the inputs and the checklist. Do not automate any other step.
You are the reviewer named in the attached plan. I will paste two versions of the same piece of work: the one a person produced and the one the AI produced. Compare them section by section against the checklist in the appendix. Mark each section accept, accept with edits, or reject, give the cause for every reject in one line, and end with a single accept rate. Do not soften the score.
Using the attached plan as the pattern, help me find the next workflow to map. Ask me, one question at a time, what work my team does that repeats every week, follows the same rules each time, and produces a document or a record. Rank the answers by hours spent and how fixed the rules are. Then draft the step-by-step breakdown for the top one in the same format as "The workflow, step by step" in the appendix.
Want a hand? Contact us.
The sprint is yours to run, and most teams can. Some want help from day one. Others run the sprint, then bring us in for what comes next. Either works.
What comes after the sprint
Where teams usually want help
- Accuracy. Setting up the reviewer check so the rework rate is measured, and tuning the prompt and context until it holds.
- Consistency. Making the output the same whoever runs it, so every analyst gets the same result.
- Large documents. Contracts, scanned archives and nightly PDFs are where a first attempt usually stalls. We have done this before.
- Security and governance. Writing down what the AI may read, where its outputs are kept, and who approves a change to the prompt, so IT and legal can sign it before the first real case.
What to send us
The two Friday deliverables. That is enough for us to scope the build and give you a fixed price. No slides needed.



