Capabilities
What do you want to stop doing?
Here is everything we build and teach, 31 things across 5 groups, one line each. Open any one and you will see an illustrative sample transcript: the kind of steps, artifacts and handoffs a working agent can produce.
Approved proof
These are approved, anonymized proof points from real engagements. The expandable transcripts below are separate illustrative demos, written as representative examples rather than client run logs.
99%
accuracy on a recurring task, using tools they already owned
$1M+
pipeline from one agent-built account research system
Days
saved per deal for a regulated M&A team
140
employees trained inside one regulated finance engagement
Work that runs without you
These finish the job on their own schedule and hand back something you can check. You keep the code, the prompts and the runbook, so your team can change any of it without us.

Illustrative transcript demos
These samples are anonymized, representative examples of how the capability works. They are not client case studies, production logs, or confidential customer evidence.
- See how this one was builtTurn last night's support inbox into ticketsAgent: opening 63 unread threads in support@ Gmail4 of them are the same outage. Merging those into one ticket.Agent: filing against the queue rules Jira Service Management
Subject Queue Owner Card declined at checkout Billing Billing rota SSO login loop Access IT Seat count change Accounts Renewals Export stalls at 90% Platform Data engineering Invoice address is wrong Billing Billing rota API key rotation question Access IT 60 tickets open in 4 minutes. 5 of them went to a person first because the ask was unreadable. - See how this one was builtPrep Thursday's pipeline reviewResearcher: pulling 12 open deals SalesforceAnalyst: scoring slip risk against last quarter SnowflakeTwo deals have gone quiet for 19 days. Those go at the top.Writer: posting the one page brief SlackDeals covered12Flagged as slipping2Run time2 min 04 sSame prep by hand45 minHandoffs waiting on a person0Brief was in the channel 26 minutes before the review, and it moved the two quiet deals up from 11 and 12 to the top.
- Draft this week's renewal repliesAgent: drafting 28 replies from the account history SalesforceAnything above $50k of annual value stops here for a named approver.A pause or a cancellation is not a reply. Those get booked as a call whatever the value.Agent: queueing the held ones Front
Account Ask Value Route Regional insurer Add 40 seats $38k Sent as drafted Logistics group Discount on renewal $120k Ops lead approves Health system Contract wording change $210k Legal approves Retail chain Pause for a quarter $44k Needs a call 22 sent as drafted, 4 held for a named approver, 2 booked as calls. Nothing above $50k left the building unread. - Pull renewal dates and caps from the Q3 vendor contractsAgent: reading 412 PDFs SharePoint9 came back as scans with no text layer. Those went through OCR first. Document IntelligenceAgent: locating the term and pricing clauses
Vendor Renewal Cap Freight carrier 2026-01-31, 60 day notice 4.0% Payroll platform 2026-03-15, auto renews CPI + 2% Cleaning services 2026-06-01, no auto renew None Data provider 2025-12-20, 90 day notice 7.5% Print vendor 2026-02-08, no auto renew 3.0% Badge supplier 2026-04-22, auto renews 5.0% Translation agency 2026-05-09, no auto renew CPI 412 of 412 parsed in 19 minutes. 6 went to a person where the cap was handwritten in the margin. - Is the contract classifier safe to ship?Agent: scoring 1,240 labeled cases from your own files Eval harnessAgent: sweeping the threshold from 0.50 to 0.950.82 is where the curve stops paying. Past it recall falls faster than precision rises.Labeled cases1,240Positives in the set310Caught at 0.82293False positives6Precision / recall98.0% / 94.5%Cost per 1,000 docs$3.40Shipped at 0.82. The 40 cases that land between 0.60 and 0.82 go to a reviewer, which is 3% of the set.
- Have the Monday numbers ready before standup06:00 ET. Agent wakes.Agent: reconciling three sources SnowflakeThe finance feed was 4 minutes behind. Held the post until it landed.Agent: posting to the channel SlackAgent: writing the result to its own run log, 14 Mondays deep Run logDelivered 06:07 ETweekly-revenue.xlsxchurn-watchlist.csvstandup-summary.mdIn the channel before standup 14 Mondays in a row. The two mornings the upstream feed was down, the notice was posted by 06:02 and the numbers landed by 06:40.
- Write down how the senior analyst scopes a renewalAgent: reading 40 past scopes Google Drive6 of them broke the pattern. Booked 25 minutes with the analyst to ask why.Analyst: five of the six are already covered by three standing rules. One real exception left.Rule 1Multi site accounts get priced per site, never per seat, even when the buyer asks for seats.Rule 2Anything touching patient data goes to legal before a number is quoted.Rule 3Quote the list rate first. Discounts come off it in writing, in a second message.ExceptionRenewals inside 30 days of a support escalation hold the current rate.Three rules and one exception, out of 40 past scopes and 25 minutes of analyst review, running as a skill anyone on the team can call.
Plugging AI into your data
Your systems hold the answers. The model needs a login to reach them.

Illustrative transcript demos
These samples are anonymized, representative examples of how the capability works. They are not client case studies, production logs, or confidential customer evidence.
- Which accounts churned last quarter and why?Agent: connecting Salesforce MCPAgent: joining ticket history Zendesk MCPTwo of the four never filled in an exit survey. Read the closing notes instead.
Account ARR Tickets 30d Reason given Regional bank $180k 14 Missed the integration date Media group $96k 3 Budget cut Logistics firm $74k 21 Support response time Retail chain $61k 2 Moved to an in house build Four accounts, $411k of ARR, and two of the four name something the support queue could have caught. Same answer whether you ask in Claude or in Copilot. - Who can reach the customer database through AI right now?Agent: listing every tool the estate can reach this month GatewayAgent: resolving each team down to the roles behind itMarketing came back with more than anyone granted it. The access is inherited from a shared role, not assigned.What the marketing role could reach before this runCampaign analyticsEmail sendingModel endpointsSupport tickets, readCustomer database, readBilling, readMost of that access was intended. Billing and customer records were not, and they were scoped out the same afternoon. The gateway log also showed which requests were refused by policy, so the question of who read what before today has an answer.
- See how this one was builtReconcile the SAP invoices against the Excel trackerAgent: pulling 2,900 invoice lines SAPAgent: matching against the tracker Excel212 lines carry no PO number. Matched those on vendor and amount, then flagged them.
Invoice SAP Tracker Difference INV-40218 $18,400.00 $18,400.00 Matched INV-40251 $7,250.00 $7,520.00 Digits swapped INV-40266 $31,000.00 Missing Never entered INV-40290 $4,180.00 $4,180.00 Matched INV-40311 $9,940.00 $994.00 Decimal moved INV-40337 $12,600.00 $12,600.00 Matched 2,900 lines in 6 minutes. 31 mismatches worth $84k, each one linked back to the source row so nobody has to go looking. - What did we commit to on data retention for EU clients?Agent: searching 1,100 documents SharePointDropped 40 drafts and unsigned copies. Three signed sources left.MSA 2024, s.9.3Client data is deleted within 30 days of termination unless retention is required by law.DPA annex IIEU personal data stays in eu-central-1 and is not replicated outside the region.Security addendum, p.4Backups roll off at 35 days. That is the number to quote, not 30.Three signed sources, each with the clause it came from. The security addendum is the one people miss, and 35 days is the number that ends up in the answer.
- What can we hand our security team before the pilot?Agent: listing the documents on file for this deploymentThe pen test is from the last release. Flagged for a re-run before go live.
Document Covers Dated Data processing agreement Retention, sub-processors, deletion Signed 14 Feb 2026 Architecture diagram Where the model endpoint sits in your tenant 02 Mar 2026 SBOM Every dependency in the shipped build Per release Penetration test report Third party, previous release 18 Dec 2025 Threat model Prompt injection and data egress paths 09 Mar 2026 Five documents your security team can read before the pilot starts. The pen test is dated 18 Dec 2025, so it gets re-run against the shipping build and go live waits on it. - Why does the CRM show three versions of the same account?Agent: scanning 44,000 account records SalesforceAgent: resolving on domain, tax ID and billing addressAnything scoring under 0.75 stops and goes to a person.
Kept Merged in Signal Score Regional health provider Regional health prov. Domain, tax ID 0.98 Regional health provider REGIONAL HEALTH Billing address 0.94 Freight operator Freight operator LLC Tax ID 0.97 Research lab Research labs Name only 0.61 1,812 merged on their own. 96 scored under 0.75 and went to a person, including the research lab match, because it matched on the name and nothing else.
The code underneath
Half of what makes an agent slow or wrong is not in the agent. It is in code that was written before any of this existed.

Illustrative transcript demos
These samples are anonymized, representative examples of how the capability works. They are not client case studies, production logs, or confidential customer evidence.
- What in here will fight us when we add agents?Agent: mapping 84 services against their call graph GitHubAgent: reading 18 months of incidents JiraBilling sits in 9 of the 12 slowest paths. Most of what looks like a separate problem is downstream of it.Agent: checking how the engineers here already work with AIBlockingNo read replica on the billing database, so any agent reading customer history competes with checkout for the same connection pool.BlockingAuth is per service and inconsistent. One agent identity cannot be scoped across the estate, so today it would need six sets of credentials.Drag, not a blockerFour of the eleven engineers use AI daily and the rest have never been shown how. The tooling is bought and installed.Start here insteadThe document store is versioned, has an API and clean ownership. Two of your six agent ideas can run against it with no work at all.Two blockers, both in the billing path, and one place to start that needs nothing done to it first. The modernization plan puts those two ahead of the four agent ideas that depend on them, and leaves the other two to start now.
- Why is the agent slow and wrong on this one service?Agent: tracing 2,400 calls through the quoting service DatadogHalf the latency is one N+1 query that predates the API in front of it.Agent: reading what the endpoint actually returnsIt returns 340 fields. The agent needs 11 and has been inferring the rest from field names.Median call4.8s → 0.6sFields returned340 → 11Accuracy on the task71% → 94%Cost per 1,000 runs$62 → $9Calls traced2,400Services touched1Eight times faster and 23 points more accurate, and none of it came from a better prompt. The service now returns what was asked for instead of everything it had.
- Can someone make this agent leak?Agent: running 340 adversarial attempts at the support assistant Red team harnessAgent: testing what the tool credentials reach if the model is talked into tryingA PDF a customer uploaded carried instructions in it. The assistant read them as if a person had typed them.CriticalInstructions hidden in an uploaded PDF were followed. Attachments were trusted at the same level as the user prompt.HighThe support token could read every ticket, not only the requester's. Asking about "my other tickets" returned somebody else's.MediumFull stack traces came back in error messages, naming internal hostnames and the model version behind the endpoint.HeldThe other 337 attempts got nowhere, including every direct request to ignore the system prompt.Three findings out of 340 attempts, and the worst one needed nothing but a PDF. All three are closed, and the harness now runs against every release instead of once.
- Six months in. Is any of this being used?Agent: reading last month's runs across the assistant, the agents and the licensed seats Usage telemetryAgent: matching seats against the people who used them in a normal week212 licenses, 118 people using them in a normal week. Finance had 31 of the quiet ones, and all 31 sat in the team that got trained in January and nothing since.Two workflows nobody asked for showed up on their own. Contract summaries ran 640 times, from 9 people.Seats in weekly use118 / 212Runs last month4,900Hours back, month310Workflows found294 seats were paying for nothing, and the fix was a refresher for one team rather than more licenses. The same report lands every month, so the next drop gets caught in weeks instead of at renewal.
- An agent approved something it should not have. What happened?Agent: pulling the run by its id, with the prompt, the tools it called and who it was acting for Audit logThe run is there in full: 14 tool calls, two of them writes, both against the vendor record.The approval came from a rule that read an amount in a field the vendor controls. The log has the exact value it read.Agent: checking whether anything else hit that path in 90 days
Run Acting for Writes Verdict 8f21c, 11:04:22 Named approver 2 Wrong, reopened 7d02a, 9 days earlier Same approver 1 Correct 6b91f, 34 days earlier Ops lead 2 Correct 5a47c, 61 days earlier Ops lead 1 Correct 4c18e, 88 days earlier Named approver 2 Flagged The answer took 20 minutes instead of a week of asking around, because the question was already recorded. Logs are immutable, kept seven years, and export to the SIEM the security team already reads, which is the same trail a SOC 2 auditor asks for. - Can we stop paying frontier prices on this one task?Agent: assembling graded examples from two years of reviewed files6,000 came back, but 1,900 were duplicates or never reviewed. Kept the 4,100 a person actually signed off.Agent: training on 3,280 and holding back 820 to score against Eval harness
Model Accuracy Cost per 1,000 Median latency Frontier, prompted 95.1% $41.00 3.1s Small, fine-tuned 94.6% $2.30 0.4s Half a point of accuracy traded for 18 times less cost and 8 times less waiting, on 820 held-out cases neither model saw during training. It only works because the task is this narrow, and we say so before anyone signs.
Things people actually open
Most AI pilots die quietly because nobody opens the thing twice.

Illustrative transcript demos
These samples are anonymized, representative examples of how the capability works. They are not client case studies, production logs, or confidential customer evidence.
- How much PTO carries over into next year?Agent: searching the handbook and the 2026 policy update SharePointThe two disagree on the cap. The 2026 update is newer, so it wins.Agent: answering in the thread SlackHandbook, s.4.2Up to 5 unused days carry into the next year and expire on 31 March.2026 updateCarryover rises to 10 days for anyone past their third anniversary.Not covered hereSabbatical accrual sits with your HR partner. This policy stops at PTO.Answered in 6 seconds, in the thread it was asked in, with the 2026 update quoted over the older handbook and the part it does not cover named out loud.
- My scanner stopped syncing after the firmware updateAgent: matching the symptom to the 4.2 release notes Zendesk GuideAgent: checking the account for an open ticket ZendeskAnswered from 3 published articlesFirmware 4.2 release notesResetting the sync pairKnown issue: 4.2 on older docksFixed in two replies, off three published articles, with no ticket opened. Had the articles come up empty it would have gone to a person with the transcript attached.
- Check this intake packet against policyAgent: reading the intake packet, 38 pages iManageAgent: testing against 61 policy rules Policy registerAgent: running the sanctions check OFAC SDN
Section Rule Finding Source of funds KYC 3.1 Missing, nothing attached Beneficial owner KYC 4.2 At 18%, under the 25% threshold Sanctions check OFAC 1.0 Run 14 Aug 2026, clear Signature Ops 2.6 Director signed, officer required Two findings in 90 seconds, before the file left the desk: no source of funds document, and a director signing where an officer is required. The other 59 rules came back clear. - Show me what is waiting on a human right nowAgent: reading the live queue PostgresWaiting on approval14Oldest item2 h 40 mMedian time to approve7 mApproved today118Sent back for a fix6118 approved today and 6 sent back. The number the team watches is the oldest item, and right now it is 2 hours 40 minutes against a 4 hour target.
People and direction
The licenses arrive on Monday. Whether anyone's Tuesday looks different is a separate project.

Illustrative transcript demos
These samples are anonymized, representative examples of how the capability works. They are not client case studies, production logs, or confidential customer evidence.
- See how this one was builtWhat did the cohort ship by week three?Agent: reading the cohort workbooks Notion11 workflows were built. 7 ran this week. Counting only those.
Team Built Runs per week Time back Ops Vendor invoice triage 120 11 hrs Sales Account research brief 80 9 hrs Finance Invoice coding 60 5 hrs HR Interview scorecard summary 35 4 hrs Support Macro suggestions 20 2 hrs Seven of the eleven were still running at week three. The four biggest hand back 29 hours a week across 295 runs. - We pay for 900 Copilot seats. Who uses it?Agent: pulling 90 days of usage M365 adminAgent: grouping by team and by taskLegal shows 2 users out of 40. Their document system was never connected, which is the whole story.Seats paid for900Opened in the last 30 days214Used more than twice a week61Spend on seats nobody opened$247k a year686 seats had gone untouched for 30 days, which is $247k a year. 38 of them are in legal, where the fix is a document system nobody ever connected rather than anything to build.
- Someone in ops built a prompt everyone wants. Now what?Agent: reading the prompt and the 40 runs behind itAgent: testing it against the policy register Policy registerIt pasted customer names into a personal account. That is the one rule it broke, and the logic underneath it was sound.Agent: rebuilding it against the approved tenant, then versioning it under the author's name Skill registrySubmitted this quarter38Approved and shipped24Sent back for one fix11Refused outright3Median time to a decision4 daysNow used by another team1724 of the 38 shipped, and 17 of those are now used by a team other than the one that wrote them. The three refusals each named the rule they broke, so nobody was left guessing why.
- What do I tell the board in March?Agent: pulling the figures from the systems the work already runs in SnowflakeAgent: attaching a source to every numberTwo claims had nothing behind them. Cut both.Where we areSeveral agents in production, with more in pilot. Source: the run log.What it costCosts split between build time and licenses. Source: finance.What came backMaterial annualized return, measured against the same task before the build. Source: the eval set.What we do nextClaims first pass, once the data fix lands in May.Four slides, with a named source behind every figure on the page. Two claims came out because nothing backed them.
- See how this one was builtRank our 41 ideasAgent: scoring return, effort and data readiness Assessment workbook17 of the 41 turned out to be three ideas worded differently. Merged.Agent: marking the ones blocked on access
Idea Return Effort Start Claims first pass $1.2M a year 11 weeks After the data fix Contract review $210k a year 6 weeks Q3 RFP drafting $140k a year 2 weeks Now Vendor invoice triage $34k a year 3 weeks Now Field service scheduling Unclear 9 weeks Not yet 41 ideas down to 27 once the duplicates came out. Two start this quarter for 5 weeks of work, one waits on a data fix, and every parked item carries the reason it is parked. - See how this one was builtWhat is on the AI agenda this month?Agent: reading the last four steering notes NotionTwo of the four raise the same duplicate build. Promoting it to a decision.Decision duePick one vendor for document AI. Three parallel pilots is two too many, and they cost $180k a year between them.In flightVendor invoice triage went live on the 3rd. Four weeks of runs before it widens past ops.WatchTwo teams are building the same intake bot. One of them should stop this month.ClosedThe 06:00 reporting job has run 14 weeks. Handed to the data team on the 11th.One decision on the table and one duplicate build to stop. Picking a vendor switches off two of the three pilots, which is about $120k of the $180k a year.
- Can I paste a client contract into ChatGPT?Agent: checking the policy, one page IntranetContract text is client data under rule 1. That settles it.Rule 1Client documents go into the approved tenant tools only, which today means Copilot on the company account and the internal assistant. Personal accounts are out.If you are unsureAsk in the ai-help channel. You get an answer inside a working day, and the answer gets added here.One page, and it answers the question outright: client documents go in the approved tenant tools, personal accounts are out. What the page does not cover gets an answer inside a working day and then gets added to the page.
- See how this one was builtWho here is cleared to build agents that touch customer data?Agent: checking the certification register HRISBuilder level is the one that clears customer data. Practitioner covers everything else.
Person Standing Ops lead Builder, passed 12 Mar 2026, renews Mar 2027 Data engineer Builder, passed 21 Jan 2026, renews Jan 2027 Finance analyst Practitioner, passed 04 Apr 2026, renews Apr 2027 Marketing manager In progress, exam booked 12 Sep 2026 Two people cleared for customer data today, both Builder level, both renewing in 2027. A third sits the exam on 12 September.

Sometimes the answer is not to build anything
The light end
Approved real proof: a membership association hit 99% accuracy on a hard recurring task without a line of custom code. They already owned the tools. Nobody had pointed them at the actual job.
The heavy end
Approved real proof: a regulated advisory team needed accuracy they could defend to reviewers. That took custom engineering, entity resolution across name variants, formal evals, and a human signing off on every finding.
We will tell you which one you are looking at, including when it is the cheap one.
Do not worry about picking the right one
Tell us what is repetitive or expensive in your business. Thirty minutes later you will know whether AI can fix it and what it would take. If the answer is that you do not need us, we say so.
