Aimasteryhub

Outcome Library

Every entry here is a build, with the numbers the team measured.

Six composite sketches. Each one names the task, the stack in a single line, the test set, and where the design stops. No organisation is identified.

These are composite sketches of the projects cohorts build. Figures come from the teams' own test sets and sandbox environments; they are illustrative and do not describe named organisations.

composite sketch

55 → 9 min

reading plus notes per file

120

tender PDFs in the test set

v7

prompt version at measurement

Tender and RFP summariser

A procurement pair needed a first pass over long tenders: parties, dates, penalties, and a list of questions still unanswered. They built a PDF text extract, a schema-constrained model call, and a markdown brief per file. Stack: Python, a hosted large language model, regex date cleanup, CSV export.

Timing used a stopwatch on the same 120 documents before and after, with two reviewers checking ten briefs for missed obligations. The design stops at a draft for a human; it does not submit bids. Measured on the team's own test set of 120 items in a sandbox environment.

composite sketch

71%

tickets routed with no agent touch

3,200

labelled historic tickets

19%

held for a person on low confidence

Customer support triage router

An operations seat owned a queue that mixed billing, access, and product defects. The build classified intent, suggested a team, and wrote a two-line handoff. Stack: existing helpdesk export, embeddings for near-duplicate search, a classifier prompt, a confidence gate.

They scored against a 3,200-row labelled history, then shadow-ran 400 live-shaped sandbox tickets. Billing edge cases and angry threads stayed with people. The router does not send customer-facing replies. Measured on the team's own test set of 3,200 items in a sandbox environment.

composite sketch

0.4%

mismatched lines after the assistant

6.1%

mismatches under manual keying

860

invoice and PO pairs scored

Invoice reconciliation assistant

Finance staff were keying vendor invoices against purchase orders. The artefact matches amounts, quantities and tax, then lists unmatched lines. Stack: spreadsheet ingest, fuzzy string match, a model call only on residual rows, an exceptions sheet.

They compared the assistant’s exception rate with the prior month’s manual log on 860 pairs in a sandbox ledger. Residual 0.4% still needed a person for credit notes and split shipments. The tool does not post to the ledger. Measured on the team's own test set of 860 items in a sandbox environment.

composite sketch

2,400

internal policy documents indexed

64%

questions with a cited correct passage

40

held-out questions in the gold set

Internal policy search

A small operations group could not find which version of a leave or vendor rule applied. They indexed 2,400 PDFs and intranet pages, then answered questions with quoted passages. Stack: chunking at 400 and 800 tokens, hybrid keyword plus embeddings, citation required in the prompt.

Hit rate was counted on 40 questions written before indexing. Failures clustered on tables and scanned images. The index is not a legal opinion and does not replace the policy owner. Measured on the team's own test set of 40 items in a sandbox environment.

composite sketch

12 → 4 min

to structure a 20-minute call

90

call recordings in the sample

3

fields still filled by the seller

Sales call note structuring

Account staff were rewriting call recordings into a CRM shape. The build transcribes, fills next step, product mentioned, and risk, then leaves three judgement fields blank. Stack: speech-to-text, a schema prompt, CRM CSV import, a review screen.

Time was measured on 90 sandbox recordings with the same people filling the CRM before and after. Names of real customers were stripped. The assistant does not send follow-up mail. Measured on the team's own test set of 90 items in a sandbox environment.

composite sketch

38

checklist items generated per change

22

items kept after engineer review

15

past incident write-ups in the corpus

QA checklist generator for an engineering team

A product engineering group wanted a first-pass test list from a change note and a handful of incident reports. The artefact drafts a checklist, maps each item to a source sentence, and waits for an engineer to delete or add rows. Stack: retrieval over 15 incident PDFs, a generator prompt, a spreadsheet the CI job can read.

They compared generated lists with the team’s last eight hand-written checklists on the same change notes. About two fifths of rows were dropped as noise or duplicates. The generator does not mark a release as safe. Measured on the team's own test set of 8 change notes in a sandbox environment.

Method

How we count

A baseline is whatever the team already did: minutes on a document, mismatch rate on a ledger, share of tickets a person had to open. We freeze that number on a labelled set before the first prompt. After a change we rerun the same set. The difference is the only figure that belongs on this page.

We do not publish a lone “model accuracy” score. A model can be fluent and still miss a penalty clause or route a billing ticket into product. The task metric (time, match, escalation) sits in front. Where a classifier is used, we still report the share that needed a person, because that is the number operations can staff.

Limits

What usually breaks first

These failure modes show up in clinic more often than model choice. Each one has a paragraph because a label without a mechanism is not useful.

The evaluation set was written after the demo

If cases are chosen to flatter the prompt, the score moves in class and collapses on Monday’s real queue. We now refuse to look at a demo until the 40-case file exists.

Tables and scans never entered the index

Retrieval looks finished until the answer lives in a screenshot or a merged cell. Hit rate then plateaus in the sixties and people blame the model.

No human gate for low confidence

A router that always emits a team will be wrong on the long tail. The 19% hold-back in the triage sketch is the part that kept the 71% honest.

The owner of the workflow was not in the room

A builder can ship a schema the desk will not use. Programmes that mix an operator and a developer on the same artefact fail less often on handover week.

Bring the workflow you want to score.

Enrolment starts with a 20-minute call and a S$400 deposit. Twelve seats, evenings in Bras Basah, live stream if you are outside Singapore.