composite sketch
55 → 9 min
reading plus notes per file
120
tender PDFs in the test set
v7
prompt version at measurement
Tender and RFP summariser
A procurement pair needed a first pass over long tenders: parties, dates, penalties, and a list of questions still unanswered. They built a PDF text extract, a schema-constrained model call, and a markdown brief per file. Stack: Python, a hosted large language model, regex date cleanup, CSV export.
Timing used a stopwatch on the same 120 documents before and after, with two reviewers checking ten briefs for missed obligations. The design stops at a draft for a human; it does not submit bids. Measured on the team's own test set of 120 items in a sandbox environment.
composite sketch
71%
tickets routed with no agent touch
3,200
labelled historic tickets
19%
held for a person on low confidence
Customer support triage router
An operations seat owned a queue that mixed billing, access, and product defects. The build classified intent, suggested a team, and wrote a two-line handoff. Stack: existing helpdesk export, embeddings for near-duplicate search, a classifier prompt, a confidence gate.
They scored against a 3,200-row labelled history, then shadow-ran 400 live-shaped sandbox tickets. Billing edge cases and angry threads stayed with people. The router does not send customer-facing replies. Measured on the team's own test set of 3,200 items in a sandbox environment.
composite sketch
0.4%
mismatched lines after the assistant
6.1%
mismatches under manual keying
860
invoice and PO pairs scored
Invoice reconciliation assistant
Finance staff were keying vendor invoices against purchase orders. The artefact matches amounts, quantities and tax, then lists unmatched lines. Stack: spreadsheet ingest, fuzzy string match, a model call only on residual rows, an exceptions sheet.
They compared the assistant’s exception rate with the prior month’s manual log on 860 pairs in a sandbox ledger. Residual 0.4% still needed a person for credit notes and split shipments. The tool does not post to the ledger. Measured on the team's own test set of 860 items in a sandbox environment.
composite sketch
2,400
internal policy documents indexed
64%
questions with a cited correct passage
40
held-out questions in the gold set
Internal policy search
A small operations group could not find which version of a leave or vendor rule applied. They indexed 2,400 PDFs and intranet pages, then answered questions with quoted passages. Stack: chunking at 400 and 800 tokens, hybrid keyword plus embeddings, citation required in the prompt.
Hit rate was counted on 40 questions written before indexing. Failures clustered on tables and scanned images. The index is not a legal opinion and does not replace the policy owner. Measured on the team's own test set of 40 items in a sandbox environment.
composite sketch
12 → 4 min
to structure a 20-minute call
90
call recordings in the sample
3
fields still filled by the seller
Sales call note structuring
Account staff were rewriting call recordings into a CRM shape. The build transcribes, fills next step, product mentioned, and risk, then leaves three judgement fields blank. Stack: speech-to-text, a schema prompt, CRM CSV import, a review screen.
Time was measured on 90 sandbox recordings with the same people filling the CRM before and after. Names of real customers were stripped. The assistant does not send follow-up mail. Measured on the team's own test set of 90 items in a sandbox environment.
composite sketch
38
checklist items generated per change
22
items kept after engineer review
15
past incident write-ups in the corpus
QA checklist generator for an engineering team
A product engineering group wanted a first-pass test list from a change note and a handful of incident reports. The artefact drafts a checklist, maps each item to a source sentence, and waits for an engineer to delete or add rows. Stack: retrieval over 15 incident PDFs, a generator prompt, a spreadsheet the CI job can read.
They compared generated lists with the team’s last eight hand-written checklists on the same change notes. About two fifths of rows were dropped as noise or duplicates. The generator does not mark a release as safe. Measured on the team's own test set of 8 change notes in a sandbox environment.