14 January 2026
Why the evaluation set is the first thing we build
The second intake of Applied LLM Foundations spent week one arguing about tone in a customer email prompt. On Wednesday night the scores looked fine because each person tested three favourite examples. On Friday the same prompt failed 11 of 40 frozen cases, most of them short billing disputes the authors had never typed into the sandbox.
Since then the Monday brief is a file, not a discussion. Forty rows, expected output, a manual baseline. The model call waits until that file exists. The habit feels slow in hour one and saves the Thursday rewrite. It also makes prompt version 4 comparable with version 1, which is the only comparison the log can support.
If you take nothing else from Foundations, take the set. A clever system without it is a demonstration. A plain system with it is something you can hand to a colleague in week six.
11 February 2026
Retrieval fails quietly: three checks before you blame the model
Retrieval & Data Engineering cohorts often arrive sure the generator is the weak link. The log usually says otherwise. First check: did the gold question even retrieve the passage that contains the answer? We print the top five chunks. If the clause is absent, no prompt will invent a trustworthy citation.
Second check: tables and scans. A 2,400-document index that skipped image PDFs will sit around a 60% hit rate and look mysterious. Third check: chunk size. We run 400-token and 800-token versions on the same 40 questions. The winner is the one that moves the score, even if it is the less fashionable size.
Only after those three do we change the generator prompt. Blaming the model first is a reliable way to spend week four on theatre.
19 March 2026
Where automation actually saves hours in an ops team
Automation Studio teams like to graph every step of a process. The hours come from a narrower place. In the invoice sketch, the model barely ran: fuzzy match cleared most lines, and the language model saw only residuals. In triage, the saving was the 71% that never reached an agent, not a drafted reply.
We now ask for a volume map on Monday of week two: how many items per week, minutes each, and which 20% of cases eat the calendar. Agents that write prose for the long tail look impressive and move the weekly total by a few minutes. Routers, extractors and exception lists move the total by hours because they sit on the bulk.
If your sponsor cares about hours, put the metric in the log as minutes per 100 items. Pretty dashboards can wait until that number bends.
8 April 2026
Prompt versioning for people who hate version control
Several Operations Leads and a few developers refused Git in week one. We still needed history. The method that stuck was a folder of text files named task-v03-2026-04-08.txt plus a single sheet with version, date, score on the frozen 40, and a one-line reason for the edit.
Copying a file feels crude. It also made rollback real: when version 6 dropped two citation cases, the room opened version 5 in thirty seconds. People who later accepted Git kept the same naming in the commit message so the sheet and the repo agreed.
The rule is small. Never overwrite the file that produced the last published score. If you hate version control, you can still keep a drawer of dated prompts. The evaluation log is what makes the drawer useful.