Aimasteryhub

Studio

Notes from the studio floor.

Short write-ups from 2026 cohorts. They describe habits that showed up in evaluation logs, not a newsletter cycle. All four pieces sit on this page.

We publish when a pattern repeats across more than one table. The tone matches the room: case counts, prompt versions, minutes. If you want the syllabus rather than the diary, start with Inside the Cohort.

14 January 2026

Why the evaluation set is the first thing we build

The second intake of Applied LLM Foundations spent week one arguing about tone in a customer email prompt. On Wednesday night the scores looked fine because each person tested three favourite examples. On Friday the same prompt failed 11 of 40 frozen cases, most of them short billing disputes the authors had never typed into the sandbox.

Since then the Monday brief is a file, not a discussion. Forty rows, expected output, a manual baseline. The model call waits until that file exists. The habit feels slow in hour one and saves the Thursday rewrite. It also makes prompt version 4 comparable with version 1, which is the only comparison the log can support.

If you take nothing else from Foundations, take the set. A clever system without it is a demonstration. A plain system with it is something you can hand to a colleague in week six.

11 February 2026

Retrieval fails quietly: three checks before you blame the model

Retrieval & Data Engineering cohorts often arrive sure the generator is the weak link. The log usually says otherwise. First check: did the gold question even retrieve the passage that contains the answer? We print the top five chunks. If the clause is absent, no prompt will invent a trustworthy citation.

Second check: tables and scans. A 2,400-document index that skipped image PDFs will sit around a 60% hit rate and look mysterious. Third check: chunk size. We run 400-token and 800-token versions on the same 40 questions. The winner is the one that moves the score, even if it is the less fashionable size.

Only after those three do we change the generator prompt. Blaming the model first is a reliable way to spend week four on theatre.

19 March 2026

Where automation actually saves hours in an ops team

Automation Studio teams like to graph every step of a process. The hours come from a narrower place. In the invoice sketch, the model barely ran: fuzzy match cleared most lines, and the language model saw only residuals. In triage, the saving was the 71% that never reached an agent, not a drafted reply.

We now ask for a volume map on Monday of week two: how many items per week, minutes each, and which 20% of cases eat the calendar. Agents that write prose for the long tail look impressive and move the weekly total by a few minutes. Routers, extractors and exception lists move the total by hours because they sit on the bulk.

If your sponsor cares about hours, put the metric in the log as minutes per 100 items. Pretty dashboards can wait until that number bends.

8 April 2026

Prompt versioning for people who hate version control

Several Operations Leads and a few developers refused Git in week one. We still needed history. The method that stuck was a folder of text files named task-v03-2026-04-08.txt plus a single sheet with version, date, score on the frozen 40, and a one-line reason for the edit.

Copying a file feels crude. It also made rollback real: when version 6 dropped two citation cases, the room opened version 5 in thirty seconds. People who later accepted Git kept the same naming in the commit message so the sheet and the repo agreed.

The rule is small. Never overwrite the file that produced the last published score. If you hate version control, you can still keep a drawer of dated prompts. The evaluation log is what makes the drawer useful.

Next

Topics we are working on next

Later 2026 cohorts will spend more studio time on evaluation packs for multilingual queues (English and Malay in the same gold set), on logging when a vendor hides token counts, and on handover notes for finance teams that will never run Python. We are also rewriting the chunking drills around tables exported from spreadsheets, because that is where hit rate stalls.

None of this is a new public programme yet. If one of those constraints is your actual week, mention it on the enrolment call so we can sit you in the run that already carries the extra drill.