Capstone: build and evaluate a triage prompt
Render a production-style prompt from a config, then grade a batch of model replies and report what to fix.
- Render a complete prompt from a versioned configuration
- Parse replies tolerantly and grade them against expected labels
- Turn results into a short report that points at the next fix
Put it all together on a realistic task: triaging support tickets into labels. A production prompt isn’t a hand-typed paragraph - it’s rendered from a configuration you can review and version: product name, label set, balanced examples, and the escaped ticket.
Then you evaluate: run the prompt over a labeled test set (here, the replies are given), parse every reply tolerantly, and report accuracy, parse failures and the most common confusion - which tells you what to change next (a clearer label definition? an example of the confusing case?).
from collections import Counter
mistakes = [("technical", "billing"), ("technical", "billing"), ("shipping", "billing")]
print(Counter(mistakes).most_common(1))[(('technical', 'billing'), 2)]Key takeaways
Render prompts from reviewed, versioned configuration.
Parse replies tolerantly and count parse failures separately from wrong answers.
The most common confusion points to the next prompt change.
Lesson quiz
5 questions · pass with 4 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Step 1: render the triage prompt
Read one JSON config with product, labels, examples (pairs of ticket and label) and ticket. Print exactly:
1<instructions>
2You triage support tickets for PRODUCT. Classify the ticket into exactly one label: L1, L2, L3. If none fits, use "other". Respond with only JSON: {"label": "...", "reason": "..."}
3</instructions>
4<examples>
5<example><ticket>TEXT</ticket><label>LABEL</label></example>
6</examples>
7<ticket>TICKET</ticket>Escape all example and ticket text with html.escape(text, quote=False).
- Acme Bikes
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Step 2: grade the replies
Each line is id|expected|reply. Extract the JSON object from the reply (first { to last }) and read its label.
- If there’s no parsable JSON with a label, print
ID: could not parse reply. - If the label is wrong, print
ID: expected E, got G.
Then print accuracy: C/N (P%), parse failures: F and, if there were wrong labels, most confused: E -> G (COUNT) for the most frequent pair (alphabetically first on ties).
- Six replies
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…