Loading
0x80Lesson 9 of 9

Capstone: build and evaluate a triage prompt

Render a production-style prompt from a config, then grade a batch of model replies and report what to fix.

30 min 5-question quiz 2 code exercises
By the end of this lesson you can
  • Render a complete prompt from a versioned configuration
  • Parse replies tolerantly and grade them against expected labels
  • Turn results into a short report that points at the next fix

Put it all together on a realistic task: triaging support tickets into labels. A production prompt isn’t a hand-typed paragraph - it’s rendered from a configuration you can review and version: product name, label set, balanced examples, and the escaped ticket.

Then you evaluate: run the prompt over a labeled test set (here, the replies are given), parse every reply tolerantly, and report accuracy, parse failures and the most common confusion - which tells you what to change next (a clearer label definition? an example of the confusing case?).

confusions.py
from collections import Counter
mistakes = [("technical", "billing"), ("technical", "billing"), ("shipping", "billing")]
print(Counter(mistakes).most_common(1))
Output
[(('technical', 'billing'), 2)]

Key takeaways

  • Render prompts from reviewed, versioned configuration.

  • Parse replies tolerantly and count parse failures separately from wrong answers.

  • The most common confusion points to the next prompt change.

Lesson quiz

5 questions · pass with 4 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Step 1: render the triage prompt

+25 XP

Read one JSON config with product, labels, examples (pairs of ticket and label) and ticket. Print exactly:

1<instructions>
2You triage support tickets for PRODUCT. Classify the ticket into exactly one label: L1, L2, L3. If none fits, use "other". Respond with only JSON: {"label": "...", "reason": "..."}
3</instructions>
4<examples>
5<example><ticket>TEXT</ticket><label>LABEL</label></example>
6</examples>
7<ticket>TICKET</ticket>

Escape all example and ticket text with html.escape(text, quote=False).

  • Acme Bikes
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Step 2: grade the replies

+25 XP

Each line is id|expected|reply. Extract the JSON object from the reply (first { to last }) and read its label.

  • If there’s no parsable JSON with a label, print ID: could not parse reply.
  • If the label is wrong, print ID: expected E, got G.

Then print accuracy: C/N (P%), parse failures: F and, if there were wrong labels, most confused: E -> G (COUNT) for the most frequent pair (alphabetically first on ties).

  • Six replies
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Did you like the lesson? 😆👍
Consider a donation to support our work: