Learner track
The hands-on lessons you work through during the session.
This is the Learner track: the hands-on lessons you work through during the workshop. Each module builds on the last, taking the Agent Arena NL→SQL chatbot through one flow: select the base model → measure offline → release and detect → investigate with a human → prove and monitor the improvement — all instrumented with Langfuse from the very first module, served over OpenRouter, and backed by ClickHouse.
The arc starts with the decision that matters most when building a serious agent: which model to use. Module 01 answers it with evidence — the contest, run as Langfuse experiments and ranked by cost per correct answer — and every module after it builds on that choice.
Work through the modules in order:
- 00 Setup — connect OpenRouter, ClickHouse Cloud, and Langfuse Cloud, and seed the database.
- 01 Select the base model — the foundational decision: run the model × prompt grid as Langfuse experiments and crown a winner by cost per correct answer.
- 02 Measure offline — go beyond "it won": read per-tier
accuracy, the
agent-arena-llm-judgescore, and individual traces to see how the winner really performs. - 03 Release and detect — release the selected configuration with a stale business policy, then see an operational evaluator pass while a real user's 👎 exposes the blind spot.
- 04 Investigate — use the feedback signal to find the authoritative production trace, complete a human annotation, verify the correction, and record its provenance.
- 05 Close the loop — promote the reviewed incident, prove the policy improvement with paired experiments, calibrate and enable a general online evaluator, then monitor future traffic for both scores and feedback.
See the top-level overview for the full module table, including the matching Instructor-track notes for each module.