Modules
Take single modules or the whole track. Two hour sessions, combined into half days where it suits your calendar.
Testing Non-Deterministic Systems
How the testing pyramid changes when outputs are probabilistic, not fixed.
- Why non-determinism breaks classic testing assumptions
- The testing pyramid for AI applications
- Unit testing LLM integrations
- Mocking APIs and testing prompt templates
Evaluation Frameworks and Metrics
Measure relevance, coherence and factuality with custom evaluators.
- Relevance, coherence and factuality metrics
- Writing custom evaluators
- Integration testing against real LLMs
- Snapshot testing and multi-turn conversations
Testing RAG and Agentic Systems
Retrieval quality, embeddings and the harder cases of tool-calling agents.
- Testing RAG retrieval quality
- Evaluating embeddings
- Testing tool-calling and agentic systems
- Production observability for AI features
Teams Often Pair This With
Related modules other teams add alongside this track.
Red-Green-Refactor in Practice
The test-first rhythm applied to single objects and clusters of objects.
View module Applied AI Engineering · Building GenAI ApplicationsTool Calling and MCP Integration
Extend an LLM with external functions and wire in the Model Context Protocol.
View moduleEnquire About This Training
Tell us which modules interest you and we will propose a plan within 24 hours. Choosing Enquire on a module above fills in the subject for you. Add more than one if you like.