AI evaluation and guardrails
Evaluate AI features, protect permissions, and enforce task budgets.
Keep AI suggestions useful, authorized and bounded
A document assistant must answer from documents the caller may read. A support agent may propose an action without being allowed to approve it. An extraction pipeline may need human review instead of repeatedly retrying malformed data.
Use the evaluation primer and local fixture first. The workbench then supplies four complete reference workflows and separate AWS adapters. A deterministic fixture proves state transitions, not real-model quality. Evaluate actual model behavior separately when you connect a provider.
Parts group related chapters. Each lesson has a chapter.lesson address, such as 4.07. Open a title below, or use Next to follow the reading sequence. Within a lesson, On this page lists its sections.
- 18.01
Evaluate an AI feature and enforce task-level limits
Concepts and examples
- 18.02
Evaluate tag suggestions without hiding rare failures or outage cost
Concepts and examples
- 18.03
Draft support replies without granting tool authority
Concepts and examples
- 18.04
Serve recommendations with safe fallback ranking
Concepts and examples
- 18.05
Build a document assistant with current permissions
Concepts and examples
- 18.06
Build and deploy the AI project workbench
Concepts and examples
- 18.07
- 18.08
- 18.09
- 18.10
- 18.11
- 18.12
- 18.13
Stage 4: Add optional AI tag suggestions
Projects and practical assessment