Name one complete job
Write the job as an input, an action and an output. For example: an operations user receives a document, reviews extracted fields and approves a draft CRM update. That is more testable than ‘an AI assistant for operations’.
Define where the task ends. Preparing a draft and writing a live customer record are different scopes. A pilot can prove the first without silently taking on the second. Our Document OCR work is an example of extraction feeding a reviewable workflow; the following scoping example is illustrative, not a new client case.
The brief before the build
| Decision | What to write down |
|---|---|
| User and task | Who performs the work, how often, and which step is slow or unreliable. |
| Inputs | Document types, source systems, expected languages and examples of difficult inputs. |
| Output | The exact fields, answer format or proposed action the user needs. |
| Access | Which systems may be read, which may be changed and who approves the access. |
| Review | Who checks exceptions and what requires approval before execution. |
| Success | The baseline, test set, pass conditions and reasons to stop. |
Keep an evaluation set outside the demo
Choose examples that represent the actual workload before tuning the system. Include incomplete documents, conflicting records, ambiguous requests and unavailable integrations. Hold back some examples from development so the final review tests more than familiarity with a handful of prompts.
Agree how correctness is judged and who resolves disagreement. For extraction, measure required fields individually. For retrieval, check whether the answer is supported by the allowed sources. For tool use, check the requested action, arguments and approval behavior.
Record the model and prompt version, input set, errors, review effort, latency and usage cost for each evaluation. A system that produces an attractive answer but needs extensive correction may not improve the workflow.
Define what the pilot can do when it is wrong
- Missing or contradictory information enters review instead of being filled in by guesswork.
- An unavailable integration produces a visible incomplete state, not a success message.
- The model cannot grant itself new tools or permissions based on instructions inside a document.
- Consequential writes require the approval agreed for the workflow.
- A repeated request should not create duplicate business actions; test recovery as well as the happy path.
Estimate the cost of operating the whole workflow
Model usage is only one line. Include retrieval, storage, background processing, provider subscriptions, human review and ongoing maintenance. Compare cost per completed task, not just cost per model request.
A longer context window, repeated tool calls or a high review rate can change the economics. Record actual pilot usage and reviewer time before projecting a monthly figure. The scope brief includes fields for these assumptions without prescribing a budget.
End with a decision
The review should say which tasks passed, which still need human handling and what a limited production rollout would require. The next step may be a narrow release, a further evaluation cycle or stopping because an existing tool is sufficient.
Download the brief, fill in the parts you know and leave uncertainties visible. A useful discovery conversation resolves those uncertainties rather than hiding them behind a long feature list.
