
A prospective customer submits a quote request. Your AI assistant summarises it, says it has forwarded the enquiry to sales, and displays a reassuring success message. The next morning, the sales representative opens the CRM. There is no record. The automation sounded convincing, but the work was never completed.
Agencies scaling their use of AI need to answer a practical question: What evidence lets us say that a workflow succeeded? Choosing a capable model matters. So does knowing whether the intended business outcome actually exists. This guide proposes an acceptance process for an illustrative workflow that classifies incoming enquiries and transfers them to a CRM.
What the source establishes, and what we add
Anthropic’s agent evaluation guide, published on 9 January 2026, distinguishes an agent’s response from the actual outcome in its environment. It discusses repeated trials for variable outputs and different evaluation methods. This is an engineering practice article, not a peer-reviewed marketing experiment. We are not presenting it as a September announcement.
The acceptance card, scenario mix, agency example and delivery plan below are our proposed application. Acceptance conditions should be agreed with the person responsible for the business process. A universal threshold borrowed from another project will not establish whether your own workflow is ready. For an agency, a good output must also be usable by sales, traceable and correctable when something goes wrong.
Write down the workflow’s promise first
“We use AI for lead management” is not a testable specification. Without a clear connection between input and outcome, different teams can celebrate different kinds of success. The account manager might expect a readable summary, sales might expect an assigned task, and engineering might simply expect the connection to return without an error.
For our illustrative project, the promise is: “A valid quote request becomes one CRM record with its source information preserved, is assigned to the appropriate team, and appears in a visible review queue if the workflow cannot resolve or complete it.” Not every message should become a sales opportunity. A job application, an existing customer’s support request and a new sales enquiry are different kinds of work.
| Field | Illustrative decision | Evidence |
|---|---|---|
| Input | A website quote request | Submission ID and receipt time |
| Outcome | One record, correct team and source | CRM record ID and field values |
| Uncertainty | Review when sales and support intent are unclear | Reason and assigned reviewer |
| Failure | No success claim for an incomplete transfer | Error status and retry record |
| Boundary | No automatic price or discount promises | Output and action history |
The person implementing the automation should not fill in this card alone. Sales must explain which team should receive which request. Account management should specify where campaign attribution belongs. Automating an ambiguous process can make the wrong work happen faster, even when the integration itself runs correctly.
Inspect the message, the action and the outcome
Ask three separate questions. What was the customer told? What action did the system attempt? What was actually created in the business application? This helps attach a failure to something the team can fix. A classification error and an unavailable CRM require different responses.
One enquiry, three checkpoints
- 01What it says
Does the response accurately describe the request and the work performed?
- 02What it attempts
Was the action sent to the right system with the right field values?
- 03What exists
Are the expected record, assignment and source really present in the CRM?
DijitalPi’s illustrative acceptance framework. Passing one checkpoint does not establish that the others passed.
A record might exist while its campaign source remains blank. Sales can partly proceed, but marketing reporting has lost information. Checking only that something arrived in the CRM misses this problem. Filling the gap with a campaign name guessed by the model is not a remedy. An explicitly unknown source is more useful than false certainty.
How should you choose the first 20 scenarios?
Do not populate the initial set exclusively with easy enquiries. Ask the team about situations that previously needed manual intervention. Remove private information or create fictional inputs that preserve the important business distinctions. Twenty is an organising choice for this starter worksheet, not a statistically sufficient sample or a reliability guarantee.
Our proposed mix contains five clear requests, five incomplete or mixed requests, four duplicates or changes, four connection or execution failures, and two authority boundaries. This covers both language interpretation and difficult operational situations. Write the expected outcome before running each scenario. Do not redefine success afterwards to match the output.
Download the 20-scenario AI workflow worksheet (CSV)
The file includes a scenario, fictional input, starting conditions, expected outcome and evidence to inspect. Actual result, reviewer and decision fields are intentionally empty. Complete them using your own trials. Open the file in a spreadsheet application and duplicate each scenario row for additional attempts. Record model version, instruction version and trial ID so comparisons can be traced.
An illustrative failure review: the same request arrives twice
The form takes too long to acknowledge submission, so the visitor presses the button again. The workflow preserves a submission identifier that connects both events to the same request. Test the events sequentially, then close together if your test environment allows it. The expected outcome is one record under the agreed duplicate rule, rather than two sales opportunities.
This does not mean merging everything that shares an email address. The same person might request a different service next week. Your acceptance card must distinguish a repeated event from a genuinely new need. That is why these cases have separate rows in the worksheet.
When a test fails, avoid closing the review with “the AI got it wrong”. Describe the observed behaviour: “Processing a second event with the same submission identifier created a new opportunity.” Then record its business effect. Two representatives might contact the same person; conversion counts might be inflated; reporting might imply success that did not occur. Engineering now has a concrete question to investigate: which duplicate check was missing when the event was processed again?
What can an overall pass rate hide?
Imagine an illustrative first run in which 18 of 20 scenarios pass. That 90% describes this scenario set in that run. It does not establish that 90% of real enquiries will be handled correctly. If the remaining failures silently lose a request and promise an unauthorised price, the overall score may not justify handover.
| Category | Example | Proposed response |
|---|---|---|
| Critical | Wrong customer record, lost enquiry or unauthorised action | Fix and rerun the relevant scenarios |
| Functional | Wrong team or missing campaign source | Assess impact against the agreed handover conditions |
| Editorial | An accurate but unnecessarily long internal summary | Track improvement separately from the business outcome |
Use the same inputs and starting conditions to compare versions. Testing the old and new versions on different customer types makes it hard to tell whether the difference comes from the model or the examples. Repeat scenarios to expose variable behaviour and report the number of attempts. Alongside examples used throughout development, retain separate examples that are only introduced during the final check.
Human handover is an outcome, not a waiting room
“Send for human review” needs acceptance conditions too. Where will the request appear? Who will see it? What response time is the team aiming for? How will the system recognise that the review is complete? Adding a label does not necessarily transfer ownership of the work.
Consider a message containing both a sales enquiry and a support problem. Create a review task containing the original message, the reason for uncertainty and the actions already completed. The reviewer should not have to collect the same information again. An unassigned task or a queue nobody monitors should count as a failed handover. For this scenario, an appropriate human handover can be a successful outcome.
A practical one-week work plan
On day one, agree on a single workflow and its acceptance card. On day two, have sales, account management and engineering review the scenarios together. On day three, run trials in a test environment and gather outcome evidence without contacting real customers or altering live records. On day four, address failures according to their impact and rerun the corrected version under the same conditions.
On day five, discuss critical failures, handover behaviour and unresolved cases before looking at the overall score. This five-day sequence is a planning example, not a promise that every integration can be completed within a week. Access, data quality and unclear business rules may need to be resolved first.
The handover pack should contain the acceptance card, version information, completed worksheet, open issues and the process owner’s decision. After launch, continue monitoring correction time, decisions reversed by people and incomplete operations that receive no attention. Rerun relevant tests when the model, instructions or CRM mapping changes. The value of an automation depends on understanding the conditions under which it can be trusted.



