Blog/Blog
BLOG8 min read

Does Your AI Workflow Work, or Just Say It Does?

DijitalPi
The DijitalPi Team
29 September 2026
Does Your AI Workflow Work, or Just Say It Does?

A prospective customer submits a quote request. Your AI assistant summarises it, says it has forwarded the enquiry to sales, and displays a reassuring success message. The next morning, the sales representative opens the CRM. There is no record. The automation sounded convincing, but the work was never completed.

Agencies scaling their use of AI need to answer a practical question: What evidence lets us say that a workflow succeeded? Choosing a capable model matters. So does knowing whether the intended business outcome actually exists. This guide proposes an acceptance process for an illustrative workflow that classifies incoming enquiries and transfers them to a CRM.

What the source establishes, and what we add

Anthropic’s agent evaluation guide, published on 9 January 2026, distinguishes an agent’s response from the actual outcome in its environment. It discusses repeated trials for variable outputs and different evaluation methods. This is an engineering practice article, not a peer-reviewed marketing experiment. We are not presenting it as a September announcement.

The acceptance card, scenario mix, agency example and delivery plan below are our proposed application. Acceptance conditions should be agreed with the person responsible for the business process. A universal threshold borrowed from another project will not establish whether your own workflow is ready. For an agency, a good output must also be usable by sales, traceable and correctable when something goes wrong.

Write down the workflow’s promise first

“We use AI for lead management” is not a testable specification. Without a clear connection between input and outcome, different teams can celebrate different kinds of success. The account manager might expect a readable summary, sales might expect an assigned task, and engineering might simply expect the connection to return without an error.

For our illustrative project, the promise is: “A valid quote request becomes one CRM record with its source information preserved, is assigned to the appropriate team, and appears in a visible review queue if the workflow cannot resolve or complete it.” Not every message should become a sales opportunity. A job application, an existing customer’s support request and a new sales enquiry are different kinds of work.

An acceptance card to complete with the process owner
FieldIllustrative decisionEvidence
InputA website quote requestSubmission ID and receipt time
OutcomeOne record, correct team and sourceCRM record ID and field values
UncertaintyReview when sales and support intent are unclearReason and assigned reviewer
FailureNo success claim for an incomplete transferError status and retry record
BoundaryNo automatic price or discount promisesOutput and action history

The person implementing the automation should not fill in this card alone. Sales must explain which team should receive which request. Account management should specify where campaign attribution belongs. Automating an ambiguous process can make the wrong work happen faster, even when the integration itself runs correctly.

Inspect the message, the action and the outcome

Ask three separate questions. What was the customer told? What action did the system attempt? What was actually created in the business application? This helps attach a failure to something the team can fix. A classification error and an unavailable CRM require different responses.

One enquiry, three checkpoints

  1. 01What it says

    Does the response accurately describe the request and the work performed?

  2. 02What it attempts

    Was the action sent to the right system with the right field values?

  3. 03What exists

    Are the expected record, assignment and source really present in the CRM?

DijitalPi’s illustrative acceptance framework. Passing one checkpoint does not establish that the others passed.

A record might exist while its campaign source remains blank. Sales can partly proceed, but marketing reporting has lost information. Checking only that something arrived in the CRM misses this problem. Filling the gap with a campaign name guessed by the model is not a remedy. An explicitly unknown source is more useful than false certainty.

How should you choose the first 20 scenarios?

Do not populate the initial set exclusively with easy enquiries. Ask the team about situations that previously needed manual intervention. Remove private information or create fictional inputs that preserve the important business distinctions. Twenty is an organising choice for this starter worksheet, not a statistically sufficient sample or a reliability guarantee.

Our proposed mix contains five clear requests, five incomplete or mixed requests, four duplicates or changes, four connection or execution failures, and two authority boundaries. This covers both language interpretation and difficult operational situations. Write the expected outcome before running each scenario. Do not redefine success afterwards to match the output.

Download the 20-scenario AI workflow worksheet (CSV)

The file includes a scenario, fictional input, starting conditions, expected outcome and evidence to inspect. Actual result, reviewer and decision fields are intentionally empty. Complete them using your own trials. Open the file in a spreadsheet application and duplicate each scenario row for additional attempts. Record model version, instruction version and trial ID so comparisons can be traced.

An illustrative failure review: the same request arrives twice

The form takes too long to acknowledge submission, so the visitor presses the button again. The workflow preserves a submission identifier that connects both events to the same request. Test the events sequentially, then close together if your test environment allows it. The expected outcome is one record under the agreed duplicate rule, rather than two sales opportunities.

This does not mean merging everything that shares an email address. The same person might request a different service next week. Your acceptance card must distinguish a repeated event from a genuinely new need. That is why these cases have separate rows in the worksheet.

When a test fails, avoid closing the review with “the AI got it wrong”. Describe the observed behaviour: “Processing a second event with the same submission identifier created a new opportunity.” Then record its business effect. Two representatives might contact the same person; conversion counts might be inflated; reporting might imply success that did not occur. Engineering now has a concrete question to investigate: which duplicate check was missing when the event was processed again?

What can an overall pass rate hide?

Imagine an illustrative first run in which 18 of 20 scenarios pass. That 90% describes this scenario set in that run. It does not establish that 90% of real enquiries will be handled correctly. If the remaining failures silently lose a request and promise an unauthorised price, the overall score may not justify handover.

Proposed failure categories: agree acceptance conditions for each project
CategoryExampleProposed response
CriticalWrong customer record, lost enquiry or unauthorised actionFix and rerun the relevant scenarios
FunctionalWrong team or missing campaign sourceAssess impact against the agreed handover conditions
EditorialAn accurate but unnecessarily long internal summaryTrack improvement separately from the business outcome

Use the same inputs and starting conditions to compare versions. Testing the old and new versions on different customer types makes it hard to tell whether the difference comes from the model or the examples. Repeat scenarios to expose variable behaviour and report the number of attempts. Alongside examples used throughout development, retain separate examples that are only introduced during the final check.

Human handover is an outcome, not a waiting room

“Send for human review” needs acceptance conditions too. Where will the request appear? Who will see it? What response time is the team aiming for? How will the system recognise that the review is complete? Adding a label does not necessarily transfer ownership of the work.

Consider a message containing both a sales enquiry and a support problem. Create a review task containing the original message, the reason for uncertainty and the actions already completed. The reviewer should not have to collect the same information again. An unassigned task or a queue nobody monitors should count as a failed handover. For this scenario, an appropriate human handover can be a successful outcome.

A practical one-week work plan

On day one, agree on a single workflow and its acceptance card. On day two, have sales, account management and engineering review the scenarios together. On day three, run trials in a test environment and gather outcome evidence without contacting real customers or altering live records. On day four, address failures according to their impact and rerun the corrected version under the same conditions.

On day five, discuss critical failures, handover behaviour and unresolved cases before looking at the overall score. This five-day sequence is a planning example, not a promise that every integration can be completed within a week. Access, data quality and unclear business rules may need to be resolved first.

The handover pack should contain the acceptance card, version information, completed worksheet, open issues and the process owner’s decision. After launch, continue monitoring correction time, decisions reversed by people and incomplete operations that receive no attention. Rerun relevant tests when the model, instructions or CRM mapping changes. The value of an automation depends on understanding the conditions under which it can be trusted.

Related reading: Turning production savings into better experiments and the DijitalPi blog. Bring a completed acceptance card and a difficult scenario to your next workflow review so the discussion starts with a concrete decision.

FAQ

Frequently Asked Questions

Do 20 tests prove that an AI workflow is reliable?

No. These 20 scenarios are a starter template. Real workload patterns, error impact, repeated trials and observations after launch must extend the evaluation. A scenario pass rate is not the same as production reliability.

Do we need another AI model to assess the tests?

No. Record creation, field values and assignment can be checked directly in the system. An authorised team member can assess wording and business suitability. Using this worksheet requires no external model call or paid API.

Which changes should trigger another test run?

Check affected scenarios whenever the model, instructions, data source, tool connection or CRM field mapping changes. Keep previously corrected critical failures in the repeat test set as well.
A quick summary with AI
You can have this content summarised by the AI of your choice, or copy the prompt.
SHARE
inXf

Related articles

When AI Writes for Everyone, Why Should Customers Remember Your Brand?
27 September 2026

When AI Writes for Everyone, Why Should Customers Remember Your Brand?

Read →
Marketing Budgets in the AI Era: More Content or Better Experiments?
27 September 2026

Marketing Budgets in the AI Era: More Content or Better Experiments?

Read →
What Is GEO? Generative Engine Optimization Explained
1 August 2026

What Is GEO? Generative Engine Optimization Explained

Read →

Let us apply these strategies to your brand.

Free Consultation →

ASK AI ABOUT DIJITALPI

Let an AI explain what DijitalPi does.

Opens your chosen assistant with a ready research prompt. It reads the site live and answers.

Message us on WhatsApp