Skip to the review
DijitalPi
TREN
Contact

INTERNATIONAL RESEARCH / PREPRINT

Is an AI employee really ready for work?

An AI system can complete the task, but how much review will your team still need? Consider accuracy, review effort and total cost together when choosing an AI solution.

ACADEMIC WORK REVIEWED

READY or Not: Reliable Enterprise Agent Deployment

ResearchersVeronica Chatrath, Bryan Zhu, Jingxuan Fan, George Pu, Soham Dinesh Tiwari, Soham Dan, Ryan Young, Yuan (Christy) Li, Yuang Yao, Apaar Shanker, Minglai Yang, Daniel Yue Zhang, Yunzhong He, Ying Liu, Chenguang Wang, Zhijun Yin and Yuan (Emily) Xue.

arXiv · September 2, 2026 · v1 · Preprint · Source language: English

Open the original publication
ENGLISH REVIEW AND COMMENTARY: DIJITALPI

We explain the paper in plain English, discuss what it may mean for organisations and identify our own examples separately.

English review: 9 September 2026 · The source date appears in the citation above.

Editorial review draft

The source has been checked; the review still carries its original editorial status. DijitalPi did not conduct a customer experiment for this paper.

LET'S READ THE RESEARCH TOGETHER

We first explain the researchers' question, method and findings. We then discuss how to interpret the results, clearly separating DijitalPi's commentary from the source.

01 / WHAT DID THE RESEARCHERS WANT TO UNDERSTAND?

If an AI system can do the job, is it ready for use?

With READY, Chatrath and colleagues propose evaluating an AI system within specific business, reliability goal, human review, and cost conditions.

Explaining the context · DijitalPi commentary

It's gratifying to see a correct result in one try. However, other work can be done before that output can be sent to the customer: finding the missing information, checking the suspicious part, correcting it and giving final approval. The question of job readiness becomes meaningful when considering how to complete these steps.

In our example, let's assume that two assistants prepare the same number of correct quotes. In the first one, each proposal is read from start to finish. In the latter, areas requiring control are found more easily. However, if we write the correct number of offers, we cannot see the difference in the report that the team will experience between these two orders. This example is not the clinical case of the study.

02 / HOW WAS THE RESEARCH CONDUCTED?

How did they test the question?

The study examines 16 AI systems using a sample of 750 clinical-record audits. The rule for accepting a result or sending it to human review is selected on development data and tested with separate examples. The primary calculation assumes a 90% success rate for human review; that workflow was not actually operated in every case.

Understanding the method · DijitalPi commentary

Testing a rule on examples you have established is different from trying it on new examples. You can get used to the files when you make constant adjustments to the same files. A separate testing section is to examine whether the layout you choose works on unseen samples.

Let's also unpack the human control assumption. Putting a success rate into the account doesn't mean that rate is measured on your team. The knowledge of the person checking, the time in hand, and the clarity of the document may vary. Therefore, when reading such an account, we distinguish which input comes from experiment and which from accepted assumption.

03 / RESEARCH FINDINGS

Similar accuracy can be combined with different burdens of human review.

In the article, the accuracy of GPT-5.4 alone is 72.8%, and Sonnet 5 is 72.5%. At the 76% reliability target determined in the study, the rates of referral to human review are 39.2% and 29.6%, respectively. This comparison depends on the conditions of the study and the assumption of human review.

GPT-5.4
39.2%The share directed to humans under the target and assumptions of the study; accuracy alone is 72.8%.
sonnet 5
29.6%In the same comparison, the share directed to humans; accuracy alone is 72.5%.

Source: READY v1: methodology, experimental results and human review assumptions ↗

04 / DIJITALPI'S EXPLANATION

How should we interpret these findings?

The referral rate allows us to consider how much work can come to the team. But it is not time itself. The review time for ten short files and ten complex files may not be the same. In order to calculate human effort, it is necessary to know the duration and difficulty of the control, as well as the number of files.

Sending fewer files to people is not a measure of success on its own. If the system accepts problematic output without detecting it, the queue may shrink while the error moves elsewhere. First define an acceptable outcome, then evaluate the work required to achieve it. The article's 76% figure is the selected target for that experiment, not a recommended threshold for every business.

The question we take from this for businesses is: Once AI is in place, when is the work really done? Looking at the vehicle fee is a start. If preparation, control, correction and waiting records are kept, we can see more clearly where the gain is made and where it is lost.

05 / CONCLUSION AND OPEN QUESTIONS

What did we learn, and what do we still not know?

READY recommends reading the model score alongside the entire workflow. The preprint's results are limited to specific clinical tasks and assumptions; they are not a certificate of general readiness.

How effective is human control in your proposal, application or support process and how long does it take? The answer to this cannot be taken from the table in the article. The examples below are designed to help you think about how you can make this information visible in your own little experiment.

DIJITALPI'S APPLICATION COMMENTARY

What could this look like in your organisation?

We created these scenarios to make the topic concrete. They are not cases from the paper or measured client results.

EXAMPLE 01

Fast quote, long check

The assistant prepared a proposal in a minute. The salesperson spent twenty minutes checking the products, prices and conditions one by one.

Can you just write the preparation time as success?

Open the recommendation for example 01

Measure the job completion time.

Record the time from the start to the final version ready for the customer. Track the employee's active work separately. Compare it with the time required for the same type of proposal before AI, including the review period.

EXAMPLE 02

Every registration comes for approval

The results of the assistant sorting the applications appear to be mostly accurate. However, almost every application awaits approval from the team manager.

Is the number of correct answers alone sufficient?

Open the recommendation for example 02

Make the approval queue visible as well.

Count how many records came for approval, how long they waited, and how many were corrected. Before removing the control, determine under what condition it is required. Sending everything to the same person can move the bottleneck elsewhere.

EXAMPLE 03

Problem-free in the trial, different in the new month

The assistant was tested with old products. The next month, the price structure and campaign conditions changed, but the team continued to rely on the old trial result.

Does the previous test cover the new conditions?

Open the recommendation for example 03

Retest the task under the changed conditions.

Record which prices, documents, models and rules you tested. Rerun the cases affected by the change. Keep the old result and add the date and difference of the new trial.

TRY IT WITH YOUR TEAM

Follow a job from start to finish.

Choose one recurring task from your team. Record the start, first AI output, human review, correction and completion times. This log helps reveal where time is spent, but it is not a scientific measure of success on its own.

Check out AI solutions that fit your business

Source and ownership

Veronica Chatrath, Bryan Zhu, Jingxuan Fan, George Pu, Soham Dinesh Tiwari, Soham Dan, Ryan Young, Yuan (Christy) Li, Yuang Yao, Apaar Shanker, Minglai Yang, Daniel Yue Zhang, Yunzhong He, Ying Liu, Chenguang Wang, Zhijun Yin and Yuan (Emily) Xue.
ScaleAI; University of California, Santa Cruz; Vanderbilt University Medical Center.

arXiv · September 2, 2026 · v1 · Preprint. Open the original publication

The academic work belongs to the researchers named above. This page contains DijitalPi's explanatory review and original business examples; it is not a full translation of the paper.

← Return to all research articles

ASK AI ABOUT DIJITALPI

Let an AI explain what DijitalPi does.

Opens your chosen assistant with a ready research prompt. It reads the site live and answers.