You are upgrading the shopping assistant in your app. The new version gives more detailed answers, and you expect it to help customers weigh their options. Your team tracks chat usage, but you want to know what happens to purchases. Does an assistant that explains more also sell more? In an SSRN working paper that has not undergone peer review, researchers found that users randomly assigned to the reasoning assistant on China’s Ctrip platform made 2.5% fewer hotel bookings than those assigned to the existing version. DijitalPi commentary: Track chat usage, product search and completed orders separately. The experiment compared two AI versions, with no assistant-free group. It does not establish the same outcome for Türkiye or Turkish-language services.
ACADEMIC WORK REVIEWED
Reasoning AI, Consumer Search, and Purchase: Evidence from an Online Platform Field Experiment
English review: 11 October 2026 · The source date appears in the citation above.
Editorial review draft
The source has been checked; the review still carries its original editorial status. DijitalPi did not conduct a customer experiment for this paper.
LET'S READ THE RESEARCH TOGETHER
We first explain the researchers' question, method and findings. We then discuss how to interpret the results, clearly separating DijitalPi's commentary from the source.
01 / WHAT DID THE RESEARCHERS WANT TO UNDERSTAND?
How does a shopping assistant that reasons more extensively change customers’ searches and purchases?
The researchers examined how a model that uses more deliberation before answering affects consumer behaviour. They investigated whether reduced searching reflected slower responses or answers that provided enough information to make further searching seem unnecessary.
Explaining the context · DijitalPi commentary
DijitalPi commentary: Treating answer quality and commercial performance as the same measure can mislead an assistant upgrade decision. A more comprehensive explanation may appear helpful while also changing how customers compare products. Evaluating a version should therefore include the search and purchase steps beyond the chat screen.
02 / HOW WAS THE RESEARCH CONDUCTED?
How did they test the question?
A user-level randomised field experiment ran for 8 days, from 3–10 April 2025, in the mobile app of Ctrip, an online travel platform in China. Of 514,402 users who entered the Wendao assistant interface, 497,565 remained after exclusions. The control group contained 248,581 users assigned to the existing DeepSeek-V3 version, while 248,984 users were assigned by default to the DeepSeek-R1 reasoning model. Users in the latter group could switch back to the existing model using a button.
Understanding the method · DijitalPi commentary
The comparison group also had an AI assistant. Only 6.89% of users entering the assistant interface chatted at least once. The main analysis kept users in their assigned groups regardless of whether they chatted, measuring the effect of assignment to the new version rather than an effect specific to active chat users. Total search queries, listing browsing, clicks and hotel orders were tracked over 8 days. Detailed search data were available only for hotels.
The researchers excluded users in the top 1% of search, browsing or click activity and the top 0.1% of booking activity; results were qualitatively similar under alternative thresholds. The main reported percentages represent changes in expected activity counts relative to the control group. To investigate possible explanations, the analysis also included chat time and the amount of information in the answers. This additional analysis did not directly measure users’ beliefs.
03 / RESEARCH FINDINGS
Assignment to the new version reduced hotel searching and bookings.
The researchers found that, compared with the existing DeepSeek-V3 group, the DeepSeek-R1 group made 0.8% fewer hotel search queries, browsed 0.9% fewer listings, clicked on 2.6% fewer hotels and made 2.5% fewer hotel bookings; all four results were statistically significant. These are percentage changes in activity counts, not percentage points. Among users progressing through the search process, browsing per search, click-through rates and conversion rates showed no significant change, although these conditional comparisons were descriptive. Accounting for search, browsing and clicks reduced the estimated order decline to 1.1%, which was no longer statistically significant. No effect was detected on flight orders or orders for other travel products, and orders for hotels recommended within the chat did not increase. Across the three months after the experiment, users initially assigned to the new version made 11.0% more chat requests. Since the new assistant became broadly available during follow-up, this reflects persistence from initial exposure rather than continued differences in access to the two versions.
Decline in hotel booking counts
2.5%A significant decline relative to the existing assistant group over the 8-day experiment, covering all assigned users rather than only those who chatted.
Increase in total post-experiment chat requests
11.0%A result reflecting persistence from initial assignment across post-experiment months 1–3, when the new assistant was broadly available.
DijitalPi commentary: Greater assistant usage may not mean better performance against a purchasing objective. When testing a new version, examine where searching declines alongside what happens to orders. Concise answers with optional detail could be tested separately; this experiment did not establish which answer format produces better sales outcomes.
The reasoning version was slower, but users did not spend less total time searching within the hotel channel. Accounting for time left the main effects largely unchanged, whereas accounting for the information in answers made the search effects statistically insignificant. The authors consider this pattern more consistent with answers seeming sufficiently complete to reduce the perceived need for further searching. This is an interpretation of the mechanism, not a direct measurement of users’ beliefs, and the separate effects of explanation, structure and length were not identified.
05 / CONCLUSION AND OPEN QUESTIONS
What did we learn, and what do we still not know?
This working paper, which has not undergone peer review, shows that changing an assistant version on a single Chinese platform could reduce hotel searching and booking counts in the short term. A change in booking counts cannot be read as a change in revenue, and the same effect was not found for flights. The order effect remained negative in the first two follow-up months, then became small and statistically insignificant. The three-month follow-up does not replace a longer experiment with continued randomised access to different versions. This design does not establish applicability to Türkiye, Turkish-language interactions or other sectors. The authors report no funding and declare no relevant competing interests; their affiliations list no company employees. However, the data and experiment depend on Ctrip, and exclusion decisions drew on internal discussions with the platform. The text does not explicitly identify who designed and conducted the experiment.
Our open question at DijitalPi: how do more detailed answers from your Turkish-language assistant affect product exploration, and which purchasing measure will you use to evaluate that separately from chat volume?
We created these scenarios to make the topic concrete. They are not cases from the paper or measured client results.
EXAMPLE 01
Tourism: room searches after detailed answers
An illustrative accommodation business wants to evaluate a new version of its assistant for explaining room options.
Do customers explore rooms and book after reading the answer?
Open the recommendation for example 01+
Compare assistant usage alongside room searches and bookings.
Plan a comparison of the existing and new versions with random assignment of users. Record room searches, listing views, room clicks and completed bookings over the same observation period for each group. Keep all assigned users in the main evaluation, including those who never chat. Do not use Ctrip’s 2.5% decline as this business’s expected result.
EXAMPLE 02
E-commerce: comprehensive product recommendations
An illustrative electronics retailer is considering an assistant version that compares products through lengthy explanations.
Are more comprehensive recommendations replacing product exploration and orders?
Open the recommendation for example 02+
Evaluate answer format in a separate experiment.
Design a test within the same assistant version comparing concise answers with optional detail against lengthy answers. Record visits to product listings, product clicks and completed orders separately. This would investigate answer length; the source study did not isolate its effect from other differences between versions. Do not transfer the hotel finding directly to electronics.
EXAMPLE 03
B2B software: more chats, uncertain quotation demand
An illustrative B2B software company plans to upgrade its assistant for choosing subscription packages.
Does more chat activity mean more requests for a quotation?
Open the recommendation for example 03+
Define chat requests, package exploration and quotation requests as separate outcomes.
Track visits to package pages and submitted quotation requests for users assigned to the existing and new assistant versions. If chat usage persists later, also record which version each group can access. Do not count a quotation request as a signed contract. The study did not test B2B purchasing processes, so do not assume the direction of the result.
TRY IT WITH YOUR TEAM
Define the assistant’s success measure.
Define separate measures for chat usage, product search and completed orders in a comparison of the existing and new versions. Specify user assignment and the observation period in advance; this is a proposed pilot that DijitalPi has not conducted.
Se Yan; Han Zhong; Zemin (Zachary) Zhong; Wenyu Zhou; Nitin Mehta Peking University, Guanghua School of Management (Se Yan); University of Toronto, Rotman School of Management (Han Zhong, Zemin Zhong, Nitin Mehta); City University of Hong Kong (Zemin Zhong); Zhejiang University, International Business School (Wenyu Zhou).
The academic work belongs to the researchers named above. This page contains DijitalPi's explanatory review and original business examples; it is not a full translation of the paper.