INTERNATIONAL RESEARCH / ACADEMIC PUBLICATION
Can AI give your customers incorrect information?
An assistant may quote the correct price but misrepresent the return policy. We explain with examples which questions you should try before introducing the assistant to your customers.
Türkçe soru cevaplama için büyük dil modelleri üzerinde geniş ölçekli etki analizi
ResearchersZekeriya Anıl Güven · İzmir Bakırçay University
Journal of Gazi University Faculty of Engineering and Architecture · 2025 · Source language: Turkish
Open the original publicationWe explain the paper in plain English, discuss what it may mean for organisations and identify our own examples separately.
English review: 9 September 2026 · The source date appears in the citation above.
The source has been checked; the review still carries its original editorial status. DijitalPi did not conduct a customer experiment for this paper.
LET'S READ THE RESEARCH TOGETHER
We first explain the researchers' question, method and findings. We then discuss how to interpret the results, clearly separating DijitalPi's commentary from the source.
01 / WHAT DID THE RESEARCHERS WANT TO UNDERSTAND?
How do we know if a Turkish answer is correct?
Zekeriya Anıl Güven examined the success of different models in answering questions in Turkish and the effect of how the answers are scored on the results.
Explaining the context · DijitalPi commentary
When testing an assistant, it can be tempting to look only at whether it speaks natural Turkish. The user also needs the information in the answer to be correct. That requires defining what counts as a correct answer. A single character can matter in a product code, while different wording can express the same meaning.
Our own example: “Where is the meeting?” Let the expected answer to the question be "Ankara". It is also reasonable to accept the answer "in Ankara". However, we cannot accept the statement "not in Ankara". In order to understand the scoring issue of the research in daily life, let's keep in mind that similarity in spelling and meaning are not the same thing.
02 / HOW WAS THE RESEARCH CONDUCTED?
How did they test the question?
The researcher trained and evaluated models in the BERT, ALBERT, DistilBERT, mDeBERTa and mT5 families with SQuAD data machine translated into Turkish. There were 8,291 questions in the test: 2,346 were answered in the given text, 5,945 were not.
Understanding the method · DijitalPi commentary
The task here is to answer the question based on the given text. Just because the answer is not found in the text does not mean that it is not known at all in the world. For example, if you are given the address of only one store, you cannot remove the working hours from this document. Noticing a deficiency in such a question is also a skill that should be evaluated.
Fine-tuning means training a pretrained model further with examples for a specific task. Scoring defines which answers count as correct. These stages must be separated: a higher score after the evaluation rule changes does not mean that the assistant learned new information. Semantic similarity compares meaning rather than requiring the exact same wording.
03 / RESEARCH FINDINGS
Overall success can hide the difference between question groups.
The overall accuracy of mDeBERTa-TrSQuAD is 74.50%. When re-scored by semantic similarity, it becomes 79.09%. In the second evaluation, the accuracy of the questions with answers was 62.83%, and the accuracy of the questions with no answers was 85.50%. The similarity threshold is used as 0.5.
- Questions with answers
- 62.83%Re-evaluation with semantic similarity in a set of 2,346 questions.
- Questions with no answers
- 85.50%The same evaluation for 5,945 questions; the ability to say "no answer" correctly.
Source: Data and method, PDF page 4 ↗
04 / DIJITALPI'S EXPLANATION
How should we interpret these findings?
There are not the same number of questions in these two groups. So the overall score is not the simple average of two percentages. A larger group has a greater impact on the overall result. That's why we want to see how the questions are distributed when interpreting a success rate.
Consider a separate illustrative calculation: imagine that 80 of 100 questions have no answer in the supplied text. A system that says "I don't know" every time could score 80 correct answers under this rule, while helping with none of the 20 answerable questions. This is not the researcher's experiment; it shows why an overall score alone may be insufficient.
We have two expectations from a business perspective: to use existing information correctly and not to make up missing information. Counting them separately makes it easier to understand which problem is due to lack of documentation and which is due to the assistant's answer. If we only look at the total percentage, we may miss room to improve.
05 / CONCLUSION AND OPEN QUESTIONS
What did we learn, and what do we still not know?
In the comparison of the study, mDeBERTa stands out; semantic evaluation affects the scores. However, translation data and specific task scope preclude using the result as the overall success rate of current business assistants.
How does the outcome change with your own products, exceptions and updated documentation? This requires a business-specific test. The 30-question preliminary test below makes that need concrete; its responses have not yet received human evaluation.
Return to the researchers' original publication
Data and method, PDF page 4 Accuracy results, tables 4 and 6Source and ownership
Journal of Gazi University Faculty of Engineering and Architecture · 2025. Open the original publication
The academic work belongs to the researchers named above. This page contains DijitalPi's explanatory review and original business examples; it is not a full translation of the paper.
← Return to all research articles