We tested six locally run language models with trap questions whose answer was not in the document. The best models never invented an answer. The weakest did so 41 percent of the time. Almost every invented answer was a real value from the same document that belonged to a different question. The best models made the opposite mistake: they answered "information not found" even when the answer was there.
This article goes deeper into our test in which six language models read Finnish invoices, contracts and meeting minutes. That test measured how often the models answer correctly. Here we look at what kind of errors they made. In automation, the kind of error matters as much as the number of errors: an invented due date slips through unnoticed, while after a needless "information not found" a person has to look up the answer themselves.
Trap questions at a glance
- What was measured:
- Whether the model recognises that the requested information is not in the document.
- Traps:
- One per document across 50 documents, 44-50 completed depending on the model.
- Decoy:
- In every trap, the document contained a plausible but wrong value nearby.
- Best result:
- Gemma 4 12B and Qwen 3.8 27B, no invented answers.
- Weakest result:
- Mistral 7B, invented an answer to 41% of traps.
- Most common error among the leaders:
- Needless refusal, 20 of Qwen 27B's 30 errors.
Does AI invent information that is not in the document?
Some models do, some do not. Gemma 4 12B and Qwen 3.8 27B correctly answered every trap question with "information not found". Mistral 7B invented an answer to 20 of 49 traps, and the other models did so 7-8 percent of the time.
In a trap question, the requested information is missing from the document, and its absence was verified by searching the source text. The correct answer is "information not found". Every model received the same instruction: answer only from the document, do not guess, and say so if the information is not there.
The hallucination rate alone is misleading, though. A model that answers "information not found" to almost everything passes the traps but is otherwise useless. That is why the table shows a second number: how often the model answered correctly when the answer was in the document.
| Model | Invented an answer to a trap | Correct when the answer was in the document |
|---|---|---|
| Gemma 4 12B | 0% (0/49) | 93.1% |
| Qwen 3.8 27B | 0% (0/47) | 86.3% |
| Phi-4 | 6.8% (3/44) | 82.5% |
| Llama 3.2 3B | 8.0% (4/50) | 18.1% |
| Qwen3 8B | 8.2% (4/49) | 54.4% |
| Mistral 7B | 40.8% (20/49) | 46.9% |
The right-hand column is calculated on the shared sample of 202 questions, excluding trap questions.
Llama 3.2 3B invented an answer to only 8 percent of traps, close to Phi-4. Yet when the answer was in the document, it answered correctly only 18 percent of the time, the weakest result in the test. Its low hallucination rate therefore says nothing about reliability: the model often answered "information not found" even when the answer was there. A good trap score is valuable only together with good answer accuracy, and in this test only Gemma and Qwen 27B achieved both.
Why is AI hallucination hard to spot?
Because an invented answer does not look invented. Almost every hallucination in the test was a number or date that really was in the document but belonged to another question. The model did not make things up from nothing. It grabbed the nearest plausible value.
Examples from the test:
- The buyer's business ID. Mistral answered confidently, giving a company name and a business ID. Both belonged to the seller on the invoice.
- Warranty period. Mistral answered "78 months", which was the estimated duration of the procurement. For another contract's warranty period, Llama answered "3 years", which was the contract term.
- Notice period for complaints. Qwen3 8B answered "14 days net". That is the payment term.
- Date the minutes were approved. Phi-4 and Qwen3 8B gave the publication date. Phi-4 even quoted the sentence stating that the minutes were published on that day.
Each of these would pass a spot check. The business ID has the right format and its check digit is valid, because it was taken from the same invoice, just from the wrong party. The warranty period is a plausible length. The answer is wrong only in relation to the question, and you notice it only by reading the document yourself.
This leads to a practical rule: when automation extracts data from a document, the check cannot rely on whether the value looks reasonable. The more useful question is where in the document the value came from. When the model is asked to state the source passage alongside its answer, a grab at the wrong field is easier to catch.

Can AI be too cautious?
Yes, and in our test that was the most common error among the best models. Of Qwen 3.8 27B's 30 wrong answers, 20 were refusals to questions whose answer was in the document. The same instruction that kept the model flawless on traps made it too cautious when the answer had to be calculated or inferred.
A typical example came from meeting minutes. According to the minutes, one participant left at 16:34 and the meeting ended at 19:08. Asked whether that person was present at the end of the meeting, Qwen 27B answered: "Information not found in the document." Comparing two times would have been enough. The model likewise refused to calculate late-payment interest and to reconcile VAT lines on an invoice, even though everything needed was on the invoice. Gemma refused less often, and its refusals mostly concerned interpretation questions, such as what waiving a right of call-in means.
Excess caution is a safer error than hallucination, because it is visible. It is not free, though. Every needless refusal hands the task back to a person, and if there are many of them, the automation saves less than it was supposed to.
A related pattern was the model correcting itself mid-answer. Qwen 27B and Phi-4 often opened with the wrong yes or no and reached the right conclusion only at the end: "Does not match. [...] The due date is 7 days later. This matches the payment term." If the automation reads only the first word, it gets the wrong result. In the test, these answers were scored as wrong.
How do you stop AI from guessing without needless refusals?
By testing on your own documents and measuring both error types separately. The instruction given to the model affects both errors. When the model is strictly told not to guess, it invents answers less often. At the same time, it more often declines to answer even when the answer is in the document. Choose how strict the instruction is based on which error costs more in your process.
In our test, every model received the same instruction, and no alternative instruction was tried. So we do not know how much a different instruction would have reduced Qwen 27B's refusals. We do know that the same instruction produced very different balances for different models, so the instruction and the model are worth testing together.
Two examples show how to make the choice. If the model extracts a bank account number from a purchase invoice for payment, an invented value costs money, and caution pays off. If the model compiles decisions from meeting minutes for internal use, a needless refusal hurts more than an occasional inaccuracy. The risk grows when AI is allowed to act without a human check, which is the key difference between an AI agent and traditional automation.
Four things to include in your test before going live:
- Trap questions with decoys. Ask for information that is not in the document but has a similar-looking value nearby. Traps without a decoy are too easy.
- Two numbers instead of one. Measure the share of invented answers and, separately, the share of correct answers to questions that do have an answer. A single average hides exactly what you want to see.
- A review of refusals. Check how many "information not found" answers were actually wrong. If there are many, try adjusting the instruction before switching models.
- The conclusion in its own field. Ask the model for its final answer as a separate field, so the automation does not take the first word of the response as the conclusion.
A locally run model is a natural starting point for this kind of test, because the documents stay in your own environment during the trial too. If your documents are confidential, our AI security checklist for businesses helps you work out where the documents may be processed and which services they may be sent to.
The results apply to six models run on CPU only, to short and medium-length documents, and to a single run per question. The test did not measure cloud-hosted models.
If you want to know where AI can work on its own in your document processes and where a human check is needed, a fixed-price Automation Assessment reviews your processes and calculates the value in euros for each opportunity.
Where can you trust AI in your document work?
The Automation Assessment reviews your document processes and shows which steps AI can handle reliably and where a human check is needed.
Book an Automation AssessmentEmpirica Finland is a Finnish provider of operational AI and automation that builds automations from the data produced by a company's systems as well as its devices and sensors, and is responsible for keeping those automations running. Empirica is a Claude Partner Network member and a Microsoft partner.



