We tested six locally run language models with 250 questions about 50 Finnish invoices, calls for tender, contracts and meeting minutes. Three models answered 85-95 percent of the questions correctly on an ordinary computer, with no hardware bought for AI. Extracting a single value is essentially a solved problem. Applying a condition is not: even the best model reached only 83 percent there.
A local LLM is an attractive option for a company whose documents must not leave the building. Public benchmarks, however, almost always measure English and general knowledge, not whether a model can correctly read the payment terms on a Finnish invoice or the deadline in a call for tender. So we measured it ourselves.
The study at a glance
- Material:
- 50 Finnish documents: 16 invoices, 14 calls for tender, 10 contracts and 10 sets of meeting minutes.
- Questions:
- 250, five types: extraction, calculation, applying a condition, cross-reference and a trap where the information is not in the document.
- Models:
- Gemma 4 12B, Qwen 3.8 27B, Phi-4, Qwen3 8B, Mistral 7B and Llama 3.2 3B.
- Setup:
- Run locally in LM Studio, CPU only, temperature 0.
- Best result:
- Gemma 4 12B, 94.6% correct on a shared sample of 202 questions.
- Biggest gap:
- Extraction 94-100%, applying a condition at most 83%.
What was tested and how?
Each document got five questions, one of each type, and every question has an answer key with references to the source lines. The calls for tender, contracts and minutes are real public documents. The invoices are synthetic, because real sales invoices are not publicly available, and they were generated from templates by a script, not by a language model. No customer data was used.
Every model received the same instruction: answer only from the document, do not guess, and say so if the information cannot be found. Scoring had two stages. An automatic scorer settled the clear-cut cases, and the rest of the answers were read by hand.
The trap questions ask for information that is missing from the document, with a decoy nearby. For example, when the question asks for the buyer's business ID, the document contains the seller's. The correct answer is then "not found".
Which local LLM understands Finnish best?
Gemma 4 12B was clearly the best, and three models formed the top tier. The comparison uses the 202 questions that every model managed to process, so all models were measured on exactly the same questions.
| Model | Parameters | Correct | Share |
|---|---|---|---|
| Gemma 4 12B | 12 bn | 191 / 202 | 94.6% |
| Qwen 3.8 27B | 27 bn | 180 / 202 | 89.1% |
| Phi-4 | 14 bn | 171 / 202 | 84.7% |
| Qwen3 8B | 8 bn | 125 / 202 | 61.9% |
| Mistral 7B | 7 bn | 103 / 202 | 51.0% |
| Llama 3.2 3B | 3 bn | 70 / 202 | 34.7% |
Two findings stand out. First, size is not decisive. Gemma is less than half the size of Qwen 27B, yet it solved 14 questions that Qwen got wrong, while Qwen solved only three that Gemma got wrong. The difference is statistically significant (p ≈ 0.01). Second, there is a gap of more than 20 percentage points between the top tier and the middle tier. In this test, models under ten billion parameters were not good enough for business documents.

Is it enough that the AI finds the information in the document?
No. Five models out of six found a single value, such as a business ID or a deadline, 94-100 percent of the time. When the answer had to be calculated or inferred, results fell apart: 66-98 percent for the top models, 19-48 percent for the middle tier. Answering from a document is not the same as understanding it.
The hardest task was applying a condition. A typical question: "The buyer pays the invoice on 10 June 2026. Is the payment late, and what interest rate applies to the delay?" Answering it means finding the due date, counting days and choosing the right term. Gemma got 83 percent right, Qwen 27B 71 percent and Phi-4 66 percent.
This is exactly what you would expect from an accounts receivable or procurement assistant, and the mistakes were instructive:
- Phi-4 added the payment term to the due date in five invoice questions and concluded that a late payment was not late. The error repeated in the same form, so it is systematic.
- A step was skipped. 3 hours times 52 euros was correctly calculated as 156 euros, but VAT was left out even though the question asked for the price including VAT.
- The answer corrected itself midway. Qwen 27B and Phi-4 often opened with the wrong yes or no and reached the right conclusion at the end. A reader who only reads the first sentence gets the wrong answer.
The practical conclusion is clear: extraction can be automated almost directly, but where conditions and calculations are involved, the process needs a review step. If you are deciding which document work to automate first, the same logic applies as in scoring processes to choose your first automation target: start with work that is repetitive and rule-based.
Does the model make up an answer when the information is missing?
The top models did not. Gemma and Qwen 27B correctly answered every trap question with "not found". Mistral 7B invented an answer for 20 traps out of 49, or 41 percent, and the other models did so 7-8 percent of the time.
Almost every invented answer was a value that sat close by in the document but belonged to a different question. Asked for the buyer's business ID, Mistral gave the seller's. Asked for the warranty period, it answered "78 months", which was the estimated duration of the procurement. Hallucination here is not random invention but latching onto a wrong yet plausible number. That is why it is hard to spot.
The top models had the opposite problem. Two thirds of Qwen 27B's wrong answers were refusals to questions whose answer was in the document, such as calculating late payment interest. The same instruction that made the model flawless on traps made it too cautious at reasoning. This is worth knowing when you design instructions: caution and accuracy are a trade-off, and you do not get both for free.
What do the results mean for document automation?
A local LLM is good enough for extracting information from Finnish documents, and for much of the calculation, on an ordinary computer, provided the model is chosen well. The biggest advantage of running locally is that documents stay in your own environment. That is one of the questions we covered in our AI security checklist for SMEs.
The limits show up in three places:
- Long documents. On CPU alone, some models could not process the longest documents, around 72,000 characters, in a reasonable time. Phi-4's context holds about 40,000 characters. The results therefore describe short and medium-length documents.
- Applying conditions. 83 percent is a good result, but not good enough without review when payments or contract terms are at stake.
- Model choice. The gap between the middle tier and the top was more than 20 percentage points. The wrong model produces errors that look correct.
Cloud models were not part of this test, so we make no claims about them. The choice between a local and a cloud model depends on the data, document length and volume. The monthly cost of a cloud model can be calculated in advance, as we showed in what running AI costs per month.
Four things worth doing before deployment:
- Test the model on your own documents, not on public benchmarks.
- Measure extraction, calculation and applying conditions separately. An average hides the difference.
- Add trap questions where the correct answer is "not found".
- Build a review step where the model applies conditions, and read answers to the end.
If you want to know which of your document processes are ready for automation and where a human review is needed, a fixed-price Automation Assessment goes through them and ranks the opportunities by their value in euros.
Which document processes should you automate first?
The Automation Assessment goes through your processes and shows which steps AI handles reliably and where a human review is needed.
Book an Automation AssessmentEmpirica Finland is a Finnish provider of operational AI and automation. It builds automations from a company's own data, both system data and data produced by devices and sensors, and is responsible for their continuous operation. Empirica is a Claude Partner Network member and a Microsoft partner.



