9.4 MB
42 files
Updated 4 days ago
Name
Size
results
results_ext
results_robust
README.md17.5 kB
xet
analyze.py4.12 kB
xet
analyze_ext.py2.46 kB
xet
analyze_robust.py1.7 kB
xet
blindspots_harness.py10.1 kB
xet
figure.png89.4 kB
xet
figure.py3.02 kB
xet
figure_ext.png121 kB
xet
figure_ext.py3.71 kB
xet
items.jsonl606 kB
xet
items_ext.jsonl471 kB
xet
items_robust.jsonl216 kB
xet
make_items.py9.02 kB
xet
make_items_ext.py10.3 kB
xet
make_items_robust.py2.97 kB
xet
run_all.sh709 Bytes
xet
run_ext.sh782 Bytes
xet
run_ext_rest.sh527 Bytes
xet
run_rest.sh619 Bytes
xet
run_robust.sh800 Bytes
xet
README.md

Small models read French invoices the American way

When the instruction is in English and the document is French, small frontier models read 1,250 kg as one thousand two hundred and fifty kilograms and 2.500 kg as two and a half kilograms. The same numbers with a French prompt are read correctly much more often (Gemma 4 E2B goes from 97% with a French prompt to 8% with an English one). The models know the French convention, but they let the language of the instruction decide which convention to apply, instead of the origin of the document. The errors only happen when the wrong reading also gives a well-formed number, so the output looks fine and nobody notices. The same thing happens beyond numbers, with the French word billion (10^12), with family names written first on French forms, and with dates, and in some cases the models make the opposite mistake on US documents.

Accuracy per condition

Results

I ran 1,600 paired prompts per model. The table shows accuracy on the ambiguous numbers (1,250 and 2.500, compare and sum tasks, 160 items per cell, 95% intervals of about 5 points).

model English, US numbers (control) French prompt English prompt, French document French prompt, rule stated
Gemma 4 E2B 97% 97% 8% (88% read the US way) 97%
Qwen3 4B Instruct 100% 80% 44% (56% US) 90%
Ministral 3 3B 99% 71% 33% (66% US) 58%

On the same numbers, going from a French prompt to an English prompt costs 89 ± 5 points for Gemma, 36 ± 7 for Qwen and 38 ± 8 for Ministral.

For Gemma and Qwen the knowledge is clearly there. Gemma reads 1,250 correctly 97% of the time with a French prompt and again 97% when the rule is spelled out, but only 8% of the time when the same invoice lines come with an English instruction. The failure is also invisible in an average score: on numbers that have no valid US reading (3,75, 1 234,5) all three models stay between 94 and 100% in every condition. Only the ambiguous numbers fail, and when they fail the answer is almost always exactly the US reading (56 to 88% of these items), which rules out random mistakes. Gemma even explains its error. Reading 3.400 kg from a French invoice, it writes "in the context of a French invoice, the period (.) is the standard decimal separator [...] A = 3.4 kg". It knows the document is French and invents the wrong rule.

The model from a French lab is the most fragile of the three. Ministral 3 3B misreads French numbers even with a French prompt (71%), and stating the rule does not help much (58%). It also has the opposite bias on dates: on a ticket from a US airline it reads 05/11/2025 as 5 November in 18 of 80 cases. So the real problem is that the models do not follow the convention of the document, in either direction. Dates also depend on the model. Gemma reads French DD/MM dates correctly in every condition, while Qwen reads 34% of them the US way when the instruction is in English.

The two number tasks measure two different things. Comparing two weights needs the actual magnitude, so it tests reading. With an English instruction Gemma gets 16%, Qwen 38% and Ministral 45%, where guessing would give 50% and the US reading gives 0% by design. Adding two weights tests what the model commits to when it has to write a value. On paper 1,619 + 7,706 = 9,325 is correct in both conventions, and the mistake only shows up when the model writes 9325 instead of 9.325 in the answer field, which is exactly what a document pipeline does. Gemma writes the integer reading explicitly (1619 + 7706) in 25 of its 40 failures on these sums, while Qwen and Ministral keep the ambiguous string and convert it the US way at the end.

Robustness

I reran the same 160 ambiguous items with three more English instructions: two rewordings (a supplier in Lyon, and a packing list table from a Paris warehouse) and the original prompt with the French convention stated in English.

model French prompt English, original rewording 1 rewording 2 English + rule stated
Gemma 4 E2B 97% 8% 6% 9% 76%
Qwen3 4B Instruct 80% 44% 32% 32% 51%
Ministral 3 3B 71% 33% 46% 23% 40%

Both rewordings reproduce the drop, and for Qwen they make it worse, so this is not an artefact of one prompt. Naming a French city is not enough context. Stating the rule in English only fixes part of the problem. Gemma recovers on comparisons (98%) but not on sums (55%), Qwen reaches 51% and Ministral 40%. The same rule written in French worked much better for Gemma and Qwen (97% and 90%). Adding one sentence to the system prompt, which is the first fix most teams would try, does not close the gap.

Beyond numbers

To check whether this is about number formatting only, I added five more conventions that differ between France and the US, with the same four paired conditions and the same scoring (1,200 more prompts per model, 60 items per task). The French word billion means 10^12 (a milliard is 10^9), so "9 billions d'euros" is 9,000,000 million euros and not 9,000. French forms write the family name first and in capitals, so in LAURENT Simon the first name is Simon, and I chose names that are common as both first names and family names. A French 3e étage is 3 levels above the street, while a US "3rd floor" is 2 levels above it. The last two tasks use DD/MM dates: which of two deadlines comes first, and how many days separate two dates.

Accuracy per condition, beyond numbers

task model English, US document (control) French prompt English prompt, French document French prompt, rule stated
French billion (10^12) Gemma 4 E2B 97% 57% 7% (67% US) 100%
Qwen3 4B Instruct 93% 17% 0% (100% US) 100%
Ministral 3 3B 67% 83% 43% (50% US) 100%
French milliard (10^9), control all three 100% 97 to 100% 97 to 100% 77 to 100%
first name in NOM Prénom Gemma 4 E2B 85% 43% 13% (87% give the family name) 97%
Qwen3 4B Instruct 98% 57% 48% (50%) 98%
Ministral 3 3B 98% 80% 55% (45%) 100%
which deadline comes first Gemma 4 E2B 97% 100% 95% 100%
Qwen3 4B Instruct 97% 100% 78% (22% US) 100%
Ministral 3 3B 30% 55% 93% 52%

The French billion is the clearest case. The control word milliard is read correctly almost everywhere, so the models understand the question, but the false friend billion fails as soon as the instruction is in English, and for Qwen even with a French prompt (17%). Stating the rule fixes it completely for all three models. The name order shows the same pattern for Gemma, which gives the family name as the first name in 87% of the cases under an English instruction, while stating the rule brings all three models to 97% or more. Deadlines confirm the date results of the main run: Gemma reads French dates correctly, Qwen reads 22% of them the US way under an English instruction.

The floor task and Ministral's results show the other side of the same problem. On a US listing, Gemma and Qwen answer that a "7th floor" is 7 levels above the street in 60 of 60 cases, so they never apply the US convention, and their 100% on French floors only means they always repeat the floor number. Ministral does the reverse on French floors (30% with a French prompt, 70% of answers follow the US rule) and on US dates (30% on the US control, consistent with the 18 of 80 US dates it read as DD/MM in the main run). Taken together, the models do not track the convention of the document. Which convention they apply depends on the model and on the language of the prompt, and it can be wrong in either direction.

The task that counts the days between two dates did not work as a test. Even the US control is at 8 to 23%, many answers were cut off at 512 tokens while counting month by month, and almost no error matched the US reading, so it measures arithmetic rather than date reading. It stays in the files but I do not draw conclusions from it.

Limitations

The items are synthetic: random numbers in short invoice and ticket templates, not real scanned documents. The extension uses 60 items per task, so its intervals are wider (5 to 25 points on the paired drops). Other conventions (German 1.234,5, Swiss 1'234.5, Indian 1,23,456, units, week numbering) are untested. Qwen3 4B and Ministral 3 3B ran in 8-bit (LLM.int8) to fit on an 8 GB GPU, which is close to lossless on most benchmarks but is not the same as bf16. I only tested small models, as the form asks for models between 0.6 and 6B parameters, so I do not know whether the larger models of the same families share the bias. Decoding is greedy with one answer per prompt, so the error bars cover the choice of items and not decoding randomness. The main run uses one prompt template per condition, and the rewordings only cover the English instruction case.

1. The blind spot and why I care about it

I build document AI in France. One tool I work on reads shipment files (air waybills, packing lists, commercial invoices) and fills in French customs clearance instructions. In pipelines like this one the instruction to the model is easily written in English, because prompt templates, tools and teams often default to it, while the document itself is French: Poids net : 1 234,50 kg, Valeur : 2.500,00 €, Date : 03/04/2026. French uses a comma as the decimal separator and a space or a dot to group thousands, and so do most of continental Europe, most of Latin America, Russia, Indonesia, Vietnam and much of Africa. That is hundreds of millions of people whose documents follow this convention. A model that quietly applies the US convention to these documents does not produce an obvious error. It produces a clean number that is wrong by a factor of a thousand, so a 1.25 kg parcel becomes 1,250 kg and a declared value of 2 500 € becomes 2.50 €. Standard benchmarks miss this for three reasons. Multilingual benchmarks translate the question and adapt its numbers consistently, or keep all numbers in one convention, so the mixed case of an English instruction over a local document is rarely tested (I could not find a benchmark that isolates it), even though it is the everyday case for document AI in Europe. Math and numeracy benchmarks are mostly written with US formatting. And an average score hides the problem, because models do well on numbers that cannot be read the US way, so the failure only appears when you look at the numbers whose wrong reading is also valid. This is close to something I spent the summer measuring in physics, where a generated gravitational lensing image can look exactly like the requested lens without being a real lensing of anything. Looking right and being right are different things, and a metric that only checks the first one will miss the second.

2. Evaluation

I wrote 400 base items and rendered each one under four conditions that use exactly the same values, so every comparison is paired (1,600 prompts per model).

condition prompt language number format what it tests
us_en English US control: can the model do the task at all?
fr_fr French French a French user
fr_in_en English French, and the document is said to be French the customs invoice case
fr_explicit French French, with the convention stated is the knowledge there when asked?

There are three tasks: compare two weights (answer A or B), add two invoice weights (numeric answer), and read the month of a ticket date. There are four number types.

type example value US reading
amb3 1,250 1.25 1250 ambiguous, the wrong reading is a valid number
dotk 2.500 2500 2.5 ambiguous
spacek 1 234,5 1234.5 none unambiguous
dec12 3,75 3.75 none unambiguous

Dates use DD/MM/YYYY with both day and month at most 12, so the US reading is also valid. In the comparison items the second number is chosen so that the correct reading and the US reading give opposite answers, and in the sums the two readings differ by a factor of 1000. Positions A and B are balanced (80 each).

Scoring uses no LLM judge. Each reply must end with ANSWER: ... and the script takes the last one. An answer is strictly correct if it is right and in the requested format (a plain number with a dot as decimal separator). It is leniently correct if it is right under any reading of what the model wrote, so a failure cannot come from how the model formatted its number. It counts as a US misreading if it is exactly the value the US convention gives, which is the test of the mechanism, since a random error would not land on that value. Decoding is greedy with the chat template and up to 512 new tokens. Error bars are binomial for a single condition and paired over the same items for the difference between two conditions.

I chose three small members of frontier model families and ran them locally on a laptop RTX 5070 with 8 GB of memory: google/gemma-4-E2B-it in bf16 (its per-layer embedding table stays on the CPU, as the architecture intends), Qwen/Qwen3-4B-Instruct-2507 in 8-bit to fit in memory, and mistralai/Ministral-3-3B-Instruct-2512 (BF16 release, run in 8-bit), which comes from a French lab and lets me check whether a French model does better on French formatting.

3. A path forward

For Gemma and Qwen the knowledge is already there, as the fr_explicit condition shows, and the problem is a wrong default when the language of the instruction and the convention of the document differ. This situation is cheap to generate at scale for training data. You can take real documents from comma-decimal countries, keep their numbers as they are, and pair them with instructions in another language. The answers can be computed, so the labels are free. The ambiguous cases (1,250, 2.500, 03/04) should be over-represented because they are the only ones where the default matters. Ministral, which gets it wrong even when told the rule, probably also needs more of the convention itself in its training data. For fine-tuning I would try two things. The first is a consistency objective: the same value written in two conventions must give the same answer, which targets exactly the variable that matters and penalises the plausible but wrong answer directly. The second is to reward reasoning that states which convention the document uses before computing anything. Gemma's failures rarely mention the convention (18 of 160 ambiguous items), and when they do they sometimes invent it, which a verifier could catch. The robustness results show that prompting alone is not enough, since stating the convention in the prompt only recovers part of the gap (Gemma from 8% to 76%, Ministral from 33% to 40%). On the architecture and tooling side, numbers are tokenised as text, so 1,250 in a French document and 1,250 in an American one are the same tokens and only the context can tell them apart, and the evaluation shows that the context is not weighted enough. Two fixes follow from this. Numbers could be normalised when the document is ingested, with an explicit tag for the convention (⟨num locale=fr-FR⟩1,250⟨/num⟩) inferred from the document rather than from the prompt. Or number extraction could go through a convention-aware parser (ICU or Babel) exposed to the model as a tool, which document pipelines can do today. Finally, multilingual and document AI benchmarks should include the mixed case and report the ambiguous subset and the US misreading rate separately, because an average over all numbers hides the failure.

Files

file what
items.jsonl the 1,600 prompts with gold answers and the US misreading value
make_items.py generates items.jsonl (seeded, deterministic)
blindspots_harness.py loads a model, runs the prompts and scores every reply
analyze.py builds the tables above (results/summary.json)
results/<model>/outputs.jsonl every raw reply with its score
results/<model>/report.json accuracy per condition and paired differences
make_items_robust.py, items_robust.jsonl the 480 robustness prompts (same numbers)
analyze_robust.py, results_robust/ robustness outputs and summary
figure.py, figure.png, figure_ext.py, figure_ext.png the two figures above
results/tables.md full tables per cell, including dates and unambiguous numbers
make_items_ext.py, items_ext.jsonl the 1,200 prompts of the extension (billion, names, floors, dates)
analyze_ext.py, results_ext/ extension outputs and summary
run_all.sh, run_rest.sh, run_robust.sh, run_ext.sh, run_ext_rest.sh the exact commands used (run_ext_rest.sh resumed Ministral after the machine shut down, and the two parts are merged in results_ext/)

To reproduce, run python make_items.py and python make_items_robust.py, then the three run_*.sh scripts (an 8 GB GPU is enough), then python analyze.py, python analyze_robust.py and python figure.py. For the extension, run python make_items_ext.py, run_ext.sh, python analyze_ext.py and python figure_ext.py.

Total size
9.4 MB
Files
42
Last updated
Oct 2
Pre-warmed CDN
US EU US EU

Contributors