A review of school notices translated into Chinese (Simplified) came back with a short note: the documents were mostly in Traditional characters. Not all of them, and not all of each one. The glossary terms were in Simplified; the text around them was not.
That mix was the first clue. Glossary enforcement runs after translation and writes the approved term, which was Simplified. So the translation itself had been asked for the wrong script, and enforcement had patched only the words it knew about.
Where the language went missing
Every job carries a set of language rules for its target: script, punctuation, number and date formats, quotation marks. To pick the rules, the pipeline normalizes the language name it receives and matches it against a list of known names.
Normalization removed anything in parentheses. “Chinese (Simplified)” became “chinese”. So did “Chinese (Traditional)”. Two different rule sets now matched the same word, and the pipeline took the first one in the list.
Which one came first depended on the order the server’s file system returned the files in. On one of our servers, Simplified came first and every job was right. On another, Traditional came first. The same job, with the same settings, could be correct or wrong depending on which machine picked it up. That is why it looked intermittent, and why our own tests had passed.
How we proved it
We rebuilt the instructions a failing job had sent and compared them with the instructions from a clean run of the same file. The difference was 1,029 characters, and it was exactly the difference between the Traditional and the Simplified rule blocks. That turned a theory into a cause.
Not only Chinese
The same collapse could happen to any language whose variants are distinguished in parentheses: Spanish (Spain) and Spanish (Latin America), Portuguese (Portugal) and Portuguese (Brazil), French (Canada) and French (France), English (UK) and English (US). Whether it did depended on the same file-order lottery. An older part of our system had been made region-safe in July; this newer pipeline had its own copy of the logic and never received that fix.
The fix
- Rules are chosen by language code first, which every job already carries, and only fall back to names when there is no code.
- The text in parentheses is kept and used for matching, so “Simplified” and “Traditional” stay different words.
- The order is deterministic and the same on every server, so a job cannot depend on where it runs.
We verified it on each server by running Simplified Chinese end to end and counting Traditional characters in the output: zero.
The lesson we keep relearning: a language is a code, not a name. Names are for people.
Want this in your workflow? Try Fily with one file — no card, no demo form.
