Blog

Engineering · September 9, 2026 · 5 min read

The tags were right. Our cleanup moved them.

The report was one line: the tags came out wrong. They had. In 182 of 190 segments, every inline tag was piled up at the end of the target. Bold that should have wrapped two words wrapped nothing, and a link sat after the full stop.

The model had them right

We keep every model response for a job, and in those responses the tags were exactly where they belonged. So the damage happened somewhere between the answer and the file, which is to say in our code.

Two safety nets pointed at each other

When a segment goes to the model, its inline tags are replaced by numbered placeholders: {1}, {2}. On the way back, the placeholders are swapped for the real tags, and then a cleanup step removes any leftover placeholder the model made up. Invented tags are a real problem, and that cleanup exists for good reasons.

In this particular Phrase export, the original tags are themselves written as {1}, {2}. Identical to our placeholders. The swap changed nothing, and the cleanup then deleted every real tag as if the model had invented it.

The validator saw a target with no tags against a source with several, and retried. Same result. Then the last resort ran: a Phrase file with missing tags is rejected on import, so a repair step put the missing tags back at the end of the segment, where they would at least be accepted. The file imported. It was wrong.

Each of those steps is reasonable on its own. Together they turned a correct translation into a broken one and reported it as repaired.

We had already fixed this, somewhere else

The uncomfortable part: the same collision had been found and fixed in our memoQ pipeline in June, where the restore now works in three phases and cannot confuse an original tag with a placeholder. The fix was never carried over to Phrase. It is now. Trados and XLIFF files were never affected, because their tags are XML and never look like our placeholders.

The damage was not only tags

The false retry regenerated the whole chunk, not just the tagged segments. Segments with tags shipped from the second attempt and segments without tags from the first, so the same product term came out several different ways in one document. A tag bug had quietly become a terminology problem.

We repaired the delivered file from the saved responses, without retranslating anything, and unified the terms. If you received a Phrase Word file from us before September 8 and its tags sit at the end of segments, tell your Account Manager: it can be repaired the same way.