What it is
When someone says a wording is fixed, something has been counted. A conventional unit has no instrument behind it and no body willing to declare it correct, so the only thing on offer is text: the sequence was found, found often, or found in many places. Those are three claims and not one.
Each has its own measure. One asks whether the sequence occurs and how often. Another asks whether its parts turn up together more often than their separate frequencies predict. A third asks whether the occurrences are spread through the corpus or piled into a corner of it. The three come apart, because a sequence can be frequent and concentrated, or uncommon and tightly bound.
Two neighbours hold the ground on either side: the record on the cline states that a count needs a rule before it can count, and the record on unithood that a figure travels no further than the rule that produced it. This one begins where those end, with the rule declared and the number in hand, and asks what the number is evidence of. The terminology section weighs the evidence behind a designation; what is weighed here is the evidence behind a wording.
Where the fixity comes from
Nothing prescribes a conventional wording, so the whole warrant is attestation, and attestation comes in grades. The literature supplying the measures is more careful about their limits than the reviews that quote them.
- Raw frequency is the weakest grade. Gries observes that frequencies in isolation can mislead, because they take no account of how far the item is dispersed, and that the dispersion measures already proposed are neither widely known nor applied. He offers DP instead.
- Association asks whether co-occurrence beats chance. Church and Hanks adapt mutual information from Fano and state two limits themselves: the ratio becomes unstable when the counts are very small, so they leave out pairs occurring together five times or fewer, and an improvement would use t-scores but would need a variance estimate beyond the paper's scope. The score can also mislead, they add, because it takes only distributional evidence into account.
- Rychlý's logDice is 14 plus the base-two logarithm of twice the joint frequency over the sum of the two marginal frequencies. Its theoretical maximum is 14, values usually stay below 10, and negative values mean no statistical significance. It does not depend on the size of the corpus, and its values are corpus-specific, so pairs drawn from different corpora cannot be compared through it.
No cut-off appears here, because none of the three sources supplies one: a threshold is declared by whoever runs the study.
Consequences for translation
A measure is quoted in review far more often than it is run, and supports less than it is asked to.
- A frequency claim arrives with its corpus, its threshold and its direction, or it is an assertion wearing a number.
- A corpus-specific score cannot be carried between collections. Two studies reporting one value on different corpora have agreed about nothing.
- A count counts strings. Where a target version states the same thing as a noun in one paragraph and an adjective in the next, a search on either form under-reports, and the shortfall leaves no trace.
- Attestation in the document at hand is the weakest grade of all. It shows what this text prints and settles nothing about what the target community writes, so an objection resting on it reaches the client as a preference.
Examples
EN>NL
- Source: if it is grossly unfair to the creditor
- Target: indien zij een kennelijke onbillijkheid jegens de schuldeiser behelzen
- Comment: Directive 2011/7/EU, Article 7(1). Nothing in the act fixes either wording; the act attests both. Dutch carries the concept as a noun here and as an adjective at 7(2) and 7(3), so a search on kennelijk onbillijk finds two occurrences out of three.
EN>DE
- Source: if it is grossly unfair to the creditor
- Target: wenn sie für den Gläubiger grob nachteilig ist
- Comment: Same provision. German holds the adjective at all three paragraphs, so the search that under-reports in Dutch is accurate here. Whether a count can be trusted is a property of the version counted.
EN>FR
- Source: if it is grossly unfair to the creditor
- Target: lorsqu'elle constitue un abus manifeste à l'égard du créancier
- Comment: The same provision once more. French does what Dutch does, running a noun phrase here and manifestement abusive at Article 7(2). One article in three target versions yields three counts of one wording, none of them wrong.
References
- Church, K. W. and Hanks, P. (1990). "Word association norms, mutual information, and lexicography." Computational Linguistics 16(1), 22 to 29.
- Gries, S. Th. (2008). "Dispersions and adjusted frequencies in corpora." International Journal of Corpus Linguistics 13(4), 403 to 437.
- Rychlý, P. (2008). "A lexicographer-friendly association score." In P. Sojka and A. Horák (eds), Proceedings of Recent Advances in Slavonic Natural Language Processing, RASLAN 2008, 6 to 9. Brno: Masaryk University.
- Directive 2011/7/EU of the European Parliament and of the Council of 16 February 2011 on combating late payment in commercial transactions (recast), Article 7, English, Dutch, German and French language versions.