Join early access

Formal or informal “you”? The English string cannot say

English has one word for the person reading your interface and French, Spanish, Hindi and Korean do not, so every string that speaks to the user needs a decision the source never recorded, and machine translation follows it only where an engine offers a switch.

A funnel with a top hat and a flat cap sitting in its mouth and one small plain pebble below the spout, drawn in white line on a slate ground

English has one word for the reader, and your target languages do not

A settings screen says Your changes have been saved. Send it to a French translator and the first thing they need to know is whether the product says vos modifications or tes modifications. The English does not say. Nobody wrote it down, because in English there was nothing to write.

That is not a quirk of French. The World Atlas of Language Structures classifies how 207 languages treat politeness in the pronoun for “you”. In 136 of them the pronoun encodes no politeness distinction, and English is one of those. 49 have a binary one, the French tu and vous kind. 15 have several levels. In 7, including Japanese, speakers mostly avoid the pronoun for politeness reasons and use titles, status and kinship terms instead 1.

So most of the world’s languages do not make this choice. The ones software usually ships in mostly do. WALS files French, Spanish and Italian under the binary type, Hindi under several levels, and Japanese and Korean under pronoun avoidance, where respect lives in the verb endings and the nouns rather than in one word 1.

It is also not one rule applied the same way everywhere. The chapter points out that “a single grammatical distinction (familiar versus polite) corresponds to a complex set of pragmatic rules and social contexts”, and that those rules “are not identical to those of other languages with a binary politeness distinction” 1. Between two people, the rules settle the question on the spot. An interface talks to everyone at once, so somebody has to settle it in advance, once per language.

That answer is a fact about the product, not about the sentence. It is the same kind of fact as the number of plural forms a message needs, which we argued is a property of the target language that an English source cannot carry.

One publisher, one sentence, four forms of address

The clearest evidence that this is a product decision comes from one publisher answering it differently, language by language, for the same sentence. Microsoft’s localization style guides each work through the same English example, the setup line that begins “Give your PC a name”, and show the approved translation.

LanguageApproved translationForm of address
French (France)Donnez un nom à votre PC (celui que vous voulez).Formal, vous
Spanish (Spain)Dale a tu PC el nombre que quieras.Informal, tú
ItalianDai al tuo PC il nome che preferisci.Informal, tu
Portuguese (Portugal)Dê o nome que quiser ao seu PC.Pronoun left out
Microsoft's approved translation of the same English setup string, from each language's style guide.

French gets vous 2. Spanish and Italian get the informal form 3,4. The European Portuguese guide avoids the question: “you” “should be avoided in Portuguese, except when the information conveyed becomes unclear (and even in those cases, usage of the word ‘você’ should be avoided by rephrasing the sentence altogether, e.g. by using the passive voice)” 5. Same company, same product family, same English, four answers.

A second publisher can answer the same language the other way. Mozilla’s Spanish (Spain) guide says product translation uses the formal form of address, and its example turns “Choose an option:” into “Elija una opción:” 6. That is usted, where Microsoft chose tú. Neither is a mistake. They are two brands deciding how they want to sound in Madrid.

Where publishers land on the same answer, it is still a decision somebody wrote down. Mozilla’s Hindi guide says “it is encouraged to use honorific pronoun in Hindi. So, it is better to use words like आप, यह, वह instead of तुम, ये, वे respectively” 7. A translator who never read that page would have to guess.

One product does not always use one register

A single choice per language is already too coarse for these publishers. The same Mozilla guide continues: web content, even when it is about a product, uses the informal form. Its example is “When you use Firefox, you browse safer”, translated as “Cuando usas Firefox, navegas más seguro” 6. Same organization, same language, formal inside the browser and informal on the website selling it.

Microsoft’s Korean guide does the same by document type. It prefers the more conversational -세요 ending over -십시오 for the product’s voice, then adds that translators should use -십시오 “when localizing EULA or legal contents to express a formal tone” 8. The European Portuguese guide, which avoids address in the product, says to “use an informal tone of voice and form of address when translating Copilot predefined prompts” 5.

So the real unit is a language and a surface together: product interface, website, legal text, assistant prompts. None of that is visible in a string called settings.saved.

Ask for nothing and the engine picks for you

If nobody decides, the machine translation engine decides, and it does not decide the same way twice. Researchers at AWS had professional translators label the output of two general-purpose commercial systems on 300 random segments per language, with no formality setting applied. Each output was marked formal, informal, neutral, or “other”, which they used for “inconsistent formality or incorrectly omitted formality markers” 9.

Target languageSystem A, formal / informalSystem B, formal / informal
Spanish26.8% / 67.4%28.0% / 66.1%
French68.6% / 24.6%72.7% / 18.6%
Italian3.7% / 74.9%1.3% / 93.3%
Hindi81.7% / 3.2%87.7% / 5.2%
Japanese29.0% / 42.2%73.8% / 1.7%
Share of outputs labelled formal and informal by professional translators, two unnamed commercial MT systems, no formality control. Neutral and inconsistent outputs make up the rest. Nădejde et al., 2022, Table 4.

Read down the columns. Neither engine has a register. Each has a different habit per language: mostly informal in Spanish and Italian, mostly formal in French and Hindi. On Japanese the two systems disagreed outright, 29.0% formal against 73.8%, and more than 20% of each one’s Japanese output was labelled inconsistent or missing its formality markers 9.

The habit comes from the training data. The same paper found its own uncontrolled models leaning formal for several languages because “the generic training data biases the models towards formal” 9. The IWSLT 2023 organizers saw a regional version of it in Portuguese, pointing to the “dialectal influence of Brazilian Portuguese dominant in the pre-training corpora” 10. An engine left to itself gives you the register most common in text it has seen, which may not be the one your market expects.

These were conversational segments (customer support, chat and phone calls) run through 2022 systems the paper does not name. We found no published measurement of the same question on interface strings. The point that holds is the shape: the default differs by language and by vendor, and nobody chose it.

The formality switch exists, for some engines and some languages

Some engines let you ask. What happens when you ask differs more than the feature lists suggest.

EngineSettingValuesTarget language not supported
DeepLformalitydefault, more, less, prefer_more, prefer_lessThe request fails, unless a prefer_ value is used
Amazon TranslateFormalityFORMAL, INFORMALThe setting is ignored and AppliedSettings is null
Google Cloud Translation v3NoneNoneNot applicable
Azure Translator v3.0NoneNoneNot applicable
Azure Translator 2026-06-06targets.toneformal, informal, neutralNot stated on the page
Formality controls as documented on each vendor's API reference, read September 2026.

The two established switches fail in opposite ways. DeepL says that setting the parameter “with a target language that does not support formality will fail, unless one of the prefer_... options are used” 11. Amazon says that if it “doesn’t support formality level for the target language, or you don’t specify the formality parameter, the translation job ignores the formality setting”, and the response reports a null applied setting 12. Amazon lists 11 target locales, and Portuguese appears only as Portugal. A pipeline that asks for formal Brazilian Portuguese gets a translation back, no error, and whatever register the model happened to produce.

Google’s v3 translateText request has no formality field at all 13, and neither does Azure’s v3.0 translate method 14. Azure’s 2026-06-06 version adds a tone described as the “desired tone of target translation”. Its own reference page shows how loose “desired” is. One example asks for Spanish with "tone": "formal" and an adaptive dataset, and the response printed below it is “¿Quieres programar una cita?”, the informal form. The same formal request without the dataset, and the plain examples with no tone set at all, return the formal “¿Desea programar una cita?” 15. A documentation sample is not a benchmark, but it is the vendor showing a requested register that did not come back.

Where a switch exists, it is a request with a failure rate. The CoCoA-MT models, trained for exactly this, reached “82% in-domain and 73% out-of-domain” accuracy 9. On the out-of-domain call center data they produced the requested formal register 91.4% of the time and the requested informal register only about half the time, because the pull of the training data toward formal was still there 9. We have argued before that a request is not a constraint. A formality parameter is a request.

The switch also has two positions, and some languages have more. Amazon maps its two values onto Japanese as Kudaketa for informal and Teineigo for formal 12. The CoCoA-MT data itself carries a third Japanese level, respectful, with its own example translation 9. The IWSLT 2023 organizers found Korean harder to control than Vietnamese and put it down to “the variation in formality expression of Korean honorific speech” 10. Microsoft’s Japanese guide does not reach for a formality level at all. It says “you” can often be omitted, and otherwise should become “an appropriate word representing the target customer”, such as ユーザー or 管理者 for a role and お客様 in customer-facing text 16. That is a choice of noun, not a setting, and no two-value switch can express it.

The string file has nowhere to write the decision down

So the decision exists, in a style guide, usually a PDF. The files the strings actually travel in have no place for it.

XLIFF 2.1, the interchange standard, has no attribute for register or formality. The only place its specification mentions formality is an example of a free-text note 17:

text
<notes>  <note category="instruction" id="n1">The translation should be formal</note></notes>
Trimmed from an example in the XLIFF 2.1 specification. The note is prose for a person; nothing in the format gives it meaning.

A person can read that. An engine only sees it if someone copies it into the request. We made the general point about interchange formats when we wrote that a constraint nobody wrote down arrives as no constraint. Register is a very ordinary example.

The platforms show what it looks like when this kind of fact does get a home. Android 14 added an API so an app can store the user’s grammatical gender and pick matching resources 18. Apple’s Foundation has TermOfAddress, “the type for representing grammatical gender in localized text” 19. Both treat how to address the user as a runtime fact. Neither page mentions formality. Android’s own example of the gender problem is French, “Vous êtes abonné à...” against “Vous êtes abonnée à...” 18. It solves the gender and has already picked vous without saying so.

The quality frameworks agree that the error only exists against a written decision. MQM files register under Style, with a subtype for “using informal pronouns or verb forms when their formal counterparts are required”. Its note on the Style category is blunt: “Without style guides and proper specifications, such errors can give rise to discussions on subjectivity, as opposed to errors in Linguistic Conventions that are undisputedly right or wrong” 20. With no spec there is no register error, only two reviewers disagreeing.

MQM also names what happens when several people guess separately: the same text “translated in three different ways, which can involve different style as well as terminology or register differences”, which “often occurs due to multiple translators contributing to the target content” 20. A memory that holds both forms of the same sentence is the failure we described in same source, two targets, and it comes from the same place: a fact that nobody wrote down.

The case for writing around it

The strongest argument against all of this is that most interface strings never need the choice. The guides push that way. Mozilla asks Spanish translators to use impersonal forms “siempre que sea posible”, whenever possible 6. Microsoft steers European Portuguese away from você 5 and Japanese away from “you” 16. Buttons are verbs and menus are nouns. If most strings avoid address, a register slip is a rare style error, not an accuracy error, and a reviewer fixes it in seconds. That is a fair position, and for a small catalog in one or two languages it may be enough.

It fails in two places. First, leaving out the pronoun does not remove the register. Mozilla’s own example, “Elija una opción”, has no pronoun and is still formal; the informal imperative would be Elige. In Spanish, French and Italian the verb carries the choice. In Korean the sentence ending carries it. Any sentence that tells the user to do something has picked.

Second, “write around it where possible” is itself a per-language decision somebody has to make and write down, with a fallback for when it is not possible. It does not remove the decision. It moves it into the translator’s head, one string at a time, which is exactly how a catalog ends up with both registers in it.

Decide it per language, write it down, pass it on

A few habits make the decision once instead of a thousand times:

  • Decide the register per target language and per surface: product interface, website, legal text, assistant prompts. Write it with an example sentence, the way the style guides above do. For Japanese and Korean, write the level and the example, not the word “formal”.
  • Keep the decision where a pipeline can read it, next to the locale configuration, not only in a PDF. The translator, the reviewer and the machine translation call should all read the same value.
  • Pass it to the engine explicitly, and check that the engine supports it for that exact locale. One vendor rejects an unsupported language and another ignores the setting, and the second is the one that ships a register nobody chose. Log what was actually applied.
  • Treat the setting as a request. Sample the output per language and check the form of address held. A terminology check will not catch it, because no term is wrong.
  • When a memory or a catalog holds both forms for one source, treat it as a conflict to resolve, not as two valid variants.

The picture at the top is the whole problem. A top hat and a flat cap go into the funnel and one plain pebble comes out. English “you” is the pebble. Nothing downstream can tell which hat it came from, so somebody has to say.

References

  1. 1.Helmbrecht, 2013 Politeness Distinctions in Pronouns WALS Online, chapter 45
  2. 2.Microsoft French Style Guide Microsoft localization style guides
  3. 3.Microsoft Spanish (Spain) Style Guide Microsoft localization style guides
  4. 4.Microsoft Italian Style Guide Microsoft localization style guides
  5. 5.Microsoft Portuguese (Portugal) Style Guide Microsoft localization style guides
  6. 6.Mozilla Spanish, Spain (es-ES) Documentation for Mozilla localizers
  7. 7.Mozilla Hindi (hi-IN) Documentation for Mozilla localizers
  8. 8.Microsoft Korean Style Guide Microsoft localization style guides
  9. 9.Nădejde, Currey, Hsu, Niu, Federico and Dinu, 2022 CoCoA-MT: A Dataset and Benchmark for Contrastive Controlled MT with Application to Formality Findings of NAACL 2022
  10. 10.Agarwal et al., 2023 Findings of the IWSLT 2023 Evaluation Campaign IWSLT 2023
  11. 11.DeepL Translate text DeepL API reference
  12. 12.Amazon Web Services Setting formality in Amazon Translate Amazon Translate developer guide
  13. 13.Google Cloud Method: projects.translateText Cloud Translation API v3 reference
  14. 14.Microsoft Translator Translate Method (v3.0) Azure Translator reference
  15. 15.Microsoft Azure Translator 2026-06-06 translate method Azure Translator reference
  16. 16.Microsoft Japanese Style Guide Microsoft localization style guides
  17. 17.OASIS XLIFF Version 2.1 OASIS Standard, 2018
  18. 18.Android Developers Personalize your app's UI with grammatical gender Android 14 features
  19. 19.Apple TermOfAddress Foundation documentation
  20. 20.MQM Council The MQM Full Typology themqm.org