AI-Powered · 120+ Languages

TXT Translator

Plain text carries no markup to protect and no schema to follow, which makes it the simplest thing to translate and the easiest to quietly wreck. Upload manuscripts, notes, transcripts or exports and get them back with the character encoding, the line breaks and the indentation intact.

Max. file size 1 GB Keeps original formatting
Sign Up Free

Upload or drop document to translate

Max. file size 1 GB

.PDF .DOCX .PPTX .XLSX .TXT .JPG .PNG .IDML .EPUB .HTML
Afrikaans (Afrikaans)
Shqip (Albanian)
አማርኛ (Amharic)
العربية (Arabic)
Հայերեն (Armenian)
Azərbaycan dili (Azerbaijan)
Euskara (Basque)
Беларуская (Belarusian)
বাংলা (Bengali)
Bosanski (Bosnian)
Български (Bulgarian)
မြန်မာဘာသာ (Burmese)
Català (Catalan)
Cebuano (Cebuano)
Chichewa (Chichewa)
中文 简体 (Chinese Simplified)
中文 繁體 (Chinese Traditional)
Corsu (Corsican)
Hrvatski (Croatian)
Čeština (Czech)
Dansk (Danish)
Nederlands (Dutch)
English (English)
Esperanto (Esperanto)
Eesti (Estonian)
Suomi (Finnish)
Français (French)
Frysk (Frisian)
Galego (Galician)
ქართული (Georgian)
Deutsch (German)
Ελληνικά (Greek)
ગુજરાતી (Gujarati)
Kreyòl Ayisyen (Haitian)
Hausa (Hausa)
ʻŌlelo Hawaiʻi (Hawaiian)
עברית (Hebrew)
हिंदी (Hindi)
Hmoob (Hmong)
Magyar (Hungarian)
Íslenska (Icelandic)
Igbo (Igbo)
Bahasa Indonesia (Indonesian)
Gaeilge (Irish)
Italiano (Italian)
日本語 (Japanese)
Basa Jawa (Javanese)
ಕನ್ನಡ (Kannada)
Қазақ тілі (Kazakh)
ខ្មែរ (Khmer)
Ikinyarwanda (Kinyarwanda)
한국어 (Korean)
Kurdî (Kurdish)
Кыргызча (Kyrgyz)
ລາວ (Laotian)
Latina (Latin)
Latviešu (Latvian)
Lietuvių (Lithuanian)
Lëtzebuergesch (Luxemb)
Македонски (Macedonian)
Malagasy (Malagasy)
Bahasa Melayu (Malay)
മലയാളം (Malayalam)
Malti (Maltese)
Te Reo Māori (Maori)
मराठी (Marathi)
Монгол хэл (Mongolian)
नेपाली (Nepali)
Norsk (Norwegian)
ଓଡ଼ିଆ (Odia)
فارسی (Persian)
Polski (Polish)
Português (Portuguese)
ਪੰਜਾਬੀ (Punjabi)
Română (Romanian)
Русский (Russian)
Gagana Samoa (Samoan)
Gàidhlig (Scottish)
Српски (Serbian)
Sesotho (Sesotho)
Shona (Shona)
سنڌي (Sindhi)
සිංහල (Sinhala)
Slovenčina (Slovakian)
Slovenščina (Slovenian)
Soomaali (Somali)
Español (Spanish)
Basa Sunda (Sundanese)
Kiswahili (Swahili)
Svenska (Swedish)
Tagalog (Tagalog)
Тоҷикӣ (Tajik)
தமிழ் (Tamil)
Татарча (Tatar)
తెలుగు (Telugu)
ไทย (Thai)
Türkçe (Turkish)
Türkmençe (Turkmen)
Українська (Ukrainian)
اردو (Urdu)
ئۇيغۇرچە (Uyghur)
O'zbekcha (Uzbek)
Tiếng Việt (Vietnamese)
Cymraeg (Welsh)
isiXhosa (Xhosa)
ייִדיש (Yiddish)
Yorùbá (Yoruba)
isiZulu (Zulu)
Afrikaans (Afrikaans)
Shqip (Albanian)
አማርኛ (Amharic)
العربية (Arabic)
Հայերեն (Armenian)
Azərbaycan dili (Azerbaijan)
Euskara (Basque)
Беларуская (Belarusian)
বাংলা (Bengali)
Bosanski (Bosnian)
Български (Bulgarian)
မြန်မာဘာသာ (Burmese)
Català (Catalan)
Cebuano (Cebuano)
Chichewa (Chichewa)
中文 简体 (Chinese Simplified)
中文 繁體 (Chinese Traditional)
Corsu (Corsican)
Hrvatski (Croatian)
Čeština (Czech)
Dansk (Danish)
Nederlands (Dutch)
English (English)
Esperanto (Esperanto)
Eesti (Estonian)
Suomi (Finnish)
Français (French)
Frysk (Frisian)
Galego (Galician)
ქართული (Georgian)
Deutsch (German)
Ελληνικά (Greek)
ગુજરાતી (Gujarati)
Kreyòl Ayisyen (Haitian)
Hausa (Hausa)
ʻŌlelo Hawaiʻi (Hawaiian)
עברית (Hebrew)
हिंदी (Hindi)
Hmoob (Hmong)
Magyar (Hungarian)
Íslenska (Icelandic)
Igbo (Igbo)
Bahasa Indonesia (Indonesian)
Gaeilge (Irish)
Italiano (Italian)
日本語 (Japanese)
Basa Jawa (Javanese)
ಕನ್ನಡ (Kannada)
Қазақ тілі (Kazakh)
ខ្មែរ (Khmer)
Ikinyarwanda (Kinyarwanda)
한국어 (Korean)
Kurdî (Kurdish)
Кыргызча (Kyrgyz)
ລາວ (Laotian)
Latina (Latin)
Latviešu (Latvian)
Lietuvių (Lithuanian)
Lëtzebuergesch (Luxemb)
Македонски (Macedonian)
Malagasy (Malagasy)
Bahasa Melayu (Malay)
മലയാളം (Malayalam)
Malti (Maltese)
Te Reo Māori (Maori)
मराठी (Marathi)
Монгол хэл (Mongolian)
नेपाली (Nepali)
Norsk (Norwegian)
ଓଡ଼ିଆ (Odia)
فارسی (Persian)
Polski (Polish)
Português (Portuguese)
ਪੰਜਾਬੀ (Punjabi)
Română (Romanian)
Русский (Russian)
Gagana Samoa (Samoan)
Gàidhlig (Scottish)
Српски (Serbian)
Sesotho (Sesotho)
Shona (Shona)
سنڌي (Sindhi)
සිංහල (Sinhala)
Slovenčina (Slovakian)
Slovenščina (Slovenian)
Soomaali (Somali)
Español (Spanish)
Basa Sunda (Sundanese)
Kiswahili (Swahili)
Svenska (Swedish)
Tagalog (Tagalog)
Тоҷикӣ (Tajik)
தமிழ் (Tamil)
Татарча (Tatar)
తెలుగు (Telugu)
ไทย (Thai)
Türkçe (Turkish)
Türkmençe (Turkmen)
Українська (Ukrainian)
اردو (Urdu)
ئۇيغۇرچە (Uyghur)
O'zbekcha (Uzbek)
Tiếng Việt (Vietnamese)
Cymraeg (Welsh)
isiXhosa (Xhosa)
ייִדיש (Yiddish)
Yorùbá (Yoruba)
isiZulu (Zulu)
ARABIC PORTUGUESE RUSSIAN ITALIAN KOREAN DUTCH POLISH TURKISH SWEDISH ENGLISH SPANISH FRENCH GERMAN CHINESE JAPANESE HINDI BENGALI VIETNAMESE THAI GREEK HEBREW ARABIC PORTUGUESE RUSSIAN ITALIAN KOREAN DUTCH POLISH TURKISH SWEDISH ENGLISH SPANISH FRENCH GERMAN CHINESE JAPANESE HINDI BENGALI VIETNAMESE THAI GREEK HEBREW

The extension promises almost nothing

Every other document format carries a description of itself. A word processor document names its fonts and declares its character set. A page of markup wraps its content in tags that say what each part is. Plain text declares none of that. It is a run of bytes with an agreement that somebody will read them as characters, and the whole agreement is unwritten.

Three things therefore have to be worked out rather than read: which character encoding the bytes belong to, which sequence the author used to end a line, and whether the line breaks are meaningful or merely the width of somebody's editor window. Software guesses all three every time you open such a document, and it is usually right, which is precisely why the failures are surprising when they come.

The upside is real. There is no styling to preserve, no layout engine to fight, no embedded object to lose. If the encoding and the line structure survive, everything has survived. The rest of this page is about those two things.

What people actually keep in plain text

The extension gets attached to things that are not prose at all, and each of these needs a different kind of care. Knowing which you are holding decides whether the job takes a minute or needs preparation.

  1. Continuous prose. Manuscripts, notes, article drafts, interview transcripts, scraped article bodies. This is the straightforward case, and the only thing to watch is whether paragraphs are separated by a blank line or by nothing at all.
  2. Subtitles and captions. Numbered cues, a timing line and one or two lines of dialogue. The timings are coordinates and must be reproduced exactly, while the dialogue lines have to stay short enough to read at speed, which matters because translated lines are often longer than the ones they replace.
  3. Settings and property lists. A key, a separator and a value on each line. The key on the left is an identifier that software looks up by name; touching it silently disables the setting. Only the value on the right, and only where that value is a message rather than a path or a switch, is language.
  4. Logs and machine output. Timestamps, severity words, identifiers and stack traces. Translating a severity word breaks any filter that greps for it, and a stack trace is not prose in any language.
  5. Tabular exports in disguise. Values separated by tabs, saved with the wrong extension. If your content is really a grid, it belongs in the CSV translator instead, which understands records and columns.
  6. Read-me notes and licence text. Prose with headings faked out of capital letters and rows of dashes, numbered clauses, and hard-wrapped paragraphs. The faked structure is the part that breaks.

Three invisible properties that decide the outcome

None of these is visible when you look at the content on screen, and each of them can turn a good translation into an unusable one.

Which encoding the bytes belong to

Before Unicode became the default, every language region had its own single-byte or multi-byte scheme: separate ones for Cyrillic, for Western European accents, for Greek, for Japanese, for simplified and traditional Chinese, for Korean. Documents written under those schemes are still in circulation, and nothing in them says which one applies. Open one under the wrong assumption and the accented characters turn into runs of nonsense while the plain letters look fine, so the damage is easy to miss on a quick glance. Check in an editor that lets you choose the encoding explicitly, and reopen until the accents and the non-Latin characters read correctly.

How a line ends

Windows tools end a line with two control characters, everything descended from Unix uses one, and very old Macintosh documents used the other one on its own. Most modern editors accept all three, but plenty of processing scripts do not, and mixing conventions inside one document is what produces stray marks at the end of every line or a document that displays as a single unbroken paragraph. Keep one convention throughout, and prefer whichever the system that will consume the result expects.

Whether a line break means anything

Text wrapped by hand at a fixed width has a break at the end of every line, and those breaks are an artefact of the width rather than part of the writing. A paragraph is normally marked by an empty line between blocks. Treating each wrapped line as a unit chops sentences into fragments and translates them out of context, which reads badly in any language. Keeping the paragraph as the unit is right, and it means the output will rewrap at different points from the original. That is expected, not a defect.

Anything aligned with spaces will come out misaligned

Why the columns drift

Without styling, the only way to line anything up is to count characters. Tables drawn with pipes and dashes, headings underlined with rows of equals signs, ledgers padded out with spaces, indented outlines and boxes made of line-drawing characters all depend on every entry occupying the same width it did before.

Translation changes the width of nearly everything. German and Russian commonly run a fifth to a third longer than English; Chinese and Japanese take fewer characters but each occupies two columns in a fixed-width display. Either way the alignment moves. Expect to repad hand-drawn tables afterwards, and consider whether the content wants a real spreadsheet instead.

Leading whitespace is content

In a settings document, in a snippet of code, and in any outline, the spaces at the start of a line carry meaning: they show nesting, or they mark a block as literal. A tab and a run of spaces look identical on screen and are different bytes, so a tool that normalises one into the other changes the document without appearing to.

The same applies to the small escape sequences written as a backslash and a letter, which stand for a line break or a tab inside a single line of a settings entry. They are two ordinary characters that some later program will interpret, and they have to arrive on the other side unchanged.

Cost for a plain text job

Charged on the words in the document, so a long log with a few unique messages costs far less than its size suggests.

7-Day Trial

MOST POPULAR
$2.00 today

then $14.99/month after trial ends

  • 7-day full access trial
  • Trial limit: 10 pages or 3,000 words
  • $0.005/word AI translation
  • 120+ languages
  • PDF, DOCX, XLSX, PPTX, IDML, TXT, JPG, PNG, CSV, JSON
  • Team access & custom glossaries
  • Email support

Monthly

POPULAR
$14.99/month

Regular price $29.99, now 50% off

  • 100 pages or 30,000 words per month
  • $0.005/word AI translation
  • 120+ languages
  • Unlimited file storage
  • PDF, DOCX, XLSX, PPTX, IDML, TXT, JPG, PNG, CSV, JSON
  • Team access & custom glossaries
  • Priority email support
🎉 Best value: save $44.88/year

Annual

SAVE 25%
$135/year

~$11.25/month, save 25% vs monthly

  • 100 pages or 30,000 words per month
  • $0.005/word AI translation
  • 120+ languages
  • Unlimited file storage
  • PDF, DOCX, XLSX, PPTX, IDML, TXT, JPG, PNG, CSV, JSON
  • Team access & custom glossaries
  • Priority email support

Sample documents, side by side

Click a thumbnail to open the full view. Reading across a pair shows:

  • Paragraph breaks and blank-line spacing land where they did before.
  • Accented and non-Latin characters render rather than degrading into symbols.
  • Indentation and numbered sequences keep their nesting.
  • The result opens as editable text, not as a picture of text.
A plain text document in English, the sourceThe same plain text document rendered in PortugueseA second plain text document in English, the sourceThe same second plain text document rendered in Spanish

Each pair came from a single upload with no manual repair afterwards.

See all translation examples →

Five minutes of preparation, then the upload

Nearly every complaint about plain text translation traces back to something that was already true of the document before it was uploaded. These checks catch all of them.

  1. 1

    Confirm the encoding reads correctly

    Open the document in an editor that shows and lets you change the encoding. If accented or non-Latin characters look like sequences of unrelated symbols, you are reading it under the wrong scheme. Fix that first and save a Unicode copy, because no translation can recover characters that were already misread.

  2. 2

    Decide what is prose and what is not

    Where the content is settings, logs or subtitles, identify the parts that must stay untouched: keys on the left of a separator, severity words, cue numbers and timing lines. If most of the document is machine content with a little prose in it, extracting the prose first is faster than repairing the result.

  3. 3

    Create an account and upload

    Sign up with an email address and send the document in. Single uploads run to 1 GB, which covers manuscripts, transcript archives and log exports without splitting them into parts.

  4. 4

    Set both languages and translate

    Naming the source matters more here than in a formatted document, because there is no markup and no metadata for the system to infer it from. Then choose the target and start the job.

  5. 5

    Reopen the result and check the same three things

    Encoding, line endings, and whether anything aligned by hand needs repadding. If the document is going into a build pipeline or a media player, run it through that pipeline once before you rely on it.

TXT translation FAQ

Common questions about plain text

My document came back full of strange symbols instead of accented letters. Why?

That is an encoding mismatch and it almost always predates the translation. Bytes written under an older regional scheme, read as though they were Unicode, produce exactly that pattern: plain letters look correct while accented and non-Latin characters turn into runs of unrelated symbols. Open the original in an editor that lets you pick the encoding, find the one that renders it correctly, save a Unicode copy and start from there.

Will my line breaks be in the same places afterwards?

Paragraph boundaries will be. Individual wrapped lines will not, and cannot be, because the translated wording occupies a different number of characters. Where the original was hand-wrapped at a fixed column width, the result will break at different points. If a specific width matters for where the document is going, rewrap it after translation with the tool that consumes it.

Can I translate a subtitle document this way?

Yes, and two things need attention. The cue numbers and the timing lines are coordinates rather than words and have to survive exactly, and translated dialogue is often longer than the original, which can push a line past what a viewer can read in the time available. Review the longest lines against the timings and shorten where a cue has become too dense.

What about a settings or properties document?

Only the right-hand side is language, and not all of it. The key before the separator is an identifier that software looks up by name, so translating it disables the setting silently. Values that are file paths, switches, numbers or references to other keys are not prose either. Where a document is mostly keys with a few messages, pulling the messages out first is quicker than checking the whole result.

My hand-drawn table lost its alignment. Can that be avoided?

Not in plain text, because alignment there is made of counted spaces and translation changes how many characters a phrase occupies. German and Russian typically expand, Chinese and Japanese contract in characters while each one takes two display columns. Repadding afterwards is the fix. For anything genuinely tabular, a spreadsheet or a comma-separated export is a better home.

How large a document can I upload?

A single upload can reach 1 GB or five thousand pages of equivalent content on a paid plan, so full-length manuscripts, transcript collections and long exports go through in one piece. Sending it whole rather than in chunks also keeps recurring terminology consistent from the first page to the last.

Is the result editable, or an image?

Editable. What comes back is a text document you can open in any editor and keep working in, with the same character content and structure as the original. Nothing is rasterised and nothing is locked.

What does it cost for a long document?

Pricing follows the word count rather than the byte count. Seven days of access costs $2 and allows 10 pages or 3,000 words; a month costs $14.99 for 100 pages or 30,000 words; a year costs $135. Work beyond an allowance is billed at $0.005 per word, so a manuscript can be estimated straight from its word count before you start.

Send the document, keep the structure

Check the encoding, upload, choose the languages, and get back editable text with the paragraphs, the indentation and the non-Latin characters where they belong.

Our Partners

Accenture
Bloomberg
Citrix
P&G
SAP