Your text

Detection

Loading the language profiles (twelve Latin-script languages, hand-curated function words, letters and endings). The data is imported on this page only.

What this does and cannot do

The profile table itself is shown above: every point comes from a listed function word, a distinctive letter, or a characteristic ending. There is no model, no network call and no stored text — the profiles load in the browser and your input never leaves it.

Short samples are refused rather than guessed: at least 5 words and 24 letters are required, and the result is reported with a confidence band. Text that is mostly another script (Cyrillic, Han, Arabic, …) is reported as unsupported, because only Latin-script languages are profiled.

Closely related languages are deliberately narrowed to one profile per pair — Swedish but not Danish or Norwegian, Czech but not Slovak — because scoring cannot separate them reliably. Proper nouns, loanwords, and text that switches language midway can still mislead any profile scorer.

This language detector scores a sample against twelve hand-built Latin-script profiles and shows the evidence behind the answer — which function words, distinctive letters and word endings matched, and by how much the winner led.

What is Language Detector?

A language detector usually answers with a name and nothing else. This one shows its work: paste a sample and the text is scored against twelve Latin-script profiles — English, Spanish, French, German, Italian, Portuguese, Dutch, Swedish, Polish, Czech, Turkish and Finnish — and the full table is printed with the result.

All of the language detection here is arithmetic on three kinds of evidence a reader can inspect. Function words are the first: closed-class words a language cannot avoid, such as the, der, que and och. Distinctive letters are the second: ñ, ß, ı and ł belong to one orthography far more than to its neighbours. Word endings are the third: zione, ung and ção carry a language's shape even in a short fragment. Each matched function word is worth two points, each distinctive letter one and a half, and each ending one.

The weighting matters. Function words are the strongest signal in ordinary prose, letters survive in names and one-line labels where grammar cannot show, and endings rescue samples that are too terse for either. The score table lists every candidate language with its three evidence columns, so when English wins on the, and and of, you can see exactly why rather than take it on faith. There is no model, no training data and no network call: the profiles load lazily on this page only, and your text never leaves the browser.

What the tool refuses to do is as important as what it does. Fewer than five words or twenty-four letters and it answers "not enough text" instead of offering the three least-bad guesses. Text that is mostly Cyrillic, Han, Arabic or another script is reported as unsupported rather than forced into a Latin profile. When the top two languages are close, the confidence band drops to medium or low and the notes name the runner-up, so a near-tie is labelled rather than hidden.

The language list is deliberately narrower than a full identifier could be. Closely related pairs — Danish and Norwegian, Slovak and Czech — share most of their function words, so a profile scorer cannot separate them with the same confidence it separates German from Spanish. Rather than add profiles that would amount to a coin toss, one member of each pair is included and the limits are stated on the page. Every profile is described in the source; adding a language means curating the same three kinds of evidence for it.

Read the limits before you rely on the result. Proper nouns, loanwords and brand names can be mistaken for the language they look like; a single word is never enough, which is why short input is refused; text that switches language midway is scored as one sample; and transliterated or heavily misspelled text confuses any profile scorer. The detector says nothing about dialect, formality or quality. It answers one narrow question — which of twelve languages does this sample most resemble — and attaches the evidence.

In practice that covers the cases the tool exists for: checking a pasted paragraph to detect the language, triaging mixed-language content, or confirming that a translation is really in the language it claims. Read the confidence beside the name, then the evidence table below it. Where the answer matters legally or operationally, treat a medium or low result as a prompt to find a native reader rather than as a decision.

How to use Language Detector

  1. Paste a sentence or two of ordinary prose — five words and twenty-four letters is the minimum for a result.
  2. Read the detected language and its confidence band; medium or low names the runner-up in the notes.
  3. Check the evidence table: function-word hits are the strongest column, letters and endings support it.
  4. Treat refused samples as refusals, not failures — add more text if you need a verdict.
  5. Copy or download the report when you need the score table alongside the text.

When to use Language Detector vs related tools

Use this detector when the question is what language a sample is written in, and reach for different tools when the question is about the text itself. The Spell Checker checks words against a real English dictionary once you know the sample is English, and the Character Encoding Converter answers the adjacent question of what bytes the text is made of.

If a sample is refused as too short, a word-count tool will not help — the detector needs prose, not statistics.

Privacy & Security

This tool runs entirely in your browser — no data ever leaves your device. There is no server round-trip, no upload, no logging, and no account required. Your input is processed locally using client-side JavaScript and is never stored, transmitted, or accessible to anyone else. When you close the tab, everything disappears.

Frequently asked questions about Language Detector