How We Test Language Apps (and What We Refuse to Claim)
We test language apps by hiring each one for a job and seeing whether it does that job — not by scoring everything against everything and averaging the soul out of it. This page is the whole method: what we weigh, where our facts come from, and the claims we refuse to make even when they'd convert better. Judge our roundups against it.
The five things we weigh
Every app on the bench gets examined on the same five axes, in this order of importance:
- Catalogue fit. Does it teach your language at all, by its own current course list? This kills more recommendations than any other test — Babbel, for instance, teaches no Korean, Mandarin or Japanese as of September 2026, which no amount of course quality fixes if Korean is why you came.
- Learning shape. Is it a built course, an audio method, a vocabulary drill, an AI conversation partner, a community? None of these is "best" — but each fits a different learner, and an app honest about its shape beats a better app fighting yours.
- Speaking reality. What actually makes you produce the language — out loud, from memory — and who or what checks you? Recognising a word in a multiple-choice grid is not speaking, and we say which side of that line each app's exercises live on.
- The free tier, truthfully. "Free" spans everything from Duolingo's complete ad-supported course to a single trial lesson. We report exactly what each vendor's own pages promise for nothing, dated, and we distinguish free from free-to-start every time.
- Price transparency. Not the price itself — the vendor's willingness to publish one. Most language-app prices hide behind app checkouts and regional tests; we note who publishes readable prices (in our current roster, only Rosetta Stone) and refuse to print remembered or second-hand numbers for the rest.
Where every fact comes from
Vendor facts come from vendor pages, read on a stated date — for the current roundups, 23 September 2026 — and each fragile fact keeps its date next to it in the text. When a vendor's site won't serve a fact to a plain page-load (Mondly's site blocks fetches outright; Babbel's prices render only in-app), the page says so instead of guessing. When a vendor contradicts itself, we print the contradiction rather than picking the flattering half: Ling's own site says "70+ languages" in one place and "80" in its footer, so we say "more than 70".
Claims about learning speed get one source only: the US Foreign Service Institute's published difficulty categories, which budget roughly 600–750 classroom hours for the languages closest to English (its Category I) and about 2,200 for its hardest group (Category IV — Arabic, Mandarin, Japanese, Korean). Those are figures for full-time diplomatic classroom study; we quote them as FSI's, for comparison between languages, never as a promise about your evenings. Marketing timelines — "conversational in three weeks" — appear only as quoted vendor claims, clearly labelled, or not at all.
What we refuse to publish
- Fluency timelines, retention percentages or "users who stuck with it" statistics we didn't source. An invented number reads like a citation, which makes it worse than no number.
- Single-app review pages and star scores. An 8.7-out-of-10 on a language app is theatre: the same app is a 9 for a Spanish commuter and a 2 for someone who needs Korean it doesn't teach. We publish roundups matched to jobs instead.
- Prices from memory or from other reviews. If we can't read it at the vendor's own page on a date we can print, it isn't here.
- Coupon codes and "exclusive discount" pages. They rot, they mislead, and they are the fastest way to stop being trustworthy about the thing that matters.
How the money works, and how we keep it out of the rankings
Some vendors on the bench run partner programs; their links are marked, routed through /go/, and can pay us a commission — the full list is in the affiliate disclosure. Two structural habits keep that honest. First, the free option is always evaluated and named, even though free earns us nothing: Duolingo — no affiliate program at all — appears in nearly every roundup we publish. Second, a vendor's troubles travel with its recommendation: we tell you on the page, mid-recommendation, that Busuu's parent company cut roughly 45% of staff in October 2025, because that belongs in your subscription decision more than in ours.
The bench in practice: one claim, start to finish
Here's how a single sentence earns its place. Suppose a draft says "Babbel teaches Korean." Step one: the claim gets typed against Babbel's own homepage course list, not a review, not memory. The list read on 23 September 2026 runs Spanish, French, German, Italian, Portuguese, Russian, Danish, Dutch, Indonesian, Norwegian, Polish, Swedish, Turkish — no Korean. The sentence dies, and its opposite is published with the date attached, because "Babbel doesn't teach Korean" is exactly the kind of negative fact shoppers need and marketing pages never volunteer.
Now a harder case: "Pimsleur teaches Korean." Very probably true — the vendor advertises more than fifty courses — but the specific Korean course didn't appear on any page we fetched, so the claim doesn't ship. Instead the Korean roundup says precisely that: likely, unconfirmed, check the vendor's catalogue. That distinction — between what we believe and what we verified — is the entire method in miniature. A site that only prints what it checked can be wrong about far fewer things, and when it is wrong, the date on the fact shows you exactly how it happened.
The same discipline runs the other way, protecting apps from us. We don't repeat viral complaints we can't source, we don't mark an app down for prices we never saw, and when a vendor does something well — Rosetta Stone publishing readable prices while every rival hides theirs — it gets credited by name even if a hidden-price rival would have paid us more. Symmetry is what makes scepticism trustworthy rather than just grumpy.
What this method can't do
It can't tell you whether you'll practise — the variable that dwarfs every difference between apps — and it can't inspect features vendors keep behind checkouts and regional pricing tests. What it can do is make sure that when you read our ranking of the best language learning apps, every claim in it is either dated to a source or absent — starting with the roundup this method matters most for, the genuinely free language apps, where the temptation to blur "free" is strongest.
Method questions
Do you actually use the apps you rank?
Why don't you publish scores out of ten?
Why are so many prices missing from your pages?
What would make you pull a recommendation?
Weighing more than one language, or starting from zero? The full bench sits in the best language learning apps, ranked — and how we test language apps shows the scale we weigh them on.