Skip to content
Built in-house · Toronto

An engine that hears
what the examiner hears.

Hilingo scores speaking and writing with its own engine: word by word, sound by sound, on each exam’s real scale. Not a rented pronunciation API with an exam label on top.

  • Pronunciation scored per sound, shown in IPA, with the fix
  • Stress, pitch, pace and pauses measured from your audio
  • Grammar, vocabulary range, cohesion and relevance for writing
  • PTE, IELTS, CELPIP, DET and TCF Canada today; OET, TEF and LanguageCert next
  • French scored by a French marker, not an English model pointed at French
Speaking · engine view
Sample
  1. Theðə
  2. weatherˈwɛð.ər
  3. isɪz
  4. changingˈtʃeɪn.dʒɪŋ
  5. quickly.ˈkwɪk.li
Inside “changing”, sound by sound
  • tʃ 96
  • eɪ 91
  • n 88
  • dʒ 61
  • ɪ 93
  • ŋ 90

The /dʒ/ came out closer to /z/. One sound, one fix.

Pitch and stresscontent words up, function words down
Pace and pauses2.9 words a second · one pause of 0.6 s
  • Pronunciation82
  • Fluency78
  • Content90
On each exam’s scale
  • PTE0/90
  • CELPIP0/12
  • DET0/160
A fixed sample of the engine’s layers on one spoken sentence. The same layers score PTE, CELPIP and DET; each exam gets its own scale.

What the engine hears.

Six measurements on every speaking answer, from a two-second reply to a two-minute talk. Pace, pitch, stress, content and fluency run on every exam; the word-by-word alignment and the IPA layer need a reference text, so they run on PTE Read Aloud and Repeat Sentence and on French read-aloud. Each one is shown to the student with the fix.

  1. 01

    Which words you actually said

    Your recording is transcribed and aligned to the reference word by word, so a skipped word, a replaced word and a mispronounced word are three different findings, not one lower number.

  2. 02

    Every sound, in IPA

    Each word is broken into its sounds and each sound is scored against the expected pronunciation. You see the expected IPA next to what you produced, so “changing” coming out with /z/ instead of /dʒ/ is a one-sound fix, not a vague “work on pronunciation”.

  3. 03

    Stress in the right places

    Content words should carry stress and function words should not. The engine reads pitch, loudness and length inside every word to check where your stress landed, and marks the function words you stressed by mistake.

  4. 04

    Pitch and voice

    Your pitch range and average pitch are measured across the answer. A flat, low delivery or an over-projected one gets called out, with the numbers.

  5. 05

    Pace and pauses

    Words per second, the long pauses, the hesitations and the restarts, timed from the audio itself. Fluency is scored the way the exam scores it: rhythm and continuity, not speed for its own sake.

  6. 06

    Whether you answered the question

    Content is checked against the key ideas of the prompt and its sample answer, not a keyword list. A fluent answer about the wrong thing scores as not answering the task, exactly as it would in the exam.

What the engine reads.

Writing is scored the way an examiner reads it: form first, then language, then whether it makes the case.

  1. Grammar and spelling, inline

    Every error is marked where it happens, with the correction, in the variant of English the exam expects.

  2. Six dimensions of writing

    Cohesion, syntax, vocabulary, phraseology, grammar and conventions are scored separately, so a strong argument with weak linking words is told exactly that.

  3. Vocabulary range

    How much of your vocabulary is repeated, how much is precise, and whether the level matches the task.

  4. Argument and relevance

    Essays and emails are judged on whether they make the case the prompt asked for, in the register the prompt asked for. Off-topic writing scores as off-topic.

  5. Meaning, not matching

    Summaries are compared with the source for meaning, so a paraphrase in your own words scores as well as it should, and copying a sentence does not.

  6. The form rules

    Word limits, single-sentence summaries, email conventions and capitalisation are enforced first, because the exam enforces them first.

Every exam on its own scale.

The measurements are shared. The rules on top are not: each exam has its own weights, its own scale and its own timers.

ExamScaleHow the score is built
PTE Academic and PTE CoreLive10–90 per skillEach task is scored against the criteria the exam publishes for that task type, then combined into skill scores and an overall on the exam’s scale. Pearson does not publish how it combines them, so the overall is our estimate rather than a reproduction of theirs.
CELPIP-General and General LSLiveLevels M to 12Writing and speaking are placed on the 12-level scale per task; listening and reading are marked per item and converted as the test converts them.
Duolingo English TestLive10–160 with subscoresEvery task feeds the four integrated subscores, and the overall follows from them.
IELTS Academic and General TrainingLiveBands 0–9 in half stepsWriting is marked on all four criteria and Speaking across Parts 1 to 3; Listening and Reading are marked per item and converted on the published indicative tables for the variant you sat.
TCF CanadaLive100–699 with CEFR, and out of 20 with NCLCCompréhension orale and écrite are marked per item and reported on 100–699 with the CEFR level beside them. Expression orale and écrite are scored out of 20 by a marker built for French, and shown with the NCLC level IRCC reads.
UK CAS Interview (pre-CAS interview practice)LiveStrong / Developing / Needs workThe same speech measurements from our engine (what you said, pace, pauses and voice variety), and an AI language model judges whether the answer covers its key points, and a separate rule-based check flags any amount, university, course, payer or plan after study that disagrees with the student’s profile. Each answer gets a band on 10 parts, and a mock gets Green, Amber or Red. Eye contact is measured in the browser; the video is never scored. There is no official scale, so these are bands, not a score.
OET, TEF Canada, LanguageCertNextSub-test grades, CEFR and NCLC levelsNext on the roadmap: the same engine with each test’s criteria and scale.

French is marked by a French marker.

An English essay model pointed at French is not a French marker. We measured the difference and rebuilt that path.

What went wrong with the English model

Scored on the same content written in both languages, at the same quality, the English-trained model marked the French version down across every dimension. It reads accents as noise and French discourse markers as padding, so it marks French writing down for reasons that have nothing to do with the writing.

What replaced it

Expression écrite is built from signals that are real in French: a French rule pack for grammar and conventions, sentence geometry, opener variety, lexical range, and French discourse markers for cohérence. Expression orale is transcribed and scored with French as the recognition language, so French speech is not read as broken English, and sound-by-sound IPA feedback is available on the French read-aloud drill. English scoring is untouched.

The NCLC levels are the published correspondence table, not a formula. We had been computing them from a single anchor, which disagreed with the official table on 7 of the 21 possible marks and erred upward at the top — reporting 14/20 as NCLC 10 when it is 9, and 15/20 as an NCLC 11 that does not exist. The real tables are in the product now, with every boundary pinned by a test. It is the number you copy onto an IRCC form, so it has to be the published one.

Why we built it instead of renting it.

Most practice platforms license a general speech-scoring API and put an exam label on the number. We started from the exams instead: what each task rewards, what each examiner listens for, and what a student needs to hear to fix it by tomorrow.

Owning the engine is what lets us show every layer, change a rule the week an exam changes, keep your audio on our own servers, and score without counting clicks.

See what the report shows
Rented scoring APIHilingo engine
Where the score comes fromA general pronunciation API, licensed per callAn engine we built for exam scoring and keep tuning per task type
Speaking feedbackA number per wordWords, sounds in IPA, stress, pitch, pace and pauses, with the fix
Off-topic answersOften still score for fluencyScore as not answering the task, like the exam
Exam scalesOne generic estimatePTE 10–90, CELPIP levels, DET 10–160, each with its own rules
Your recordings and essaysSent to a third partyStay on our servers, deleted on our schedule, never used to train anyone else’s product
Scoring limitsMetered by the vendor’s billUnlimited speaking and writing scoring on every paid plan
When the exam changesWait for the vendorWe change the rule ourselves

Feedback a teacher would give, in minutes.

  • For every answer

    Components on the exam scale, the words coloured by how they were heard, one tip per component — and on PTE and TCF read-aloud, the expected IPA next to yours.

  • For every mock

    Accuracy per task type, marks lost and where, time per question, the three things to fix first.

  • For the student

    An AI teacher that has already read the results and answers questions about them, in lessons, not chat.

  • For the teacher

    A one-page briefing per student: standing against the target, the problems costing marks with evidence, what to assign this week.

Questions about the scoring

Try it on a free mock

Or see the scale each exam uses: PTE, CELPIP, DET.

How accurate is the scoring?

Your reading and listening answers are marked against the answer key, item by item, so a right answer is right. The conversion from raw marks to a band or level is our own estimate, because the test owners do not publish theirs. Speaking and writing are scored per component following each exam’s published guide. We do not publish an accuracy percentage we cannot substantiate.

What do you mean by IPA-level feedback?

Every word is broken into its sounds, written in the International Phonetic Alphabet, and each sound is scored. You see the expected pronunciation next to what the engine heard, so you know which sound to fix rather than which word to repeat.

Does the same engine score CELPIP and DET?

Yes. The measurements are shared; the rules on top are per exam. CELPIP speaking and writing come out as levels, DET tasks feed the four subscores on the 10–160 scale, PTE tasks are scored on Pearson’s published per-task criteria and combined into an overall we report as an estimate.

Do you use a third-party speech API?

No. The speaking and writing engine is ours, run on our own servers. That is why we can show every layer of the analysis, change a rule when an exam changes, and offer unlimited scoring without a per-click bill.

Is an off-topic answer punished?

Yes. A fluent answer about the wrong subject scores as not answering the task, the same as in the exam. The content component is checked against the question and its sample answer first.

Can I see why I lost points?

Every scored answer shows its components, and every mock report ranks the task types where marks leaked with the questions that cost them. Teachers get a briefing built from the same data.