Skip to content
PTE

PTE speaking: the microphone is the only witness

Aman Batth · 16 min read · · Updated

There are seven speaking tasks on PTE Academic. Three of them pay into listening as well, and a fourth pays into listening instead — it cannot move your speaking score at all. That is the first thing a speaking guide should tell you. The second is what unites all seven: nothing in the room hears you. A file is captured, and the file is what gets marked.

Most "everything you need to know about PTE speaking" pages are a list of the seven tasks followed by thirty pieces of advice. The advice is usually fine. The list is usually wrong in one specific way, and it is the way that costs people study weeks: it presents all seven tasks as speaking tasks.

They are not. On PTE Academic, three of the seven pay into listening as well as speaking, and a fourth pays into listening instead of speaking. That is published, it is checkable, and it decides what "practising speaking" even means for you.

This page is the map and the one principle underneath it. The individual tasks each have their own guide, linked where they come up — this is not a replacement for those, it is the thing to read before you pick one.

We build the engine that marks speaking and writing answers on Hilingo, so this is written from the marking side rather than the coaching side.

The seven tasks, and which scores they actually feed

Timings are from Pearson's own PTE Academic speaking and writing format page. Item counts and the "communicative skills scored" column are from the PTE Academic Test Taker Score Guide.

TaskOn a testPrepareSpeakSkills it feeds
Read Aloud6–730–40 svaries with the textSpeaking
Repeat Sentence10–12—15 sListening + Speaking
Describe Image5–625 s40 sSpeaking
Retell Lecture2–310 s40 sListening + Speaking
Answer Short Question5–6—10 sListening
Summarize Group Discussion2–310 s2 minListening + Speaking
Respond to a Situation2–310 s40 sSpeaking

Read the last column rather than the first. Three things fall out of it.

Answer Short Question is a listening item. You hear a question, you say a word into a microphone, and on PTE Academic the mark lands in listening. There is no pronunciation score on it and no fluency score on it. If speaking is your low skill, this task cannot help you; if listening is, it is five or six listening items you have probably never drilled as a listening task. The full case is in PTE short answer questions.

Repeat Sentence is the largest task on the test by frequency, and half of it is listening. Ten to twelve items, more than any other question type, paying into two columns. Nothing else on the speaking part comes close for leverage.

Retell Lecture and Summarize Group Discussion are comprehension tasks wearing a microphone. They are scored for listening as well as speaking because that is what they measure. PTE Retell Lecture works through the consequence in detail: the content mark on that task is a listening score, which is why fluency work does not move it.

Only three of the seven — Read Aloud, Describe Image and Respond to a Situation — feed speaking and nothing else.

What changes on PTE Core

Different test, different map. The differences that matter for speaking, verified against the PTE Core Score Guide:

  • Read Aloud is scored for Reading and Speaking on PTE Core. On PTE Academic it is Speaking only. This single row is copied wrongly across most of the internet in both directions, so check the guide for the test you are actually sitting.
  • Answer Short Question is scored for Listening and Speaking on PTE Core, not listening alone.
  • Retell Lecture and Summarize Group Discussion do not exist on PTE Core.
  • Describe Image appears 3–4 times rather than 5–6.

If you are sitting Core for Canadian immigration, the whole task-to-skill table for both tests, with Pearson's page numbers, is in PTE marks distribution.

The three traits, and the one that is missing

Six of the seven tasks are scored on exactly the same three traits. Not similar traits — the same three, and two of the three on identical scales.

TraitScaleScored by
Content0–6 on Describe Image, Retell Lecture, Respond to a Situation and Summarize Group Discussion; 0–3 on Repeat Sentence; on Read Aloud the maximum depends on the length of the textAI, with human expert review on the four 0–6 tasks
Pronunciation0–5AI only
Oral Fluency0–5AI only

The seventh, Answer Short Question, has one trait: Vocabulary, worth a single point, correct or incorrect.

Now look at what is not in that table.

There is no separately named Vocabulary trait on any of the six open speaking tasks. That is not the same as word choice being unscored. It is scored — inside Content rather than beside it. The Content bands on Describe Image, Retell Lecture, Summarize Group Discussion and Respond to a Situation rank it directly: the top band wants "a variety of expressions and vocabulary ... with ease and precision", the middle bands accept a range sufficient for basic description but leaning on repetition, and the lowest ones call the range narrow, then limited, then highly restricted. Vocabulary as a trait of its own appears exactly once in the whole speaking part — as the one point on Answer Short Question, which on PTE Academic is a listening item.

Read those bands for what they measure, because it decides what is worth doing mid-answer, and this is the most expensive habit in this part of the test. They grade range and precision across the whole response — a variety of expressions used throughout, against simple expressions used repeatedly — not individual word choices. One upgraded noun does not carry a repetitive answer into the band above it. So the two seconds you spend reaching for a better word buy you almost nothing on Content, and they come straight off Oral Fluency, which is a five-point trait measuring exactly that.

The trade is one-way. This is the arithmetic behind the advice to keep going, and it is more precise than "be fluent".

The thing that unites all seven tasks

Here is the principle, and once you have it, most speaking advice either follows from it or falls apart.

Nothing in a PTE test centre listens to you. There is no examiner, no interview, no one to ask what you meant. What happens is that a microphone opens, a file is captured, and the file is scored. Everything else about your performance — that you knew the answer, that you were nervous, that the sentence was going somewhere good — exists only inside your own head, and your own head is not what gets marked.

Pearson publishes what the engine is — Versant for speech, locating and evaluating segments, syllables and phrases, with statistical models trained on expert raters' judgements — but not the model, the weights, or how the recorder behaves. This part is not a claim about any of that. It is what any recorded-and-marked assessment must be doing, because after you leave the room the recording is the only thing that still exists.

Four consequences, and each one is a real and common way to lose marks.

A perfect answer that arrives late scores nothing

The window opens and closes on a schedule. Fifteen seconds on Repeat Sentence, forty on Describe Image, ten on Answer Short Question. A well-formed sentence that lands after the recorder stopped is not a slightly worse answer. It is not on the file, so it did not happen. Answer Short Question makes the point sharpest: ten seconds, no partial credit, so a correct word that lands after the window is worth exactly what no word is worth.

Candidates who think while the microphone is open are spending marks on silence. The habit to build is starting immediately with something imperfect, because an imperfect thing on the tape outscores a perfect thing that is not.

A restart is genuinely less usable speech, not a clean slate

There is no undo. Beginning a sentence, abandoning it and starting again does not remove the abandoned fragment — it adds it, and it spends window on it.

Pearson's guide is explicit about one narrow case: on Repeat Sentence, hesitations, filled or unfilled pauses, and leading or trailing material are ignored when scoring Content. Read what that carve-out implies. It is one trait on one task. The guide does not extend it, and Oral Fluency is by definition a measure of rhythm and pausing, so a restart lands there whatever it does to Content. That inference is ours; the carve-out is Pearson's.

And Oral Fluency's published bands count occurrences rather than seconds: one band allows no hesitation, repetition or false start, the next allows one, the one below that allows two or three. A single false start is the step between them, however long it took.

Hesitating to find a better word costs more than the better word gains

Covered above from the trait side; here is the same thing from the recording side, and it applies to the six open tasks. The file captures duration, and silence is duration with nothing in it. Two seconds of deliberation on a forty-second Describe Image is a twentieth of the answer, and it buys one better word inside a band that is grading the whole response.

The instruction that falls out is narrower than "be fluent". It is: say the first acceptable word and move on.

Answer Short Question is the one place this reverses. There the word is the entire mark, scored right or wrong on a single Vocabulary point, so accuracy is worth more than speed — on that task and nowhere else in speaking.

You cannot hear your own pronunciation

This is the one that makes speaking feel unfixable from the inside, and it is not a confidence problem.

When you speak, you know which sound you were aiming for. Your ear then reconciles what arrived with what you intended, and the intention wins. On playback you can hear your own grammar mistakes perfectly well, because grammar lives in the words and the words are recoverable. You cannot hear your own vowel, because you already know which vowel you meant.

A machine has no such access. It has the waveform. Pronunciation and Oral Fluency are the two traits Pearson's guide marks AI-scored only. Content on the four 0–6 tasks adds a human expert review, with a second human settling it when the AI and the human disagree; no equivalent review is published for the other two. Humans are not absent from speaking altogether — they train the engine on every speaking response, and they score live recordings the machine cannot handle, the guide's example being ones that are quiet or unclear. But on a recording that scores normally, Pronunciation and Oral Fluency are the machine's alone. They are also the two traits you are least able to assess for yourself.

That is a measurement gap, not a knowledge gap, and it is the reason candidates reach for scripts. A script is the one part of the answer you can control by memorising. It is also the part that scores least — see why PTE templates do not work for the Content-zero gate that makes a memorised answer worse than a clumsy real one.

A word about the "3-second rule"

You will read on prep sites that Pearson follows a three-second rule: go silent for three seconds and the item advances. The number gets repeated confidently, usually with no source.

It is not in the Test Taker Score Guide, and Pearson does not publish the behaviour of the recorder. We are not going to assert someone else's internals as fact, and neither should the page you read it on.

What you can do is notice that it changes nothing. Whether or not a timer exists, a silence is empty duration on a file that is scored partly on rhythm, inside a window you cannot extend. Start immediately, keep going, finish before the cut. That instruction is identical under both versions of the world, which is a reasonable sign it is the right instruction.

What to actually check, when you cannot hear it yourself

Three checks, in this order, on a recording of your own voice. They work on any task with an open microphone.

  1. Read the transcript, not your memory of what you said. Mark every word that could only have come from this prompt — a number, a name, the specific claim. If there are three or four, your Content is the problem and no amount of delivery work will move it.
  2. Count the pauses and find where they cluster. Pauses in the first seconds mean you did not fix your opening during preparation. Pauses late mean you ran out of material, which is an upstream problem.
  3. Check whether your last clause finished. Cut off mid-clause means your time budget is wrong, not your English.

Check three you can do by ear. Check two you can roughly approximate. Check one you cannot do at all on your own, because transcribing yourself and listening to yourself are the same act — you will write down the word you meant.

What our report gives you, and what it is not

This part is ours, and we label it that way deliberately.

On PTE Read Aloud, a Hilingo answer comes back with the transcript, your speaking pace, a pause count, and a per-word pronunciation breakdown in phonetic notation — the sound the word is supposed to carry, set against the sound the engine detected in your recording, word by word. When a word scores low, you can see which segment of it did.

That last one is the specific blind spot described above. It is the only way we know of to inspect a sound you are structurally unable to hear, because it converts an acoustic judgement into something you can read.

Two honest limits.

On the open tasks the phonetic breakdown is comparing something different. It runs on all of them — Describe Image, Retell Lecture, Summarize Group Discussion and Respond to a Situation get the same per-word expected-against-detected rows as Read Aloud, because the alignment is to the dictionary pronunciation of whatever you actually said rather than to a script. The difference is what a row can tell you. On Read Aloud and Repeat Sentence there is a sentence you were given, so a row can say you produced the wrong word. On the open tasks there is no such reference: the rows say how cleanly you pronounced the words you chose, and nothing about whether they were the right ones. The practical answer is to learn your own habits on Read Aloud, where the reference makes them unambiguous, then carry what you found to the open tasks.

This is our engine's report, not Pearson's. Nobody outside Pearson runs Pearson's engine, and any platform telling you its number is the same number is telling you something it cannot know. What we will say is what the method is: we score the traits the published guides name, separately rather than collapsed into one figure, with the evidence shown so you can disagree with it. The band boundaries are our own, calibrated against real scored recordings rather than derived from Pearson's descriptors.

Where to start

If you do not yet know which of your four scores is the low one, start there — a self-assessment will not tell you, and practising the wrong column is how a preparation month goes by with nothing to show for it.

Hilingo gives you one full scored PTE mock free, with no card. Beyond that there is a paid library of full-length mocks and section tests for both PTE Academic and PTE Core — a section test being one section of a full paper offered on its own, so a shorter sitting rather than an extra paper. On a paid plan you can retake them as often as you like, which matters more on speaking than anywhere else, because the drill that works is the same item repeated until the habit changes. There is a single-question practice mode alongside the full mocks, so you can sit six Read Alouds back to back rather than a whole paper to reach six. Objective questions come back with a written reason the answer was wrong, with the evidence sentence quoted where the passage supplies one, speaking and writing are marked per criterion rather than as one number, and any question can be shown in 50 languages if the English of the prompt is getting in the way of the English being tested.

Take a free scored mock

Frequently asked questions

How is PTE speaking scored?

Six of the seven speaking tasks are scored on three traits: Content, Pronunciation on a 0 to 5 scale, and Oral Fluency on a 0 to 5 scale. The seventh, Answer Short Question, is scored on a single Vocabulary point, correct or incorrect. Content on Describe Image, Retell Lecture, Respond to a Situation and Summarize Group Discussion is reviewed by a human expert alongside the AI, with a second human deciding if they disagree; Pearson's guide states that Pronunciation and Oral Fluency are AI-scored only, though human scorers still handle live recordings the machine cannot score, such as quiet or unclear ones. You never see any of these traits on your report — Pearson removed the old enabling skills in November 2021, and what a test taker receives now is one overall score, four communicative skill scores, and a Skills Profile grouping performance into eight descriptive categories that only you see.

Which PTE tasks count towards the speaking score?

On PTE Academic, six of the seven. Read Aloud, Describe Image and Respond to a Situation feed speaking alone. Repeat Sentence, Retell Lecture and Summarize Group Discussion feed listening and speaking together. Answer Short Question feeds listening only and cannot move your speaking score at all. PTE Core differs: there Read Aloud is scored for reading and speaking, Answer Short Question for listening and speaking, and Retell Lecture and Summarize Group Discussion do not exist.

Can I restart my answer in PTE speaking?

There is no undo. Starting again does not remove what you already said; it adds a fragment and spends part of a window you cannot extend. Pearson's guide does say that on Repeat Sentence, hesitations and leading or trailing material are ignored when Content is scored — but that is one trait on one task, and Oral Fluency measures rhythm and pausing on every task that carries it. Carrying on from where you are is almost always cheaper than going back.

Does pausing reduce your PTE speaking score?

Oral Fluency is a 0 to 5 trait on six of the seven speaking tasks, and it is measuring the rhythm of what is on the recording. Silence is time with nothing in it. The habit that costs most is pausing to find a better word. Word choice is not a separate trait on the open tasks; it is graded inside Content, as the range and precision of the whole answer rather than as any one upgraded word — so a single better word gains you very little while the pause costs something.

Is there a 3-second rule in PTE speaking?

It is widely repeated on preparation sites and it is not stated in Pearson's Test Taker Score Guide. Pearson does not publish how the recorder behaves, so treat the number as folklore rather than fact. It also makes no practical difference: whether or not a timer exists, a long silence is empty duration inside a fixed window on a response scored partly for fluency. Start immediately and keep going either way.

How can I check my own PTE pronunciation?

Not by listening to yourself, which is the real problem. You know which sound you intended, and your ear reconciles the recording with the intention. You need the judgement converted into something you can read rather than hear — a transcript, a pause count, and a per-word comparison of the expected sounds against the sounds detected in your recording. Hilingo produces that on every open speaking task, and on Read Aloud the sentence you were given makes the comparison sharpest. It is our engine's assessment, not Pearson's; nobody outside Pearson runs theirs.

Practise with the engine this article describes.
One full scored mock, free. No card.
Take a free mock

All articles