CELPIP Speaking: eight long turns, and nobody to rescue you
Aman Batth · 17 min read · · Updated
Every page about this section prints the same eight task names. The fact that decides your level is not on the list — there is no examiner. You get a prompt, a short prep window and a fixed recording window, and if you dry up at thirty seconds the machine simply records the remaining thirty.
Search "CELPIP speaking test" and you get the same eight task names, usually in the same order, usually in a table someone copied from someone else. Giving advice, personal experience, describing a scene, predictions, comparing and persuading, difficult situation, opinions, unusual situation.
The list is correct. It is also almost useless, because it does not contain the thing that makes this section hard.
Here it is: there is nobody to talk to.
Not "it is computer-delivered", which is a format detail. The consequence. A prompt appears, a preparation clock runs down, the screen moves on by itself, a voice says "Start speaking now", and recording begins — the free Speaking Pro study pack published on celpip.ca describes exactly that sequence. Nobody asks a follow-up. Nobody rephrases the question when your face says you did not understand it. Nobody fills the gap when you stop.
That turns CELPIP Speaking into a different skill from the one most people have practised. It is not conversational fluency. It is sustained monologue against a clock, eight times inside a fifteen-minute component, and the specific way it fails is that you run out of things to say at thirty seconds with thirty seconds still on the counter.
That is not a language problem. It is a content-generation problem, and it has a fix.
This is written from the marking side — we build the engine that scores CELPIP speaking practice on Hilingo. Every format and scoring fact below is from celpip.ca's own current pages and published materials, linked where it appears, not from another blog's table.
The eight tasks, with the times Paragon actually publishes
Worth knowing before you read any table: Paragon's test format page lists the eight task names and a 15-minute allotment for the component, and stops there. It does not publish per-task timings, and neither does the test-taker guidebook — the word "seconds" does not appear in it. The per-task numbers come from the free Speaking Pro study pack published on celpip.ca, linked above — which is where a lot of competitor tables get them from without saying so, and where they also get Task 5 wrong.
| Task | Preparation | Speaking |
|---|---|---|
| 1 — Giving Advice | 30 seconds | 90 seconds |
| 2 — Talking about a Personal Experience | 30 seconds | 60 seconds |
| 3 — Describing a Scene | 30 seconds | 60 seconds |
| 4 — Making Predictions | 30 seconds | 60 seconds |
| 5 — Comparing and Persuading, Part 1 | 60 seconds | none — you do not speak |
| 5 — Comparing and Persuading, Part 2 | 60 seconds | 60 seconds |
| 6 — Dealing with a Difficult Situation | 60 seconds | 60 seconds |
| 7 — Expressing Opinions | 30 seconds | 90 seconds |
| 8 — Describing an Unusual Situation | 30 seconds | 60 seconds |
Three things fall out of that table that the copied versions lose.
Task 5 is two screens, not one. Part 1 gives you 60 seconds to read a situation, compare two options and pick one — and you say nothing at all. Part 2 then gives you another 60 seconds of preparation and 60 to persuade someone that your choice beats the one they are proposing. Tables that print "Task 5: 60 / 60" as a single row have quietly deleted a whole minute of the test. Paragon also notes that Tasks 5, 6 and 7 all offer a choice — which means part of your preparation time on three of the eight tasks is spent deciding, not planning.
Tasks 1 and 7 are the long ones. Ninety seconds each, against sixty everywhere else. If you have rehearsed a shape that comfortably fills a minute, it will leave you stranded on exactly two tasks — and giving advice and expressing opinions are the two where candidates are most likely to state a position in fifteen seconds and then stop.
Add the speaking column up: 540 seconds. Nine minutes. That is the entire audio record of your English that will exist when you finish. There is no other evidence. Ten seconds of silence is ten seconds removed from 540, and you cannot get it back on the next task, because the next task starts its own clock.
Speaking is also the last component of the test. On CELPIP-General you reach it after listening, reading and writing — roughly two and a half hours in.
IELTS gives you an interlocutor. CELPIP gives you a countdown.
This is the comparison that actually matters if you are choosing between the two, and it is usually framed as "CELPIP has no face-to-face interview, which is less stressful". Paragon frames it that way itself: its guidebook says the computer-delivered speaking component lets you demonstrate proficiency "without experiencing the anxiety or self-consciousness that often accompany the interview-style Speaking assessments employed by other testing systems."
That is true, and it is a genuine advantage for a lot of people. It is also only half the trade, and nobody sells you the other half.
In an IELTS interview an examiner is in front of you for the whole thing. Parts 1 and 3 are question and answer: if an answer dies after eight words, the next question arrives and you get another attempt at showing what you can do. Part 2 is a long turn, but it is one long turn, and the examiner is visibly there while you take it. The interaction is not scored as conversation, but it functions as a safety net — a stalled answer is refreshed by the next prompt, and a misunderstood question can be repeated.
CELPIP removes the net in both directions. There is no anxiety of being watched, and there is no rescue. You are handed eight consecutive long turns with a headset and microphone provided by the test centre (CELPIP FAQs), and each one stands or falls on whether you can keep producing relevant speech until the counter reaches zero.
If you are still choosing a test, PTE Core vs CELPIP works through the wider decision, including where each one is marked hardest.
Silence is not a gap in your answer. It is part of your answer.
Candidates treat the dead air at the end of a short response as neutral — as though the answer simply finished early. It is not neutral, and you do not have to take our word for it, because Paragon publishes the factors it assesses.
The Speaking Performance Standards have four dimensions, and each one lists the factors underneath it. Read the right-hand column of this against the factor names:
| Dimension | Factors Paragon names | What drying up does to it |
|---|---|---|
| Content / Coherence | Number of ideas; quality of ideas; organization of ideas; examples and supporting details | Direct hit. "Number of ideas" is a count, and you stopped generating them |
| Vocabulary | Word choice; precision and accuracy; range of words and phrases; suitable use | Indirect. A shorter sample simply contains less range to observe |
| Listenability | Rhythm, pronunciation and intonation; pauses, interjections and self-correction; grammar and sentence structure; variety of sentence structure | Direct hit. Pauses are a named factor, not an inference |
| Task Fulfillment | Relevance; completeness; tone; length | Direct hit. Length is a named factor |
Three of the four dimensions have a named factor that a short, gappy answer damages. That is the whole argument, and it is Paragon's own list.
It gets more specific than that. Paragon's Score Comparison Chart publishes sample responses with analysis at eleven levels from 0 to 12, and under Listenability its Level 9 entry says speakers "are also able to extend responses with fewer pauses" than at Level 8. That same extended-response-with-fewer-pauses formula reappears at level after level, each time measured against the one below — which is itself the point. Length and continuity are what the ladder is built on, at every rung. The chart does separate the levels on vocabulary too; it simply never stops describing how long you can keep going.
The study pack's response analyses point the same way. One Task Fulfillment note repeats across them, at Level 8 as much as at Level 12 — "Speaks for the full time" — and on a Level 12 Task 2 response the analysis adds "Gets cut-off at the end, but completes all task requirements before time is up."
Being cut off is not a failure. Stopping is.
CELPIP levels 4 through 10 map one-to-one onto the Canadian Language Benchmarks, so that 8-to-9 boundary is the CLB 8-to-9 boundary — the one the Comprehensive Ranking System pays most for. IRCC's table stops at CLB 10, which matters if you aim higher: a CELPIP 11 or 12 is still CLB 10 for Express Entry. The full mapping is in our CELPIP score chart, and how CELPIP is scored covers the level scale across all four skills.
The real failure: thirty seconds said, thirty seconds left
Here is the moment, and every CELPIP candidate recognises it.
Task 2 asks you to talk about a personal experience. You name the experience. You say roughly what happened. You say it was a good day. And then your brain returns nothing, the counter says 28, and you fill the rest with "so, yeah… that was, um, the experience" and a long breath.
You had the words. What ran out was things for the words to be about.
This is why "improve your fluency" is useless advice here. Fluency is a delivery property; what collapsed was supply. And supply is fixable in a way that fluency is not, because you can carry a method for generating more of it.
Four moves that always have an answer
Not a template. A template is the same words whatever appears on the screen, which by construction cannot be describing what is on your screen — and on a task scored partly on relevance, that is a real risk rather than a style preference.
These are four questions, not four sentences. Three of them cannot be answered without reading the prompt on your screen, which is what stops them hardening into a template. Carry them into the recording window and reach for the next one the instant you feel the tank empty.
- Who? Name the people, and give each one a detail. Not "my friend" but "my friend Ravi, who I have worked with for four years". This is the cheapest expansion there is — naming costs nothing and produces relative clauses for free.
- What happened? The sequence, in order, in the past. First this, then that, and then. Narration is where complex tenses appear without any effort to produce them: I had already left when she called.
- Why did it matter? Evaluation. What changed because of it, who it affected, what would have happened otherwise. This is the move that turns a description into a point, and it is the one most candidates skip entirely.
- What would you do differently? Hypotheticals. If it happened again, I would… Conditionals are the single cheapest way to produce complex grammar on demand, and this question is always answerable — there is no situation you cannot imagine handling differently.
Fifteen to twenty seconds a move fills a sixty-second task in three. On the ninety-second tasks use all four, or take two and develop each to forty-five seconds. Either way the list is not something to recite — it is there so that you always have a next question to reach for.
How the four moves land on each task
They are not only for the personal-experience task. Every CELPIP prompt is one of these four questions pointed in a particular direction.
- Task 1, Giving Advice (90s). Who is asking and what is their situation; what has happened to bring them here; why the decision matters; then the advice itself is move four — what I would do differently — and you can double its length by naming the option you rejected and why.
- Task 2, Personal Experience (60s). The natural home of all four, in order.
- Task 3, Describing a Scene (60s). Who is in the picture, with a detail each; what has just happened; why it matters to them; and what you would do if you were standing there.
- Task 4, Making Predictions (60s). Move two run forward instead of back — what happens next, then what happens after that, then why that outcome matters.
- Task 5, Comparing and Persuading (60s). Move four aimed at the other person's choice: what you would do differently from their proposal, and what goes wrong if they stick with it. Spend Part 1's silent minute choosing fast and reading the second option properly; a slow choice costs you the material, not just the time.
- Task 6, Dealing with a Difficult Situation (60s). The entire task is move four, so lead with moves one to three to earn the context first.
- Task 7, Expressing Opinions (90s). Move three carries this one. State the position in one sentence, then spend seventy-five seconds on why it matters and who it affects — not on restating the position in different words.
- Task 8, Describing an Unusual Situation (60s). Who and what, then why it is unusual, then what you would do about it.
One habit to drop while you are at it: do not announce that you have finished. "That's all I have to say" converts assessable speech into an admission that you stopped. If you truly have nothing left, go back to something you already said and add one concrete detail to it. That is what fluent speakers do, and it produces exactly the descriptive clauses the Listenability dimension is looking at.
Who actually marks it — and what Paragon does and does not say
This matters, because the internet is confidently wrong in both directions.
What Paragon states about Writing. Its FAQ answers the question "Is my test scored by artificial intelligence?" like this: the CELPIP Writing Test is scored by an AI-human hybrid system combining artificial intelligence with its human rating panel, and all CELPIP writing scores are reviewed and verified by its rating team (CELPIP FAQs). Anyone telling you CELPIP is "human-marked, unlike the machine-marked tests" is describing a test that no longer exists. Anyone telling you it is "just AI" is dropping the human verification, which is real and is worth something.
What Paragon states about Speaking. That same FAQ answer names the Writing Test and nothing else. There is no published AI-hybrid statement for Speaking, and we are not going to invent one. What Paragon does publish about Speaking is narrower and still useful:
- Its re-evaluation policy separates the components explicitly. Requesting a re-evaluation of Listening and Reading "is unlikely to result in a change in your scores as they are computer rated". Speaking sits on the side of the line where a second look can move the number. On the refund, Paragon's two pages do not agree — the FAQ says the fee comes back if the Speaking or Writing level increased, the test-results page says it comes back if the level changes for any component re-evaluated — so do not bank on a refund for a level that moves down.
- The same test-results page states that Writing and Speaking "are scored by qualified raters trained to apply consistent criteria… based on standard scoring rubrics", that trainees "must certify as CELPIP raters" before they rate, that raters get "ongoing training and regular monitoring", and that Paragon uses rater agreement statistics to check the quality of ratings.
- The Speaking Pro study pack tells test takers not to worry about background noise in the room because "the raters will be able to hear your response clearly."
So: Speaking is scored by trained, certified, monitored human raters, a re-evaluation of it can change your level, and what Paragon has not published for Speaking is whether any AI component assists, the way it has stated for Writing. That is the honest end of what is verifiable, and it is where we stop.
What does not depend on the split at all is the part that should change how you practise. Any assessment of a recording — human, machine, or a hybrid of both — can only use the recording. That is not a claim about Paragon's internals; it is what assessing an audio file means. A thought you had and did not say is not in the file. A reason you did not finish is not in the file. Thirty seconds of quiet breathing is in the file, and it is thirty seconds long.
What to practise, in order
- Time yourself honestly, once. Record one answer to each of the eight tasks at the real prep and speaking times above, then listen back with a stopwatch and write down where your speech actually ended on each one. Most people are surprised by which tasks collapse; it is usually 1 and 7, the ninety-second ones.
- Practise the two 90-second tasks separately and more often. They are 180 of your 540 speaking seconds — a third of everything you will actually say, sitting in two tasks that most preparation treats as ordinary.
- Drill the four moves against prompts you have not seen. The method only works if you can apply it cold. Practising it on a prompt you have already answered proves nothing.
- Check the answer against the four dimensions, not against a word count. "Was that 60 seconds" is the wrong question. "How many distinct ideas did I actually produce, and how many gaps were there" is the right one.
Hear what nine minutes of you actually sounds like
The thing you cannot judge from inside a 60-second answer is whether it held together, because you know what you meant. You will listen to your own recording and hear the answer you intended.
Hilingo runs full CELPIP-General and CELPIP-General LS mocks, and your first full scored mock is free with no card. Every speaking task comes back with a per-task estimate on each of the four CELPIP dimensions — content and coherence, vocabulary, listenability, task fulfilment — rather than one level for the whole section, which is the difference between knowing you got a 7 and knowing whether you got it because you ran out of ideas or because the ideas arrived in fragments.
You can also practise a single task type on its own rather than sitting the whole component every time, which is how you get enough repetitions on Tasks 1 and 7 to matter. The study centre lets you set the CELPIP level you need and shows the habits that move each section, and if English feedback about English is the wrong tool, the AI teacher works from your own result in your own language.
- How CELPIP is scored — the M-to-12 level scale across all four skills
- CELPIP Listening format — the other section where the format hides the real difficulty
- PTE Core vs CELPIP — if you have not finalised which test to sit
- CELPIP score chart — CELPIP level to CLB, IELTS and PTE Core
- Free CELPIP practice test — full mocks for General and General LS
Frequently asked questions
How many tasks are in the CELPIP Speaking test?
Eight, and the component is allotted 15 minutes. In order they are giving advice, talking about a personal experience, describing a scene, making predictions, comparing and persuading, dealing with a difficult situation, expressing opinions, and describing an unusual situation. Task 5 runs across two screens — Part 1 is a silent 60 seconds in which you choose between two options, and Part 2 is where you speak.
How long do you get to prepare and speak in CELPIP Speaking?
Between 30 and 60 seconds of preparation, and 60 or 90 seconds to speak, depending on the task. Tasks 1 and 7 give you 90 seconds of speaking; the rest give 60. Tasks 5 and 6 give 60 seconds of preparation instead of 30. Paragon's test format page publishes only the 15-minute total, so the per-task figures come from the free Speaking Pro study pack on celpip.ca.
Is CELPIP Speaking scored by AI or by a human?
Paragon's published answer to the artificial-intelligence question covers the Writing Test only, which it describes as an AI-human hybrid with every writing score reviewed and verified by its rating team. It has not published an equivalent statement for Speaking. What it does publish is that Listening and Reading are computer rated, that Writing and Speaking are "scored by qualified raters" who must certify and are monitored on an ongoing basis, and that a re-evaluation of Speaking or Writing can change your level. (Its FAQ and its test-results page describe the re-evaluation refund differently, so do not count on getting the fee back.) Anyone giving you a firmer answer than that is guessing.
What happens if I stop talking before the time is up in CELPIP Speaking?
The recording keeps running and captures the silence. Paragon names length as a factor under Task Fulfillment and pauses as a factor under Listenability, and its own Listenability descriptors describe each level as extending responses with fewer pauses than the level below it. Stopping early is not neutral — it removes evidence from three of the four dimensions at once. Being cut off mid-sentence at the end, by contrast, costs nothing as long as you covered the task.
Can you re-record your answer in CELPIP Speaking?
Paragon's description of the test is that when preparation time ends, the screen moves forward automatically, you hear "Start speaking now", and recording begins on its own. It publishes no way to restart or re-record a response on test day. The only published route to changing a speaking level after the fact is a paid re-evaluation, which you can request within six months of the test date.
What is the difference between CELPIP Speaking level 8 and level 9?
Under Listenability, Paragon's Score Comparison Chart says Level 9 speakers extend responses with fewer pauses than at Level 8 while maintaining understandable rhythm, pronunciation and intonation. The chart separates the two levels on vocabulary, content and task fulfilment as well, so it is not only a length difference — but length and continuity are the part most candidates stuck at 8 can change fastest, by generating more to say rather than learning more words.