Skip to content
Exam planning

PTE Core vs CELPIP: both use AI, so what does a human actually check?

Aman Batth · 11 min read ·

Both are computer-delivered, both are accepted for Canadian PR, both convert to CLB — and both score with AI. The real difference is how much human judgement sits on top: Paragon states every CELPIP writing score is reviewed and verified by its rating team, while Pearson's human review reaches certain traits on certain PTE tasks.

If you are sitting an English test for Canadian permanent residence and you have narrowed it to PTE Core and CELPIP-General, you have already made the sensible cut. Both are built for Canada, both are fully computer-delivered, both are accepted by IRCC, and both convert to the Canadian Language Benchmarks.

Most comparisons then list fees, durations and result times. Those matter, but they are not the difference that changes your score. Neither is "people versus machines", which is the framing you will read everywhere and which is wrong: both tests score with AI. The real difference is narrower and more useful.

Both tests score with artificial intelligence. What differs is how much human judgement sits on top of it, and on which parts of your answer.

Paragon Testing Enterprises publishes that the CELPIP Writing Test combines artificial intelligence with its human rating panel, and that every CELPIP writing score is reviewed and verified by that panel. Pearson publishes something narrower for PTE: a human expert reviews particular traits on particular tasks alongside the AI, while the rest — pronunciation and oral fluency among it — is machine-scored only.

That difference is real, and it should change how you practise. It is not, however, what decides most Canadian PR profiles. The per-skill CLB thresholds below are.

We build the AI engine that scores speaking and writing on Hilingo, so this is the part of the comparison we can speak to first-hand. For the legal requirements of any immigration program, the authority is IRCC, not us. If you have not yet ruled out IELTS, start with CELPIP vs IELTS vs PTE Core, which compares all three by which skill you are weakest at.

The short answer

  • Choose PTE Core if you speak fluently and continuously, type quickly, and want results in a day or two. Its automated scoring rewards uninterrupted delivery and content coverage.
  • Choose CELPIP if writing is your weakest skill, or if you want a target with no conversion arithmetic — CELPIP levels are CLB levels. Paragon states that every CELPIP writing score is checked by a human rater before it reaches you.
  • The trap: PTE Core's CLB 9 needs 88 out of 90 in writing but only 78 in reading. If writing is your weakest skill, PTE Core charges you more than CELPIP does for the same benchmark.

What CLB actually costs on each test

This is where most comparison tables mislead, so read this one carefully.

CELPIP maps one-to-one: CLB 9 means a 9 in each of the four skills. PTE Core does not have a single number per CLB level. It has four different numbers, and they are a long way apart.

CLBCELPIP (all skills)PTE Core ListeningPTE Core ReadingPTE Core WritingPTE Core Speaking
101089889089
9982788884
8871697976
7760606968
6650516059
5539425151

PTE Core figures are the minimum score for each benchmark, from Pearson's published PTE Core to CLB concordance. CELPIP levels map directly to CLB.

Look at the CLB 9 row. Reading asks 78. Writing asks 88 — ten points higher for the same benchmark, and only two points below a perfect score. CLB 10 writing is 90 out of 90, which means there is no margin at all.

Be careful with any comparison that gives PTE Core a single range per CLB level, such as "CLB 9 = 84–88". That collapses four different requirements into one number and hides the only one that usually matters. On PTE Core, writing decides most profiles.

The rule underneath all of this

IRCC does not average your four results. It takes your lowest skill as your overall CLB.

Three 9s and a 6 is a CLB 6 profile, not "nearly CLB 9". So when you compare these two tests, compare them on your weakest ability, not your strongest. A test that is generous where you are already strong is worth nothing.

How much of your score a human checks

This is the part worth understanding properly, because it determines how you should practise.

PTE Core: machine-scored, with human review on some traits

An automated scorer is consistent. It does not have a bad morning, it is not charmed by confidence, and it applies the same rules to your recording at 9am and at 5pm. That consistency is genuinely valuable — it removes rater variance from your result.

It is also not purely automated, whatever you read elsewhere. Pearson's current score guides state that the Content trait on a number of tasks is reviewed by a human expert alongside the AI, with a second human deciding where the two disagree. That review is targeted rather than universal: it covers particular traits on particular tasks, and pronunciation and oral fluency are machine-scored only.

The cost is that the machine-scored traits reward what a machine can measure reliably:

  • Continuity. Oral fluency is one of the traits Pearson scores by machine alone, and it is measured partly as uninterrupted delivery. Long pauses and restarts are counted as gaps rather than as thinking, and no human reviews that trait afterwards.
  • Content coverage. Integrated tasks check whether specific content from the prompt appears in your answer. Elegant paraphrase that drops key content scores worse than plainer language that keeps it.
  • Clear articulation over accent. What costs marks is mumbling, trailing off, or speaking away from the microphone, rather than having an accent.

Practical consequence: on PTE Core, keep talking. A merely good answer delivered without hesitation usually outscores a better answer delivered in fragments.

CELPIP: an AI-human hybrid, with every writing score verified

Paragon answers this one itself. Its FAQ describes the CELPIP Writing Test as scored by "an AI-human hybrid system that combines artificial intelligence with our human rating panel", and states that all CELPIP writing scores are reviewed and verified by its rating team against the CELPIP rating criteria (Paragon's CELPIP FAQ).

So CELPIP writing is not marked from scratch by a person, and any page telling you it is has not read the source. What Paragon commits to is nonetheless stronger than anything Pearson publishes: no CELPIP writing score reaches you without a human rater having verified it. On PTE, human review reaches particular traits on particular tasks. If you want a human in the loop on your writing, that is a real advantage, and it is the one to weigh.

On speaking we are not going to tell you either way. Paragon publishes its hybrid statement for the Writing Test; we have found no equivalent published statement about how CELPIP speaking is marked, so anything said about it — by us or anyone else — is guesswork. Listening and reading are machine-marked, as on any test.

A human verification pass takes time that a fully automated one does not, which is consistent with CELPIP results arriving in three to four business days rather than one or two.

Practical consequence: on CELPIP writing, write for the rating criteria, because a person is checking your score against them. Raters work from content and coherence, vocabulary, readability and task fulfilment, so covering every point the prompt asks for beats polishing the sentences of an answer that misses one.

So which is easier?

Neither, in general — and be sceptical of anyone who answers that question without asking about you first. The honest version is conditional:

  • Fluent, fast, continuous speaker who types well → PTE Core will usually read you accurately and quickly.
  • Hesitant speaker who pauses to think → this counts against PTE Core, where oral fluency is machine-scored with no human review. It is not by itself an argument for CELPIP speaking, because Paragon does not publish how that is marked. It is one known cost against one unknown.
  • Weak at writing → CELPIP, on two counts: PTE Core's 88/90 writing requirement at CLB 9 is the single steepest ask across either test, and every CELPIP writing score is verified by a human rater before it is issued.

The practical differences

These are real considerations, but they should break a tie rather than decide the choice.

CELPIP-GeneralPTE Core
DeveloperParagon Testing EnterprisesPearson
DeliveryFully computer-based, one sittingFully computer-based, one sitting
SpeakingRecorded, no interviewerRecorded, no interviewer
Writing scored byAI-human hybrid; Paragon states all scores are human-verifiedAutomated, with human review of some traits on some tasks
CLB mappingOne-to-one, no conversionPer-skill conversion, uneven
Typical results3–4 business days1–2 business days
English varietyCanadianInternational
Validity for immigration2 years2 years

On fees and test-centre availability, both change by country and over time, and several comparison pages quote figures that are already out of date. Check the current numbers on the official pages — CELPIP and PTE Core — rather than trusting a blog post, including this one.

How to choose, in order

  1. Identify your weakest of the four skills, measured under exam timing rather than guessed.
  2. Check what CLB 9 costs in that skill on each test, using the table above. If it is writing, CELPIP is very likely the better route.
  3. Account for how each test is scored. PTE Core scores oral fluency by machine with no human review; every CELPIP writing score is human-verified.
  4. Only then compare results speed, fee and centre availability.

Do not choose on "which is easier". They are easier or harder for different people, and the variable is the skill you are worst at.

Find your weakest skill before you book

Everything above turns on one number you probably do not have yet: which of your four abilities is lowest under real exam conditions. A test booking costs a few hundred dollars and weeks of waiting, and your result is capped by whichever skill you never measured.

Hilingo gives you one full scored mock free, with no card. You sit a complete test under the real timers, and our own engine scores every speaking and writing answer, reporting per task and per question — so you walk into the decision knowing which skill is holding your CLB down, and which of these two tests treats it more kindly.

Take a free scored mock

Free, no account needed:

Frequently asked questions

Is PTE Core easier than CELPIP for Canada PR?

Not in general. It depends on your weakest skill and how you perform under automated scoring. PTE Core suits fluent, continuous speakers and fast typists, and returns results faster. CELPIP suits anyone weak at writing, because PTE Core's CLB 9 writing requirement is 88 out of 90 against CELPIP's flat 9, and because Paragon states that every CELPIP writing score is reviewed and verified by a human rater.

Which test is scored by AI?

Both are, in part — the question as usually asked has a false premise. PTE Core's speaking and writing are machine-scored, with a human expert reviewing certain traits on certain tasks. Paragon Testing Enterprises states that the CELPIP Writing Test is scored by an AI-human hybrid system combining artificial intelligence with its human rating panel, and that all CELPIP writing scores are reviewed and verified by that rating team. So the useful question is not which test uses AI, but how much human checking sits on top of it — and for writing, CELPIP publishes the stronger guarantee.

What is CELPIP 9 in PTE Core?

CELPIP 9 is CLB 9, which on PTE Core is Listening 82, Reading 78, Writing 88 and Speaking 84. There is no single PTE Core number that equals CELPIP 9 — the requirement differs by skill, and writing is much the hardest.

Does IRCC accept PTE Core for Express Entry?

Yes, IRCC accepts PTE Core. It does not accept PTE Academic for economic immigration — that version is for study applications. The names are similar enough that people book the wrong test, so check before you pay.

Can I switch from CELPIP to PTE Core?

Yes. You submit whichever valid result you choose, and results are valid for two years from the test date. You cannot combine skills across two different tests — a single test must supply all four abilities.

Which test gives results faster?

PTE Core, typically one to two business days against CELPIP's three to four. Paragon states that every CELPIP writing score is reviewed and verified by its rating team, and a human verification step takes time that a fully automated one does not. If your application deadline is tight, see how long English test results take.

Which should I choose if I am targeting CLB 9?

Compare the two at CLB 9 in your weakest skill. If that skill is writing, CELPIP asks for a 9 while PTE Core asks for 88 out of 90 — a meaningfully steeper climb. If your weakest skill is reading, PTE Core is comparatively generous at 78. Measure the skill first, then choose.

Practise with the engine this article describes.
One full scored mock, free. No card.
Take a free mock

All articles