How accurate is your PTE mock score? Audit the output, not the claim
Aman Batth · 15 min read · · Updated
Every platform says its scoring is accurate and none of them shows the working, which makes the claim impossible to check from outside. The result screen is a different matter. Six things to look for on any PTE mock, including ours.
Every PTE preparation platform tells you the same thing about its scoring, in slightly different words. It is accurate. It is aligned with Pearson's criteria. Your mock score will land within a few points of your real one.
Read enough of those pages and one thing stands out: none of them shows the working. Not the sample size, not how many candidates, not how long between the mock and the real test, not who checked. "Within ±5 points" with no method behind it is not a measurement. It is a sentence that costs nothing to write.
We build the engine that scores speaking and writing answers on Hilingo, so the conflict of interest here should be stated before anything else. We are not going to tell you our engine is more accurate than anyone else's. We have not run that comparison, and we will not publish a figure we cannot show you the derivation of. Our engine is not Pearson's either — nobody outside Pearson runs Pearson's, and a platform implying otherwise is describing something it cannot demonstrate.
So stop reading the claims. Audit the output instead. A scorer that cannot show its working cannot be told apart from one that is guessing, and a result screen gives that away in about ten minutes.
Why the accuracy claim cannot be checked from where you are standing
Suppose a platform says its mock scores land within five points of the real exam. To evaluate that you would need, at minimum: how many candidates, sat in what window, how many days between the mock and the official test, whether the same answers were used, whether candidates who improved in between were excluded, and who compiled it.
None of that is ever published. And even a platform acting in complete good faith runs into a harder problem: the people who sit a mock, get a disappointing number, and then study for three weeks before the real test are exactly the people whose two scores should differ. A gap is evidence of studying at least as often as it is evidence of a bad scorer.
That is the position you are in. You cannot audit the claim. You can audit the artefact.
What any automated scorer has to do, whoever built it
This next part is reasoning, not documentation, and it is worth marking as such. Pearson names its engines and describes the approach — the Intelligent Essay Assessor for writing, built on Latent Semantic Analysis, and Versant for speech — but the scoring model itself is proprietary and unpublished (Test Taker Score Guide, page 13). Nothing here is a claim about anyone else's internals either, including any competitor's. It is a statement about what any system marking recorded and typed answers must compute in order to return a number at all.
To put a mark on a spoken answer, a scorer must first turn the audio into words. To score Content, it must compare those words against something. To award Pearson's Form trait on an essay, it must count the words. To mark Highlight Correct Summary, it must already hold which option is correct, and it has the recording's transcript in front of it.
Every one of those intermediate results exists inside the system before the score comes out. So a result screen that shows you none of them is not hitting a technical limit. It is a decision about what to display.
If a platform shows you a number and nothing that produced it, either it never computed the parts, or it computed them and does not display them. Either way there is nothing on the screen for you to check.
That gives you six checks. They work on our product and on everyone else's, and the honest ones will survive them.
1. Does it show you the transcript it actually scored?
This is the one nobody tests, and on the speaking section it is the one that matters most.
Your Describe Image answer is not scored as audio. It is transcribed, and the Content judgement is made against the transcript. That is true of our pipeline and it has to be true of any automated scorer, because a machine cannot compare a waveform to the idea of a graph.
Which means a transcription error and an English error are indistinguishable in the final number. If the recogniser heard "the chart shows a decline" as "the chart shows the climb", your Content mark just measured a speech recognition failure and reported it as your English. You will never know, because the screen shows you 61 and the word "Content".
Ask for the transcript. If the platform shows it, you can read it in five seconds and tell the difference between "I said the wrong thing" and "it heard the wrong thing" — and if it is the second, the score for that item should be thrown away rather than studied. If the platform does not show it, you are being asked to accept a content score with no way to check what content was assessed.
This is also the check that quietly explains a lot of erratic mock scores for candidates with strong regional accents. Not a scoring bias, necessarily. A transcription problem, sitting one layer earlier, that nothing on the screen exposes.
2. Does it mark per criterion, or hand you one number?
PTE does not score your essay and give it a mark. It scores it on seven separate traits, each with its own scale, and your Read Aloud on three. Those traits are published in Pearson's Test Taker Score Guide, and we have set them out task by task with the scales alongside.
A single figure per task fails twice over. It is unactionable — a candidate losing marks on spelling and a candidate losing marks on structure need different weeks of work, and both can be handed "Essay: 64". And it is unfalsifiable in the same way the marketing claim is. There is nothing in "64" you can argue with.
So look for the trait names on the screen. Content. Form. Grammar. Vocabulary. Oral Fluency. Pronunciation. If none of the words on the result match the words in Pearson's guide, the result was not built from the guide.
3. On a wrong objective answer, does it quote the evidence?
Open one reading or listening question you got wrong. You should get three things: your answer, the correct answer, and an explanation of why the right answer is right. On the multiple-choice types, single or multiple, one sentence in the passage or the transcript usually settles it and you should see that sentence. On per-blank and ordering types there is no single settling sentence, so the equivalent is a note per blank or per pair.
A red cross and the right option tells you which box to have ticked. It does not tell you what you misread, which is the only part that transfers to the next test. And the platform already has the passage, the transcript and the answer key in front of it, so there is no technical obstacle to explaining the item — only a decision not to.
4. Does it deduct where PTE deducts?
Three PTE question types take a point away for each wrong selection, flooring the item at zero: Multiple Choice Multiple Answers in reading, the same type in listening, and Highlight Incorrect Words. Everywhere else a wrong answer simply earns nothing. Reorder Paragraph goes the other way and pays one point per correctly ordered adjacent pair, so a partial answer is always worth submitting. The full map of which task feeds which skill, and where the marks move, is in PTE marks distribution.
The check is blunt: over-select on purpose. Answer a multiple-answer item with only the options you believe correct, and note the mark. Then redo the same item with every option ticked. On a platform modelling the deduction, the second attempt must score strictly lower — with every wrong option selected it floors at zero. If it scores the same or higher, nothing is being deducted, and any score built on top of it is arithmetic about a different test.
A second version of the same check, for anyone comfortable with a calculator: count your raw correct answers in a section and ask whether that count could plausibly produce the score displayed. If a mock reports a percentage converted onto the 10–90 scale by a straight line, that conversion cannot represent what PTE does. PTE's overall is not an average of the four skills, many items feed two skills at once, and three item types can score negative before the floor.
5. Does it tell you what it does not know?
This is the check that separates a scorer built by people who understand the problem from one built by people selling a number, and it is the easiest to run: read the platform's own small print.
Three specific things to look for.
Does it say plainly that it is not Pearson's engine? Pearson's scoring is proprietary and only Pearson's official scored practice runs it. Any platform whose marketing implies its algorithm is Pearson's, or is calibrated to Pearson's algorithm, is claiming access it cannot demonstrate. We wrote out the full division of labour between official and third-party practice here, including the parts where official practice beats us.
Does the vocabulary match the current score report? If a result screen presents "Enabling Skills" as though they will appear on your official report, it is describing a report Pearson stopped issuing. Enabling Skills were removed from the PTE score report in November 2021 and replaced by the Skills Profile. What you receive today is one overall score and four communicative skill scores; the Skills Profile behind them is eight descriptive categories with no numeric scores — only a description, a performance indicator and recommendations. The traits still exist inside the marking of each task — they are simply never shown to you, which is the gap a preparation platform exists to fill. A platform still using pre-2021 vocabulary in 2026 is not reading the same documents you are.
Does it admit a limit anywhere? Not as modesty. As evidence of having looked. A scorer that has genuinely been tested against real answers knows where it is weak, and a product page with no limits on it has either not looked or has decided not to say.
6. Does the same answer score the same twice?
Submit the same answer twice and compare.
This one has a harder edge to it than the others, because we can tell you what we found in our own engine. Modern scorers use language models for the judgements that are not arithmetic — content relevance, grammar, argument quality — and those models are not reproducible even with the randomness turned all the way down. In our testing the same email came back with two, three and four grammar errors on consecutive runs at temperature zero. On one summary answer the content judgement came back 0.81 and then 0.72 on identical text, which was enough to cross a band boundary and move the reported Content mark.
That is not a flaw we discovered in somebody else's product. It is a property of the tools, and it applies to any platform using a language model in its marking, ours included. Whether a given competitor does is something only they can tell you — ask them.
What a platform does about that variation is the part you can actually check. We do not take a single judgement. Content and argument quality are scored as the median of three independent runs, and a grammar error is only reported if at least two of three runs find it, on the principle that an error the model finds only sometimes is by definition not one it is certain of, and a student's mark must not depend on which run they happened to get.
So run the check. Variation of a point or two on a repeated submission is the nature of the instrument. A criterion swinging by a band is a scorer with no vote behind it, and the platform will not have told you either way — which is why this is a check and not a claim.
What our result screen actually contains
Concrete, so you can hold it against the six checks above rather than take our word for it. This is what comes back on a PTE mock, not a statement about how close the number is to Pearson's.
| Check | What is on the screen |
|---|---|
| 1. Transcript | The transcript of your speaking answer, returned with the score, so you can see what was scored |
| 2. Per criterion | Writing and speaking marked against the exam's own criteria separately — content, fluency and pronunciation as their own figures, not folded into one |
| 3. Evidence | Objective answers carry a written explanation of why the answer is what it is, where one has been authored for that item |
| 4. Deductions | Objective marking follows the published scoring criteria, including the three types that deduct |
| 5. Limits | This page, and every other page we write on the subject |
| 6. Repeatability | Median of three runs on content and argument quality; two-of-three agreement before a grammar error is reported |
On PTE Read Aloud, and on our French read-aloud drill for TCF — a practice exercise we built, since TCF itself has no read-aloud task — there is a further layer: every word broken into its individual sounds in phonetic notation, with the sound you produced next to the sound expected, plus pause count and pace. That is the level at which pronunciation is fixable — "your pronunciation is 68" is not an instruction anyone can follow, and "the /dʒ/ in changing is coming out as /z/" is. We do not offer that on IELTS or CELPIP and will not pretend otherwise. The longer technical description of what the engine measures is on our AI scoring page.
Alongside that: single-question practice as well as full mocks, an AI teacher working from your own result rather than a generic syllabus, a dashboard that surfaces your weakest skill and a mistake bank built from your own wrong answers, and any practice question or answer review viewable in 50 languages — because a question you misread costs the same mark as a question you could not answer, and most people never find out which happened.
You will not find a success rate, an average score gain or an accuracy percentage on any page where we describe our own results. We could not show you how such a figure was produced, so we do not publish one. That is the whole differentiator, and it is the only one we can actually prove to you.
Run the checklist on us
One full scored mock, free, no card, complete report. That is enough to run checks 1 through 5 in a single sitting.
Check 6 needs two attempts at the same answer, so it needs a paid plan — on which mock and section-test retakes are unlimited for the length of your validity. The free tier is one mock, and we would rather say that than advertise "unlimited free mocks" and meet you at a paywall halfway through the second one. For reference when you are comparing libraries: PTE Academic is 45 full-length mocks and 45 section tests per section, PTE Core is 30 and 30, where a section test means one section of a full paper offered on its own rather than an extra paper.
If our result screen fails any of the six for you, tell us. Those six are the entire argument for sitting one.
- Free PTE mock test: what a result has to contain — the trait tables and deduction rules from Pearson's guide, in full
- Official PTE practice vs third-party — what each is for, and what ours cannot do
- PTE marks distribution — which task feeds which skill, and where marks are taken away
- PTE mock test vs practice questions — when a mock score means nothing, for reasons that are your fault rather than the scorer's
- How our AI scoring works — what the engine measures, in detail
Frequently asked questions
How accurate are PTE mock tests?
There is no honest single answer, and any platform quoting a percentage should be asked for the method — sample size, candidates, the gap between the mock and the real exam, and who compiled it. None publishes it. What you can check yourself is whether the result shows the transcript it scored, marks each criterion separately, explains the objective answers you got wrong, applies the deductions Pearson publishes, and returns a stable mark on a repeated submission. Those are answerable in one sitting; the percentage is not answerable at all.
Why is my PTE mock score higher than my real PTE score?
Several reasons sit behind that gap and they are not all the scorer's. Different engines produce different numbers, and no third-party engine is Pearson's. The real test also has scoring paths a practice platform does not — Content on several task types is reviewed by a human expert alongside the AI, according to Pearson's score guide. And mock conditions are usually gentler than exam conditions: a paused section, a replayed clip or a retaken task inflates a result without anyone intending it. Before blaming the engine, check whether you sat the paper the way you will sit the test.
Which PTE mock test is closest to the real exam?
Nobody can answer that from outside, including us, because it would require running the same candidates' same answers through every platform and the official engine in a controlled window, and publishing the method. No such study exists. The answerable version of the question is which platform's result screen shows you what it did — the transcript, the criteria, the evidence, the deductions and its own limits. That you can establish in ten minutes for free on any of them.
Can any platform use Pearson's real scoring engine?
Only Pearson's own official scored practice tests run the real engine. Every other platform, ours included, is running its own model against Pearson's published descriptors, which is a different thing. A platform stating that plainly is telling you something true; a platform implying its algorithm is Pearson's, or "calibrated to Pearson's algorithm", is describing access nobody outside Pearson has.
Is Hilingo's PTE scoring accurate?
We do not publish an accuracy figure, because we have not run a study we could show you and we will not print a number we cannot derive. Our engine is not Pearson's. What we will state is what the output contains: the transcript of your speaking answer, per-criterion marks on writing and speaking against the exam's own criteria, a written explanation on objective answers, and per-sound phonetic detail on PTE Read Aloud and our French read-aloud drill. Sit the free mock and run the six checks above on us — that is a stronger basis for judging us than any percentage we could print.
Do PTE mock tests give the same score twice for the same answer?
Not exactly, and that is worth understanding rather than being alarmed by. Any scorer using a language model for content, grammar or argument judgements has run-to-run variation even with randomness disabled — we see it in our own engine and handle it by taking the median of three runs and requiring two of three to agree before reporting a grammar error. Expect a point or two of movement on a resubmission from any platform. A whole band of movement on one criterion means nothing is smoothing the variation, and the score you were given was one sample of several possible ones.