Skip to content
Duolingo English Test

Duolingo Speak About the Photo: silence is the most expensive thing you can say

Aman Batth · 9 min read · · Updated

You get 20 seconds to look and up to 90 seconds to talk, and the whole thing is scored by machine. That single fact changes what a good answer is — because a scorer with no audio has nothing to assess, and a brilliant observation delivered in fragments is worth less than a plain sentence that keeps going.

Speak About the Photo is one of the Duolingo English Test's spoken tasks. A photo appears, you get 20 seconds to look at it, and then you speak — for at least 30 seconds and up to 90 (DET Official Guide). One attempt, no retry.

Most advice about it is about what to say. Describe the foreground, then the background, mention the mood, use good adjectives. All of that is fine, and none of it is the thing that decides your score.

The thing that decides your score is that there is no examiner in the room. The DET is scored automatically (how the DET is scored), which means the only evidence of your English that exists is the audio file you produce in those 90 seconds.

Everything else follows from that.

A scorer cannot give you credit for something it did not hear

This is written from the marking side — we build the engine that scores speaking on Hilingo, including this task. Duolingo does not publish its algorithm and neither does anyone else, so treat the mechanism below as how automated speech scoring works in general, not as a leaked rubric.

Here is the useful split.

What an automated scorer hears wellWhat it cannot hear at all
How many words you produced, and how fastThat you noticed something subtle
What proportion of the window was speech rather than silenceThat you were searching for a word you know in your own language
How long each gap between words wasThat you were being careful
Whether the sounds you made match the words you meantThat you had a better second sentence planned
Whether the content words relate to the photoThat you were nervous

Read the right-hand column again. Every item in it is a thing candidates spend the 90 seconds doing — and not one of them produces a signal.

So the first rule of this task is uncomfortable but simple:

A pause is not neutral. It is a measurement of you, and it measures nothing.

What silence actually costs

In our engine, fluency on an open speaking answer is built from three measurable things: your words per second, the proportion of the window that contained speech, and the number and length of the gaps between words. Gaps past a short threshold count as pauses, and long ones count more heavily than short ones.

Two candidates describe the same photo.

Candidate A speaks for 80 seconds, plainly, without stopping: people, clothes, actions, place, weather, a guess about what is happening.

Candidate B says one genuinely perceptive sentence about the composition of the photograph, stops for four seconds to find the next word, says a second good sentence, stops again, and finishes at 40 seconds.

Candidate B produced a better sentence. Candidate A produced a better answer, on every measure that exists: more words, a higher speaking ratio, fewer long gaps, and more content words that relate to the image. The perceptiveness of B's observation is invisible — there is no field for it.

That is not a flaw you can strategise around. It is what scoring audio means.

The 30-second button is the floor, not the target

On the real screen the continue button stays locked until you have spoken for 30 seconds, and then it unlocks. A lot of candidates treat that unlock as a signal that they have done enough and press it.

It is not a signal. It is a minimum. The window is 90 seconds and the scorer sees the whole window. Finishing at 31 seconds means you handed over roughly a third of the evidence you were allowed to produce, and you did it voluntarily.

Treat the button as the point where quitting becomes possible, not the point where it becomes sensible.

Relevance is a gate, not a slider

Here is where this task differs from "just keep talking", and where a lot of prep advice quietly puts people in danger.

On open speaking tasks, relevance does not behave like the other measures. Fluency and pronunciation are sliders — more of them is more marks. Relevance is closer to a gate. In our engine, an answer that does not relate to the prompt at all does not merely lose content marks: it collapses. Pronunciation and fluency are floored along with content, because an answer about something else is not evidence of anything, however beautifully it was delivered.

That has one direct consequence for this task, and it is the opposite of what template sellers tell you.

A memorised paragraph is the same words whatever photo appears. If your words are identical for a photo of a market stall and a photo of a hospital corridor, they cannot be describing either one. We have written the full argument about why templates do not work for PTE, and the logic transfers exactly, because the mechanism is the same: content is scored against the thing in front of you.

The safe version is a fixed order, not fixed sentences. Which brings us to the real problem.

The actual problem: you run out of things to say at 45 seconds

Nobody's difficulty with this task is the first 20 seconds. It is second 45, when you have named the people, said what they are doing, and your brain returns nothing.

That moment is not a vocabulary problem. It is a prompt problem — you have stopped asking yourself questions. So carry questions, not sentences.

Six of them, in this order. Every photo can answer all six, which is the entire point:

  1. Who or what is in it? Count them. "There are three people" is a real sentence and it takes two seconds.
  2. What do they look like, and what are they holding? Clothes, colours, objects. This rung alone is usually four or five sentences.
  3. What are they doing? Present continuous, one verb per person or object.
  4. Where is this, and when? Indoors or outdoors, city or countryside, weather, light, time of day.
  5. Why might this be happening? A guess. "It looks like they are waiting for a bus, because…" Speculation is free and it is grammatically richer than description.
  6. What happens next, or what do you think of it? Your own reaction, in a full sentence.

Fifteen seconds a rung is 90 seconds. You will almost never need all six, and that is fine — the list exists so that you are never at zero.

Notice what is different about this and a template. A template gives you words that are the same every time and therefore describe nothing. This gives you questions whose answers are different every time, because four of the six can only be answered by looking at the photo in front of you. You never freeze wondering what comes next, and every sentence you say is about that image.

Use the 20 seconds to load nouns, not to write sentences

The preparation window is short enough that anything you compose in it, you will lose halfway through saying it.

So do not compose. Name five things in the picture, silently, as single words. Two people, a red umbrella, a wet pavement, a bus stop, an evening sky. Five nouns is five sentences you cannot run out of, and retrieving a noun under pressure is much more reliable than retrieving a clause you half-wrote.

If you finish naming early, add the verbs: waiting, checking, raining.

Three habits that feel like good speaking and are not

Self-correcting. You say a word, hear it come out wrong, and go back to fix it. The fix introduces a gap, and a gap is counted whether it was caused by confusion or by conscientiousness. A small grammatical slip inside continuous speech costs less than the silence you spend repairing it. Say the next sentence instead.

Narrating your own difficulty. "Um, I don't know what else I can say about this picture" is speech, so it fills the clock — but the content words in it are not about the photo, so it adds nothing to coverage while consuming the window you needed. If you are stuck, drop down a rung on the list above rather than commenting on being stuck.

Speaking quickly because you are nervous. Speed feels like fluency from the inside and is not the same thing. A scorer measures pace against a natural band; far above it is penalised, not rewarded, and rushing usually brings restarts with it. Steady beats fast.

How to find out what you actually sound like

Everything above is checkable, and you cannot check it by listening to your own recording — you will hear what you meant.

Hilingo runs full Duolingo English Test mocks, and your first scored mock is free with no card. Speak About the Photo comes back with three separate numbers — content coverage, pronunciation and fluency — rather than one, along with the count of long and short pauses in your own answer and a transcript of what you said with each word coloured by how clearly it came out. So "I need to be more fluent" becomes "I left four long gaps in 60 seconds", which is a thing you can practise.

The result page also carries question translation in 50 languages, if English feedback about English is the wrong tool, and an AI teacher that can see your result and explain a single answer back to you.

Take a free scored mock

Frequently asked questions

How long do you speak in Duolingo Speak About the Photo?

You get 20 seconds to look at the photo, then speak for a minimum of 30 seconds and a maximum of 90. The continue button unlocks at 30 seconds, but that is the floor rather than a target — the scorer assesses whatever audio you produced, so stopping early removes evidence you were allowed to give.

What happens if I stop talking in the middle of my answer?

Nothing good and nothing neutral. Automated speaking scores are built partly from how much of the window contained speech and how long the gaps were, so silence is not a blank space — it is a measured part of your answer, and it measures nothing. Keep going, even plainly.

Can I use a template for Speak About the Photo?

No, and it carries a real risk. Content on this task is judged against the specific photo, so wording that is identical whatever image appears cannot be covering the image. Open speaking answers that do not relate to the prompt are not just marked down slightly — they collapse. Learn a fixed order of questions instead of fixed sentences.

Is Duolingo Speak About the Photo scored by a human?

The Duolingo English Test is scored automatically. Your separate Speaking Sample is also shared with institutions, so an admissions officer can listen to that one, but the scoring itself is not a human judgement in the way IELTS or CELPIP speaking is.

What should I say when I run out of things to describe?

Move to a different question rather than a different adjective. Who is in it, what are they wearing or holding, what are they doing, where and when is it, why might it be happening, what happens next. The last two are speculation, they are always available, and they produce more complex grammar than description does.

Does my accent lower my Duolingo speaking score?

Accent is not the target; intelligibility is. What an automated scorer measures is whether the sounds you produced match the words you intended, word by word. That is why per-word feedback on a recording of your own voice is more useful than being told to work on your accent.

Practise with the engine this article describes.
One full scored mock, free. No card.
Take a free mock

All articles