On this page
Most test-takers walk into Task 4 feeling confident. They just described the image in Task 3. They know what's in the picture. So when Task 4 shows the same image again, their brain says: describe it again, but differently.
That instinct can lead you to answer the wrong task.
Task 4 rewards prediction, not description. The rubric wants logical reasoning about what happens next, not a recap of what you can already see. Describing the scene again without predicting what happens next leaves the Task 4 prompt unanswered.
Why Task 4 is harder than it looks
The trap is structural. You see a familiar image, your brain switches into description mode, and you spend 60 seconds explaining what the people are doing right now. That is the one thing Task 4 is not asking for.
The key difference between the two tasks:
- Task 3: Describe what you see happening in the image
- Task 4: Predict what will happen next, and explain why
The same four Speaking dimensions apply: content and coherence, vocabulary, listenability, and task fulfillment. Future forms and modal verbs can express your predictions, but there is no separate modal-verb score. Support what you predict with clues in the picture.
The fix is a different structure from the first sentence, not more vocabulary.
The three-part structure that works
A useful practice structure connects a likely event with a visible clue. Once you internalize it, you can apply it to any image.
Part 1: Brief scene reset (1-2 sentences)
Do not re-describe everything. One or two sentences to anchor your predictions to the image. Think of it as setting the stage, not retelling Task 3.
Example: "In this scene, a man is at the park with his bicycle near a bench, and a group of children are playing nearby."
Part 2: Two main predictions with reasons
Give each prediction a reason. Use a future form that matches how certain you are. Aim for two clear predictions. One is thin, three can feel rushed.
Strong prediction language:
- "He might decide to..."
- "It seems likely that..."
- "She could end up..."
- "This will probably lead to..."
- "Given that..., it is reasonable to expect..."
Pair each prediction with a "because" or "since" clause. That makes your reasoning easier for the listener to follow; it does not establish an official level.
Part 3: Natural wrap-up (1 sentence)
Close cleanly. You do not need a grand conclusion. One sentence that acknowledges the situation or adds a secondary outcome is enough.
Example: "Overall, this looks like it will turn into a busy afternoon for everyone in the park."
Sample scenario: use the same picture as Task 3
First describe this scene for Task 3. Then use it again below. Notice how the response changes from visible actions to possible next events.
Use the same bus-stop picture as Task 3. Predict what might happen next and explain the visible clues. Prepare for 30 seconds, then speak for 60 seconds.

Read an illustrative answer and review notes
The passengers will probably start boarding soon because they are lined up beside the bus and the woman at the front has her fare ready. The man with headphones may look down when he notices his shoes getting wet; he is focused on his phone while walking through a puddle. Meanwhile, the worker carrying cones will likely set them down somewhere nearby, perhaps to mark an area people should avoid. If the cyclist and skateboarder continue through the crowded pavement, they may have to slow down to make room for pedestrians. Overall, people will need to watch where they are going as the queue begins to move.
Each prediction has a visible basis. 'Perhaps' marks the cone placement as uncertain. Other outcomes are possible; explain a plausible next event rather than inventing a backstory.
This is an unscored teaching example, not an official CELPIP response or level prediction.
Additional text drill: weak vs strong park prediction
This is a text-only planning drill, separate from the pictured bus-stop exercise above. The responses are illustrative and unscored.
Imagine this scene: A man with a bicycle near a park bench. A group of children are kicking a soccer ball in the background. The sky is cloudy.
Weak response (sounds like Task 3): "I can see a man standing by his bicycle. There are kids playing soccer behind him. There is a bench and some trees. The weather looks cloudy and it might rain."
This response describes what is visible. The only prediction ("it might rain") has no reasoning tied to the people or action.
Strong response: "In this scene, a man has stopped his bike near a bench while children play soccer in the background. Based on his position and the fact that he seems to be watching them, he is likely to stay and rest for a while before continuing his ride. The children could move closer to him as their game expands across the field, since there is open space between them. If the clouds get darker, the whole group will probably leave the park sooner than planned."
Notice the structure: scene reset in one sentence, two predictions with reasoning, and a conditional third that fills time naturally.
Vocabulary for explaining uncertainty
Use familiar language accurately. Your prediction needs to be plausible in this picture; a sophisticated phrase cannot repair a missing visual clue.
Phrases that work:
| Prediction phrase | Reason connector |
|---|---|
| "This might happen..." | "...because / since / given that" |
| "It seems likely that..." | "...since / as / considering" |
| "She could end up..." | "...if / due to the fact that" |
| "This will probably lead to..." | "...because of / based on" |
| "It is reasonable to expect..." | "...given / since / as we can see" |
Avoid hedging too much. "It might possibly perhaps happen" weakens your prediction. One modal verb per prediction is enough.
The 30-second prep strategy
You have 30 seconds to prepare before speaking. Use it this way:
- Identify the two people or objects most likely to "do something next"
- Write one prediction for each, just a few words
- Write one reason for each prediction
That is all you need. You are not writing a script. You are picking your two anchors before speaking so you do not drift back into description mode.
Example notes: "Man + bike → will ride soon, sky darkening" and "kids → might leave early, clouds."
If you finish under 60 seconds
Task 4 responses should run close to 60 seconds. If you wrap up early, do not go silent. Add:
- A secondary prediction: "If the weather gets worse, the man might decide to take a different route home."
- A consequence: "This could mean the children finish their game earlier than they planned."
- A conditional outcome: "If more people arrive, the dynamic in the park will probably change."
These extensions are not filler. They show additional reasoning, which is what the rubric rewards.
Making the shift
If you catch yourself describing the picture again, practise switching to a prediction and explaining its visual clue. Reset quickly, predict specifically, and give a reason the listener can follow.
Practise your predictions in Speaking Exam 1, a free full mock. A free account includes one welcome AI feedback credit, so access to every task does not mean every response receives free AI feedback. Treat feedback estimates as practice guidance.
For context on how Task 4 connects to the full speaking section, read our guide on CELPIP Speaking Task 3: Describing a Scene first, since understanding the transition between the two tasks is half the work.
Put it into practice
Try a free full Speaking mock
Exam 1 is free in Speaking, Writing, Listening and Reading. Practice with the timers, then review your work. A free account includes one welcome AI feedback credit for Speaking or Writing; test access is separate.