On this page
The teaching examples on this page are original and not officially scored. Written responses illustrate content and structure; an official Speaking level also depends on the recorded delivery and the full performance.
Most CELPIP candidates look at the Task 3 picture and start naming things.
"I see a park. There is a bench. There is a woman. There is a dog. There are some trees."
That is technically a description. It repeats a sentence pattern without developing the actions or layout. That can make the description harder to follow; it does not justify assigning the answer a specific score.
The fix is a different way of looking at the picture, not harder vocabulary. This guide covers the exact Task 3 format, a 60-second structure, worked sample answers, and practice scenes you can drill this week.
Key takeaways
- Task 3 gives you 30 seconds to prepare and 60 seconds to describe a busy illustrated scene to someone who cannot see it.
- Choose a few useful details and organize them for the listener. There is no required number of people or activities to describe.
- Present continuous is the working tense: what people are doing, not what exists.
- The same picture returns in Task 4, where you predict what happens next. Describing actions in Task 3 sets that up.
- Drill it with CELPIP Speaking Task 3 practice questions under the real timer.
What CELPIP Speaking Task 3 asks you to do
The on-screen instruction is short: describe some things that are happening in the picture as well as you can, because the person you are speaking to cannot see it.
Read that carefully, because it contains three rules people miss:
- "Some things," not everything. The official pictures are deliberately overloaded, often showing a public place with several activities. Selection is part of the test.
- "Happening." The prompt asks for actions in progress. That is why present continuous carries the task: a woman is paying at the counter, two kids are reaching for the cereal shelf.
- The listener cannot see the picture. Your job is to build the image in their head, which is why spatial language matters as much as vocabulary.
The timing:
| Task 3 | |
|---|---|
| Prep | 30 seconds |
| Speaking | 60 seconds, recorded |
One more format fact worth knowing: Task 4 uses the exact same picture and asks what will most probably happen next. If you notice something about to happen while describing the scene, in the real exam you can save it for Task 4 instead of spending Task 3 seconds on it.
A 60-second structure that works
Split your response into three chunks:
- Set the scene (5 to 10 seconds). One sentence: where are we, and what is the general mood? "This looks like a busy Saturday morning at an outdoor farmers market."
- Narrate two or three focal points (35 to 40 seconds). Pick the most interesting people or actions. Describe what they are doing, not just that they exist. Move through the image in an order the listener can follow: left to right, or foreground to background.
- Wrap with an impression (10 seconds). One sentence about the overall atmosphere or how the people seem to feel.
This lines up with what the official CELPIP study materials themselves suggest for this task: start with a general statement, focus on some details rather than everything, use descriptive words, and describe people's appearance, actions, and feelings.
Memorize the skeleton. It keeps you moving forward instead of circling the same three objects.
Sample Task 3 question 1: the busy bus stop
Try describing the picture before opening the example. Start with the setting, then move through two or three areas in a clear order.
Describe what is happening at this bus stop to someone who cannot see the picture. Prepare for 30 seconds, then speak for 60 seconds.

Read an illustrative answer and review notes
This is a busy city bus stop with several activities happening around the queue. On the right, passengers are waiting beside a large bus, and the woman at the front is holding money. In the lower left, a food vendor is handing a hot dog to a customer. Nearby, a small child holding a red balloon is reaching toward an older woman. In the centre, a man wearing headphones is looking at his phone while walking through a puddle. A worker carrying traffic cones is crossing the foreground. Overall, the pavement looks crowded, with commuters, workers and children sharing the space.
Check the order: bus queue, food stall, child, then the foreground. The answer describes visible details. It does not claim to know the people's relationships, jobs or intentions beyond what the picture supports.
This is an unscored teaching example, not an official CELPIP response or level prediction.
For a thinner response, compare: “There is a bus. There are people. There is a man and there is a child.” It names objects but does not help the listener follow what is happening. The developed example adds actions and positions without claiming an official band.
Additional text drill: the farmers market
For this text-only planning drill, imagine this picture: an outdoor farmers market. A vendor behind a fruit stall is handing a paper bag to an older customer. On the left, a young man with a full basket is checking his phone while his dog pulls on its leash toward a cheese table. In the background, a musician is playing guitar and two children are dancing near an open guitar case.
Thin answer
"There is a market. I can see a man selling fruit. There is an old woman. There is another man with a dog. In the background there is a musician. There are two children."
The answer mostly names people and objects, with little detail about their actions or positions. That makes it less developed; these words alone do not establish an official score.
Stronger answer
"This looks like a lively outdoor farmers market on a weekend morning. In the foreground, a vendor is handing a bag of fruit to an older customer, and she is smiling as she takes it. Over on the left, a young man is scrolling through his phone while his dog is pulling him toward a cheese table, so it looks like he is about to be dragged away. In the background, a street musician is playing guitar, and two little kids are dancing in front of him. Overall, everyone seems relaxed and the atmosphere feels like a real community event."
Why it works
- Opens with a general statement that names the place and the mood.
- The central details show people doing something, using present continuous for actions in progress.
- Spatial anchors ("in the foreground," "over on the left," "in the background") build the layout for a listener who cannot see it.
- It covers four focal points and ignores the rest, which is exactly what the instruction invites.
- The dog pulling the man gives you a clue for a follow-up prediction. In the real exam, Task 4 reuses the Task 3 picture; for this text drill, use the same imagined scene.
Sample Task 3 question 2: the airport gate
For an additional text-based planning drill, imagine this picture: a crowded departure gate. A gate agent is scanning a boarding pass for a family at the front of the line. Behind them, a man in a suit is running toward the gate with a rolling suitcase. A teenager sitting near the window is asleep with headphones on, and a toddler next to her is dropping crackers on the floor while the mother is searching through a diaper bag.
Thin answer
"This is an airport. There are many people. A worker is checking tickets. There is a line. I can see a sleeping girl and a baby. The baby has some food."
Stronger answer
"This is a crowded boarding gate at an airport, and boarding has just started. At the front, the gate agent is scanning a boarding pass for a young family, while a long line is forming behind them. On the right, a man in a suit is sprinting toward the gate with his suitcase, so he may be worried about missing his flight. Near the window, a teenager is fast asleep with her headphones on, apparently not watching the boarding queue, and beside her a toddler is dropping crackers all over the floor while his mother is digging through her bag. It feels hectic, but in a very familiar, everyday way."
Why it works
- The opener adds one interpretive detail ("boarding has just started") that organizes everything after it.
- Contrast is doing work: the running man versus the sleeping teenager is more memorable than either alone, and raters hear that as coherence, not just vocabulary.
- Feelings are described from evidence: "may be worried," "appears to be asleep." That is one of the official strategies, and it does not require guessing anyone's biography.
Practice questions for CELPIP Speaking Task 3
You do not need official pictures to build the skill. Describe any busy scene under a 30-second prep and 60-second timer. Ten scenes worth drilling, chosen to match the settings the real test uses:
- A supermarket checkout at rush hour
- A public swimming pool on a hot day
- A classroom during a group activity while the teacher steps out
- A food court at lunch time
- A gym with several workouts happening at once
- A busy intersection with a cyclist, a delivery truck, and pedestrians
- A library study area during exam season
- A beach cleanup with volunteers and full garbage bags
- A hotel lobby during check-in, luggage everywhere
- A birthday party at the moment the cake arrives
For each one, force the structure: one scene-setting sentence, two or three narrated focal points with spatial anchors, one closing impression. Then do the same with real prompts and a real timer in Task 3 practice.
Two tools for clearer descriptions
Spatial language
Stop saying "there is." Start placing things in the scene:
- In the foreground / in the background
- On the left side / toward the right
- Next to / behind / in front of
- At the center of the image
These phrases give your description physical depth. The rater can picture the layout because you are building it for them.
Action verbs
Replace "there is a man" with what the man is doing:
- sitting: lounging, leaning back, scrolling through his phone
- walking: strolling, rushing, heading toward
- talking: chatting, laughing with, gesturing to
Specific verbs do more here than fancy vocabulary. "A couple is sharing a meal" beats "there are two people and food" every time.
Common Task 3 mistakes
Mistake 1: Listing objects instead of narrating actions
Present continuous is useful for an action in progress, but it is not a scoring requirement for every sentence. “The shelter has a curved roof” and “a customer is waiting beside the stall” can both add relevant detail. Check accuracy and usefulness rather than counting verb forms.
Mistake 2: Trying to describe everything
A busy picture can contain more detail than you can describe in one minute. Candidates who try to mention all of them produce a rushed list with no development. Three well-developed focal points beat eight name-checks.
Mistake 3: No spatial anchors
Without "on the left" and "in the background," the listener gets a pile of disconnected facts. The instruction says the listener cannot see the picture. Build the map first, then fill it.
Mistake 4: Freezing on the wrap-up
Plenty of answers describe well for 45 seconds and then trail off. Keep one closing move ready: name the atmosphere ("overall, it feels like...") or the shared mood of the people in the scene. It is ten seconds of reliable content that also matches the official strategy list.
Mistake 5: Wasting prep time writing sentences
Thirty seconds is enough for three quick choices: your opening word for the place, your two or three focal points, and the order you will take them in. Anything more and you will read instead of speak.
A 15-minute Task 3 practice routine
- Minutes 1 to 5: pick one scene from the list above. Give yourself 30 seconds of prep and record one 60-second attempt.
- Minutes 5 to 8: play it back and count your "there is" sentences. Decide which ones simply repeat information and which ones help establish the layout.
- Minutes 8 to 12: record the same scene again. Same focal points, but narrate them: who is doing what, where, and how they seem to feel.
- Minutes 12 to 15: listen once more and check the skeleton. Did you open with a general statement? Did you close with an impression? Were you still talking at 55 seconds?
Compare two recordings of the same scene to check whether your second attempt describes the actions and locations more clearly. For a structured version of this loop, how to self-review your recordings pairs well with this drill.
The skill behind Task 3
Describing a scene comes down to seeing the picture as a story rather than a list. Plan a route through the scene, then check whether that helps you connect the actions instead of searching for the next object to name.
And the payoff carries into the next task: the same picture returns in Task 4, making predictions, where the actions you noticed become the events you predict.
Drill this task with realistic scenes and the real timer in CELPIP Speaking Task 3 practice. If you are working through all eight tasks, the complete CELPIP speaking practice guide maps out where Task 3 fits and where to spend your prep time for the most improvement.
Put it into practice
Try a free full Speaking mock
Exam 1 is free in Speaking, Writing, Listening and Reading. Practice with the timers, then review your work. A free account includes one welcome AI feedback credit for Speaking or Writing; test access is separate.