CELPIP Learn · Speaking Task 3

Speaking Task 3: Describing a Scene

Give the listener a map before the details. Name the setting, follow one visual path, and connect people to their actions so the scene sounds organized rather than crowded.

Kate Feng
By Kate Feng, Language Education Specialist Published August 27, 2026 · 10 minute read
Preparation30 secScan the setting and choose a visual path
Speaking60 secDescribe selected people, objects, and actions
SourceIllustrationReport what is visibly happening
PriorityOrganizationBuild one coherent picture
Task snapshot

Thirty seconds to scan, sixty seconds to build the picture

Task 3 gives you an illustration and asks you to describe what is happening. You have 30 seconds to prepare and 60 seconds to speak. The current official study pack recommends starting with a general statement, choosing useful details rather than describing everything, and describing people's appearance, actions, and feelings.

The illustration may contain many simultaneous actions. That visual density is not an instruction to mention every object. Your job is to make the scene understandable to a listener who cannot see it. A clear setting plus several connected details is more useful than a fast inventory of unrelated nouns.

Describe relationships, not just objects. “A volunteer is handing a rake to a child near the shed” carries location, people, action, and relationship in one sentence. “Rake, child, shed, volunteer” does not build a picture.
What it measures

Selection, visual organization, descriptive range, and delivery

Content/Coherence includes enough relevant detail and a sequence the listener can follow. Vocabulary includes precise words for people, objects, actions, position, and apparent emotion. Listenability includes intelligible pronunciation, steady rhythm, grammatical control, and sentences that do not collapse under too many details. Task Fulfillment includes describing the actual scene for the full response in an appropriate neutral tone.

Accuracy matters. You may use careful language such as “appears to be,” “seems,” or “looks as though” when a feeling or relationship is inferred from visible clues. Avoid inventing an invisible backstory. Task 4 asks for predictions; Task 3 stays mainly with what is visible now.

Repeatable method

Prepare S → P → A → C

  1. Setting.Name the place, level of activity, and overall event: “a busy community garden during a weekend work session.”
  2. Path.Choose left to right, foreground to background, or one central group followed by nearby activity. Keep that path through the response.
  3. Actions.Select four to six clear action units. For each one, connect who, what, where, and any visible result.
  4. Connection.Finish by linking the details: people are preparing plots, sharing tools, waiting for help, or reacting to the same event.

Your notes might be only: S garden/work day; F tools + seedlings; M family watering; B shed/volunteer; C everyone preparing beds. Do not write full sentences during preparation.

Worked example

Move through one original scene in a visible order

Original demonstration scene: A neighbourhood garden is busy on Saturday morning. In the foreground, a woman kneels beside seedling trays while a man fills a watering can. In the centre, two children water a raised bed with an adult. To the right, a volunteer hands a rake to a new gardener beside a tool shed. In the background, two people carry a bag of soil toward an empty plot.

30-second notes: community garden / active; front seedlings + water; centre family; right shed/rake; back soil; shared work.

Model response:

This scene shows a busy community garden during a weekend work session. In the foreground, a woman is kneeling beside several trays of seedlings and appears to be deciding where to plant them. Next to her, a man is filling a large watering can from an outdoor tap.

In the middle of the garden, two children are watering a raised vegetable bed while an adult points toward a dry section of soil. On the right, a volunteer beside the tool shed is handing a rake to another gardener, who seems to have just arrived. Farther back, two people are carrying a heavy bag of soil toward an empty plot.

Overall, everyone is involved in preparing or maintaining the garden, and the scene feels cooperative and well organized.

Why it works: The response establishes the whole scene, follows foreground to background, uses precise action verbs, and ends with a supported connection rather than an invented story.

Common errors

What makes a visual description difficult to follow

  • Starting with one tiny object. The listener receives detail before knowing the place or event.
  • Jumping around the image. Left, background, centre, and left again creates an unstable mental map.
  • Listing nouns. Objects appear without people, actions, locations, or relationships.
  • Trying to describe everything. Speed increases while accuracy, grammar, and intelligibility fall.
  • Repeating “there is.” The same sentence frame hides vocabulary range and makes connected action sound flat.
  • Turning description into prediction. A long imagined future replaces what the image actually shows.
  • Claiming invisible facts. Exact jobs, names, motives, or relationships are stated without visual evidence.
Short drill

Record one path, then audit every location shift

Use the garden scene above. Prepare for 30 seconds and record for 60 seconds. This time, choose left to right instead of foreground to background. Include one general statement, four action units, and one closing connection.

Reveal the self-check

Listen without looking at the scene description. Mark each location phrase and sketch the route your words create. If the route doubles back without a reason, reorder the details. Circle every action verb; replace repeated “is” or “has” sentences with accurate verbs such as kneels, fills, waters, hands, carries, or points.

Self-review

Check whether your listener could reconstruct the scene

  • Did I name the setting and overall activity near the start?
  • Did I follow one clear visual path?
  • Did I select useful details instead of trying to mention everything?
  • Did I connect people with actions, objects, and locations?
  • Did I use precise descriptive verbs and spatial language?
  • Did I distinguish visible facts from careful inferences?
  • Did my sentence length allow steady, intelligible delivery?
  • Did I remain focused on the present scene?
Your next step

Practise a complete Task 3 scene

Open the Speaking practice catalog filtered to Describing a Scene. Choose one illustration, prepare an S → P → A → C map in 30 seconds, record once for 60 seconds, and review whether your words preserve a clear visual path.

Practise Speaking Task 3