SmoothMock
Live · turn-based · 11–14 min

IELTS Speaking, with a real examiner flow

Parts 1–3, timed cue cards, and an AI examiner that actually interrupts, follows up, and pushes back, not a script reader. Every attempt is graded against the same band descriptors a human examiner uses.

Start a free speaking mock
A test-taker mid-conversation with the AI speaking examiner

Test structure

Part 1 · Introduction

4–5 min of familiar questions about home, work, study, and interests.

Part 2 · Cue card

1 min prep, then speak for up to 2 min on a given topic.

Part 3 · Discussion

4–5 min of follow-up questions that go deeper on the Part 2 topic.

The examiner presses record, checks your passport against the name on the screen, and asks you to confirm your full name and where you're from. That's it. The test has started. There's no separate "warm-up" that doesn't count. From the first sentence you speak, you're being scored.

The whole thing is over in 11 to 14 minutes. Compare that to the reading and writing papers, which run an hour each, and listening at roughly 30 minutes with transfer time. Speaking is the shortest module of the exam by a wide margin, and also the one candidates walk away from least sure of how they did. That gap between duration and anxiety is the whole reason this page exists. Fourteen minutes isn't enough time to recover from a shaky start, and it's not long enough to hide weak vocabulary behind a good accent either. Every part is doing a specific job, and the examiner is filling in a specific rubric while you talk, not forming a vague overall impression.

This isn't a chat. It's a structured interview with three distinct parts, each testing something slightly different, each scored against the same four criteria. Knowing what's actually being measured, and when, changes how you prepare for it.

It also changes what "practice" should even mean. Most candidates practice by rehearsing answers to questions they've seen online, which trains exactly the wrong reflex: recall instead of production. The examiner isn't checking whether you know facts about your hometown. They're checking whether you can generate organised, accurate English in real time, under mild social pressure, with no chance to edit. That's a narrower and more specific skill than "being good at English," and it's trainable on its own terms.

The three parts, and what actually happens in each one

The test is delivered by a single examiner, either face-to-face in a room or over a high-definition video call at the test centre. A growing number of centres now offer the video-call format, and the content, timing, scoring and security checks are identical either way. It's always a live, real-time conversation with a certificated human examiner, never pre-recorded prompts. The whole interaction is audio-recorded for quality assurance and, occasionally, remarking.

Part 1: Introduction and interview (4–5 minutes)

The examiner verifies your identity, then asks general questions about familiar territory: your home, your work or studies, your daily routine, things you do in your free time. This part exists to settle you in and to get a first read on how you handle ordinary conversational English before the difficulty ramps up.

Don't mistake "familiar" for "unimportant." Examiners are listening for whether you can produce more than a bare answer, whether you naturally add a reason, an example, or a small extension without being asked. A one-word or one-clause response to every question caps how much range you can show, and range is exactly what's being scored.

Realistic Part 1 questions look less like "Tell me about your hometown" and more like this:

  • "Do you prefer buying things in shops or online? Why?"
  • "Is there a particular time of day you feel most productive?"
  • "How do people in your country usually celebrate a birthday?"
  • "What kind of weather do you like best, and why?"

None of these require specialist vocabulary. They reward a candidate who can talk for three or four sentences without stalling, not one who can recite a fact.

Part 2: The long turn (3–4 minutes, including prep)

This is the part everyone dreads, and it's the most mechanically specific. The examiner hands you a task card with a topic and three or four bullet points to cover, gives you a pencil and paper, and gives you exactly one minute to prepare. You may jot notes during that minute (full sentences aren't the point, quick prompts are), but the examiner collects your notes before you start speaking and doesn't look at them again; they're scaffolding for you, not evidence for the score. Once your minute is up, you're expected to talk for one to two minutes on the topic, and the examiner will stop you when the time is up, mid-sentence if necessary. After you finish, they'll ask one or two brief follow-up questions on the same topic before moving to Part 3.

A typical card:

Describe a decision you made that you were initially unsure about.
You should say: what the decision was, what made you unsure, how you eventually decided, and explain how you felt about it afterwards.

Other realistic prompts: "Describe a piece of technology you find difficult to live without," or "Describe a time you had to work with people you didn't know well." All of them follow the same shape: person, place, object, event, or experience, and all of them require you to organise a small narrative on the fly, not just list facts.

The skill being tested here isn't memory, it's structure under time pressure. Candidates who've only ever practiced answering questions, never practiced holding the floor alone for two full minutes, tend to run out of things to say around the 45-second mark and either repeat themselves or trail into silence. Both cost you on fluency and coherence, not on content. The examiner isn't grading your life story, they're grading whether you can organise and sustain one.

Part 3: The discussion (4–5 minutes)

Part 3 takes the Part 2 topic and pulls it into more abstract territory. The examiner asks you to compare, speculate, and evaluate rather than describe. If Part 2 was about a personal decision, Part 3 might move to how people in general make major life decisions, whether older and younger generations decide differently, or whether it's better to decide quickly or deliberate.

Sample follow-on questions in that vein:

  • "Do you think people today have too many choices, or does that vary by situation?"
  • "How do you think decision-making might change over the next generation, with more reliance on data and AI tools?"
  • "Should governments make certain decisions for citizens, like retirement savings, rather than leaving them to individual choice?"

This is where the examiner has the most latitude to follow up based on what you actually say. Part 1 and Part 2 run closer to a fixed script, but Part 3 genuinely responds to your previous answer. If you give a thin, one-line opinion, the natural examiner move is to push: "Why do you think that?" or "Can you give me an example?" That's not the examiner being difficult. It's the test doing exactly what it's designed to do: checking whether you can actually reason in English or whether you've only got a stock opinion with nothing behind it.

Across all three parts, a few rules hold constant: you can ask the examiner to repeat a question once, but they won't explain unfamiliar vocabulary or rephrase it to make it easier, and you can't ask to change the topic in Part 2 or negotiate the cue card. The examiner isn't there to help you succeed. They're there to observe how you perform under normal conversational constraints. Part 1 and Part 2 run closer to a fixed set of prompts the examiner works through in order; Part 3 is built around a bank of possible questions the examiner selects from and adapts based on what you've just said. That's a structural reason, not a personality quirk, why Part 3 can feel more like a real conversation and Part 1 can feel more like a checklist, because that's roughly what each one is.

The four criteria, and what actually separates the bands

Your Speaking band is the average of four criterion scores, each worth exactly 25%: Fluency and Coherence, Lexical Resource, Grammatical Range and Accuracy, and Pronunciation. There's no raw score conversion and no hidden fifth category. Examiners work from a published band descriptor scale, and because each criterion is weighted equally, one weak area drags the whole band down no matter how strong the other three are. That's the single most misunderstood fact about this test: you can't "vocabulary" your way past weak grammar, and you can't out-fluency a pronunciation problem that makes you hard to follow.

Here's roughly what each band sounds like in practice, based on the public band descriptors:

BandWhat it sounds like
5Gets the message across but leans on repetition, self-correction, and slower speech to keep going; vocabulary and grammar are noticeably limited on unfamiliar topics
6Willing to talk at length, uses a mix of simple and complex grammar, but coherence slips occasionally and complex structures often contain errors
7Speaks at length without much noticeable effort; uses some less common vocabulary and a range of complex structures, though not yet fully controlled
8Fluent with only occasional hesitation, wide and flexible vocabulary, a majority of error-free sentences, easy to understand throughout
9Fluent and precise across the board; hesitation, if any, is about content, not language; native-level control of grammar, vocabulary, and pronunciation

Fluency and Coherence

This isn't about talking fast. It's about talking without the seams showing: ideas connect to each other logically, and any hesitation comes from thinking about what to say next, not from searching for the English to say it. Band 6 candidates are "willing to speak at length" but lose coherence occasionally and overuse the same handful of connectives. Band 8 candidates hesitate rarely, and when they do, it reads as thinking, not struggling.

Weak (roughly band 5–6): "I think... um... I think that, you know, technology is... it's good but also bad, because, um, sometimes it's good and sometimes it's not good for people."

Stronger (roughly band 7–8): "Technology's a bit of a double-edged sword, honestly. It's made communication instant, but it's also cut down on how much face-to-face contact people actually get." Same basic opinion, but the second version moves in one continuous line instead of circling the same three words.

Lexical Resource

This measures range, precision, and your ability to paraphrase, not how many long words you know. A candidate who reuses "very good," "very nice," and "very interesting" for everything is capped regardless of how accurate their grammar is, because the descriptor for band 6 specifically flags vocabulary that's only "wide enough... in spite of inappropriacies," while band 8 requires vocabulary used "readily and flexibly to convey precise meaning."

Weak: "The food was very good and the restaurant was very nice and the staff were very friendly."

Stronger: "The food was excellent, really fresh, nothing overcooked, and the whole place had this relaxed, welcoming feel to it." The second version replaces one overworked intensifier with three different, more precise choices, which is exactly what "flexibility" means on the scale.

Grammatical Range and Accuracy

Two separate things are being judged here: how varied your sentence structures are, and how accurate they are once you attempt them. Candidates often over-index on one at the expense of the other: either playing it safe with only simple sentences (capping range) or reaching for complex structures that collapse under their own weight (tanking accuracy). Band 8 requires "a majority of error-free sentences," not flawless ones. Occasional slips are expected even at the top.

Weak: "I am living in this city since five years and I am liking it very much."

Stronger: "I've been living in this city for about five years now, and if I'm honest, it's grown on me a lot more than I expected." Same content, but the second version fixes the present-perfect-continuous error and adds a subordinate clause that actually works.

Pronunciation

This is the criterion candidates neglect most, partly because it's the hardest to self-assess, you can't hear your own flat intonation the way an examiner can. It's not about sounding British or American. The descriptor cares about whether you can be understood "effortlessly," and whether you use stress, rhythm, and intonation to shape meaning rather than speaking every word with the same flat weight.

Weak: a candidate reading each word with equal, even stress: "I. Went. To. The. Market. And. Bought. Some. Vegetables." No linking between words, no pitch movement to signal where the sentence is going.

Stronger: the same sentence spoken with natural linking and sentence stress: "I went to the market'n bought some vegetables." Stress lands on "market" and "vegetables," with a falling pitch at the end that signals the thought is complete. The words are identical. What changes the band is rhythm.

Where candidates actually lose marks

Reciting memorised answers. Examiners are trained to hear this within seconds: the rhythm flattens, the pace speeds up unnaturally, and the vocabulary gets oddly formal for a casual question. The real cost isn't that memorising is "cheating," it's that a scripted answer can't survive a genuine follow-up question, and when the examiner pushes and the script runs out, what's left exposed is your actual, unrehearsed fluency, which is usually worse than the polished part you led with. That mismatch between your rehearsed opening and your real answer to the follow-up is often more damaging than if you'd never rehearsed at all.

Giving minimal answers in Part 1. "Yes," "No," or a five-word answer to "Do you enjoy cooking?" doesn't give the examiner enough language to assess. It's not that short answers are penalised directly. It's that you've handed the examiner nothing to score you well on. Elaboration isn't padding; it's the only mechanism you have to demonstrate range.

Parroting the question back. Answering "What do you think about working from home?" with "I think working from home is..." wastes several seconds and shows zero paraphrasing ability, which is explicitly part of the Lexical Resource criterion. Strong candidates rephrase the prompt into their own words before answering it.

Overusing the same discourse markers. "Well, as I was saying... moreover... in my opinion... to be honest..." repeated mechanically across every answer is flagged directly in the band 5–6 descriptor as "over-use of certain connectives." It reads as a memorised toolkit rather than natural speech.

Chasing complexity you can't control. Candidates aiming to sound advanced sometimes build sentences with three subordinate clauses that collapse halfway through. Since Grammatical Range and Accuracy scores both range and accuracy, a broken complex sentence often scores worse than a clean simple one. You're better off attempting one solid complex structure per answer than forcing five shaky ones.

Going off-topic, especially in Part 2. Once you drift from the cue card's actual bullet points, coherence takes the hit, not fluency. The examiner can tell you're speaking fluently about the wrong thing.

Translating in your head before you speak. This is the quiet fluency killer for a huge number of candidates, and it's different from ordinary hesitation. The band descriptors explicitly separate hesitation that's "content-related" (pausing because you're deciding what to say) from hesitation that's about "searching for language," which is penalised more heavily the higher up the scale you go. A candidate composing a sentence in their first language and mentally converting it word by word produces exactly the second kind: long, even pauses in the middle of a clause rather than at natural thought boundaries. The fix isn't "think faster." It's building enough automatic phrases in English that you're not constructing every sentence from scratch.

How to actually prepare

Give yourself six to eight weeks if you're already close to your target band, and eight to twelve if you're trying to move a full band or more, say, from a 5.5 baseline to a 6.5 or 7. Anything shorter than that and you're mostly polishing what you already have, which is fine if that's all you need, but it won't fix a genuine grammar or fluency gap.

Record yourself. Don't just rehearse in your head. This is the single most-skipped piece of advice, and it's the one that matters most for pronunciation, because you cannot accurately judge your own intonation and stress by feel. You need to hear it played back. Record a Part 2 answer, listen to it the next day with fresh ears, and you'll catch flat delivery and filler words you didn't notice while speaking.

Drill Part 2 timing literally, with a timer and a pencil. Give yourself exactly 60 seconds to jot bullet-point notes (not full sentences, three or four words per point), then talk for a full two minutes without stopping. The habit you're building isn't "know the topic," it's "sustain two minutes of organised speech from a skeleton of notes," which is a completely different skill from writing out an answer in advance and reading it. If you're consistently running dry before 90 seconds, that's a structure problem, not a vocabulary problem. You need more sub-points, not more adjectives.

Build topic vocabulary in banks, not lists. Organise around the five Part 2 categories (person, place, object, event, experience) and around recurring Part 3 abstract themes: technology, education, the environment, government policy, generational change. For each theme, collect 10–15 useful collocations (not single words) that you can genuinely deploy on the fly, not memorised sentences you'd recite. A collocation bank survives a follow-up question. A memorised sentence doesn't.

Shadow native audio for ten minutes a day. Pick a short clip (a podcast segment, a BBC interview) and repeat it in real time, matching stress, rhythm, and where the speaker pauses, not just the words. This is the most direct way to improve pronunciation, because it trains the muscle memory for English rhythm rather than asking you to consciously think about stress mid-sentence, which never works under exam pressure.

Use a light structure for opinions, but don't let it calcify into a template. Opinion, one clear reason, one concrete example: that shape works for most Part 3 questions. The moment it becomes a rigid formula you plug every answer into regardless of the question, it starts sounding exactly like the memorised-answer problem examiners are trained to catch. Treat it as a starting scaffold you abandon once you're comfortable, not a script.

Track your own error patterns across attempts, not just your scores. A band number tells you almost nothing actionable. What you need is whether you're losing marks to the same three or four mistakes every time, tense slips, a narrow set of overused adjectives, flat intonation, because fixing a repeated pattern moves your score more than one more full-length practice test ever will.

Run at least a few attempts under genuinely realistic conditions before test day. That means no pausing mid-answer, no rereading the cue card twice, no muting yourself to think, the same conditions you'll actually face in the room or on the video call. Candidates who only ever practice in low-pressure, stop-start conditions tend to lose a noticeable chunk of fluency the moment the real clock starts, simply because they've never rehearsed recovering from a stumble in real time. Two or three full, uninterrupted mock attempts in the final two weeks matter more than another ten scattered, casual ones earlier on.

Where SmoothMock's AI examiner fits in

SmoothMock runs the same three-part structure (timed Part 1 questions, a genuine one-minute-prep Part 2 long turn, and a Part 3 discussion) as a live, turn-based conversation rather than a static record-and-submit prompt. It follows up on your answers and pushes back the way a human examiner does: if your Part 3 opinion is thin, it asks you to justify it; if you drift off the cue card in Part 2, it notices. Every attempt runs through the same mistake taxonomy examiners use, then gets translated into a coach narrative that tells you what to fix before your next attempt, not just what you got wrong.

Sources

At a glance

Graded against the 4 official criteria

Fluency and Coherence

How smoothly you speak and how logically ideas connect.

Lexical Resource

Range and precision of vocabulary, including idiomatic language.

Grammatical Range and Accuracy

Variety of sentence structures and how often errors occur.

Pronunciation

Clarity, stress, intonation, and how easy you are to understand.

Not a score. A reason for the score.

Every attempt runs through a mistake taxonomy, so you see the pattern, not just the mark.

um, I think, like, it's goodflag: overuse of fillers, Fluency and Coherence
I have went there last year→ past tense agreement, Grammatical Range
very good, very nice, very big→ vocabulary range too narrow, Lexical Resource

Sample speaking questions

Describe a skill you would like to learn. You should say what it is, why you want to learn it, and how you would learn it.

Do you think it's important to learn skills throughout your life, not just when you're young?

How has technology changed the way people learn new skills?

Questions, answered

How close is the AI examiner to a real speaking test?+

It runs the same three-part structure (introduction, cue card, discussion) as a live turn-based conversation, not a script you read into a mic. It follows up on your answers and pushes back the way a human examiner does.

Does the speaking test work on mobile?+

Yes. The speaking flow works on phone with mic access. Most test-takers practice speaking on mobile and switch to desktop for writing.

How is IELTS Speaking scored?+

Across four criteria, weighted equally: Fluency and Coherence, Lexical Resource, Grammatical Range and Accuracy, and Pronunciation. The examiner scores each out of 9, then averages them into your overall Speaking band, rounded to the nearest whole or half band.

What is a good IELTS Speaking score?+

Band 7 is generally considered strong, fluent English; band 8–9 signals near-native command. Most university and visa requirements sit between 6.0 and 7.5, so check your specific institution's threshold rather than aiming for a fixed number.

Can I ask the examiner to repeat a question?+

Yes, once, but they won't rephrase it or explain any vocabulary you didn't understand. If you miss it a second time, you're expected to answer based on what you did catch.

Take your free speaking mock test

One full attempt, fully graded, no card required.

Start your free test →