AI viva practice is best used as a high-volume drilling tool, not as your final judge. In the AI viva practice vs human mock viva debate, the practical answer is simple: use AI for repetition, timing and answer structure; use humans to judge whether you sound safe, flexible and patient-centred when the station pushes back. Current spoken assessments still reward integrated performance — communication, management, prioritisation, empathy and safety — not just factual recall.
Across current college examples, that is explicit. The MRCGP Simulated Consultation Assessment (SCA) uses 12 simulated consultations and marks domains including data gathering and diagnosis, clinical management and medical complexity, and relating to others; PACES23 consultation encounters combine focused history, examination, differential diagnosis, management and patient concerns; the MRCPsych Clinical Assessment of Skills and Competencies (CASC) tests consultation management, risk and communication; and the MRCOG Part 3 Clinical Assessment mixes simulated patient or colleague tasks with structured discussion across safety, communication, information gathering and applied knowledge.
Takeaway: don’t choose one and ignore the other. Choose the job each method does best.
Why this matters
What examiners listen for is usually more than content. They are listening for whether you open safely, select the important facts, answer the question asked, explain your plan in plain language, and stay composed when a patient, relative or examiner changes the temperature of the room. Official feedback language in current exams repeatedly points to missed cues, formulaic phrasing, weak time management and failure to reach a patient-centred management plan as ways performance falls apart.
If you only practise with AI, you may become crisp but slightly synthetic. If you only practise with humans, you may not get enough volume to iron out openings, structures and verbal habits.
The best candidates separate these tasks.
What AI viva practice does well
AI is strong when the task is repetitive, structured and time-limited. That fits a lot of modern spoken assessment design: standardised role-players in the SCA, defined station constructs in the CASC, structured discussion tasks in MRCOG Part 3, and tightly specified consultation encounters in PACES23.
Use AI for:
- first-draft answer structure
- high-volume repetition
- time-boxed drills
- spotting obvious omissions
- turning yesterday’s bad mock into five cleaner reps today
A useful way to prompt it is to make the task narrow. Ask for one stem, one timer, one interruption, and one feedback lens.
For example:
- GP-style drill: a 12-minute SCA-style consultation with a three-minute read-in, where you must get to management and safety-netting.
- Psychiatry-style drill: a CASC-style management station with 90 seconds to read, seven minutes to act, and an anxious relative who interrupts twice.
- Medicine-style drill: a PACES23 consultation where you must take a focused history, examine selectively, state a differential and explain next steps.
- O&G-style drill: a MRCOG Part 3 structured discussion with a two-minute read-in and a colleague handover halfway through.
The point is not realism. The point is reps. AI lets you do ten short openings, five differential drills, or three bad-news frameworks in the time it might take to organise one human mock.
AI viva practice vs human mock viva: where each helps
Where AI helps most
AI is best early in the cycle or immediately after feedback. If a human mock shows that you ramble, forget red flags, or never state a clear plan, AI can drill that exact weak point tonight.
It is also useful when you want uncomfortable repetition. Few colleagues want to hear you answer the same hyperkalaemia stem six times. AI does not mind.
Where AI starts to mislead
AI is a poor substitute for the human feel of the station. It may let vague empathy sound acceptable, miss that you interrupted too soon, or reward an answer that is tidy but emotionally flat. Current SCA feedback specifically warns about formulaic consulting, missing verbal and non-verbal cues, and poor timing; CASC criteria likewise emphasise professional relationship, cue recognition and control of pace.
So use AI feedback as a draft, not a verdict. If it says your answer was good, ask yourself a harder question: would a real examiner or role-player have trusted me?
What human mock vivas still do better
A good human mock tells you what the station felt like from the other side. That matters because live exams assess how you handle patients’ concerns, dignity, rapport, cueing and safety under stress, not just whether you can recite a list. PACES23 explicitly includes managing patients’ concerns and maintaining patient welfare, while MRCOG Part 3 uses lay examiners on some tasks to assess communication, patient safety and information gathering from a patient perspective.
Human partners are also better at productive unpredictability. They can go quiet, become annoyed, ask the awkward follow-up, or look unconvinced when your explanation is too slick. That is exactly where weak answers show themselves.
A few examples:
- A GP trainer can tell you that your safety-netting sounded generic rather than tailored.
- A psychiatry colleague can tell you that you asked about suicide risk competently but without warmth.
- A physician doing a PACES-style mock can stop you after two minutes and say you still have not told me what you think is going on.
- An O&G colleague can tell you that your handover was safe, but badly prioritised.
That sort of feedback is hard to fake.
Human limits
Human mocks are not perfect. They are harder to schedule, can be expensive, and vary wildly in quality. Some partners are too nice. Others over-mark based on their own pet phrases or niche practice.
That does not make human mocks overrated. It means you should use them selectively: fewer mocks, better briefs, sharper debriefs.
Build a blended practice plan
The safest approach is simple: let AI build fluency, and let humans test credibility.
A practical split
Four weeks or more out, make AI the daily tool and humans the weekly tool.
- 15 to 20 minutes of AI drilling on one narrow skill
- 1 short human mini-mock each week
- 1 recorded full mock every 1 to 2 weeks
In the final 10 to 14 days, flip that balance.
- keep AI for warm-ups and last-minute reps
- increase human mocks
- practise with people who will interrupt, challenge and debrief honestly
Protect exam content
Use invented, adapted or supervisor-written practice cases. Awarding bodies explicitly treat exam materials as confidential and prohibit reproducing or sharing content, so don’t paste recalled stations into any AI tool.
A good rule: practise the skill, not the stolen stem.
Common mistakes
- Using AI as your only source of feedback
- Asking for pass or fail judgements instead of specific critique
- Practising long monologues instead of short decisions
- Sounding scripted because you have memorised empathy lines
- Spending so long gathering data that you never reach explanation, management and safety-netting
- Choosing only friendly human mocks that never interrupt or challenge you
- Reproducing recalled exam content in notes, chats or prompts
Those failure patterns overlap closely with official feedback themes: formulaic communication, missed cues, and poor time management remain recurring problems in current spoken assessments.
Practice workflow
Pick one recurrent weakness per week. Not ten. One.
Then use a tight loop:
- Do one baseline human or recorded mock.
- Write down the exact failure point: weak opening, missed patient agenda, unsafe prioritisation, rambling explanation, poor close.
- Run 4 to 6 AI drills on that one issue over the next few days.
- Repeat the same case type with a human partner.
- Debrief with three headings only: keep, cut, add.
- Re-test 48 hours later.
If you do this consistently, the two methods stop competing. They start doing different jobs.
Summary
- AI viva practice is best for repetition, timing and answer structure.
- Human mock viva is best for realism, cueing, challenge and credibility.
- Use AI to fix a weak component; use humans to test whole-performance quality.
- Near the exam, shift towards human mocks.
- Never upload recalled exam content into AI tools.
References
- https://www.rcgp.org.uk/mrcgp-exams/simulated-consultation-assessment/introduction
- https://www.rcgp.org.uk/mrcgp-exams/simulated-consultation-assessment/marking-and-results
- https://www.rcgp.org.uk/mrcgp-exams/simulated-consultation-assessment/feedback-statements
- https://www.rcgp.org.uk/terms-conditions/sca-conditions
- https://www.rcgp.org.uk/mrcgp-exams/gp-curriculum/gp-curriculum-update-notice
- https://www.mrcpuk.org/sites/default/files/documents/PACES23%20Consultation%20scenario%20writing%20guidance.pdf
- https://www.rcpsych.ac.uk/training/exams/preparing-for-exams/preparing-for-the-casc
- https://www.rcpsych.ac.uk/news-and-features/latest-news/detail/2025/12/05/rcpsych-announces-additional-casc-exam-for-affected-candidates-in-february-2026
- https://www.rcog.org.uk/careers-and-training/exams/mrcog-our-specialty-training-exam/mrcog-part-3/mrcog-part-3-format/
- https://www.rcog.org.uk/careers-and-training/exams/exam-regulations-drcog-mrcog/
- https://www.rcog.org.uk/careers-and-training/exams/mrcog-our-specialty-training-exam/mrcog-part-3/global-expansion-and-sustainability-project/