A class in the AI Platform, “CNA – English for Healthcare AI Cases,” with two required AI patient cases listed

Practical training

AI in nursing education: which tools save faculty time

In this article

Where to start

The biggest time saver is also the riskiest

In nursing education, the AI tool that saves faculty the most hours is the one that drafts assessment material, and it’s also the one most likely to put a wrong answer in front of a class. Both halves of that sentence have numbers behind them, and a program that plans for only one of them ends up disappointed.

In a study built around a high-stakes emergency medicine exam in Hong Kong, a 100-question set written with ChatGPT-4o took 24.5 person-hours to produce, against 96 for the human-written set. Expert reviewers then found more factual errors in the AI questions (6% against 4%), more irrelevant items (6% against 0%) and far more items pitched at the wrong difficulty (14% against 1%).

That’s medicine, and the item-writing problem is the same in a nursing course. This piece stays at the faculty desk: which tasks to hand to AI tools, what to check in each one, and what goes wrong when nobody checks. What AI is changing in the profession as a whole has its own piece, on artificial intelligence in nursing.

Task by task

Six tasks on a clinical instructor’s desk

Most of what a clinical instructor does outside the classroom fits into six tasks. The middle column names the kind of tool, since brands change every semester.

Six faculty tasks, and what to review in each
TaskKind of toolWhat the educator checks
Writing NCLEX-style practice questionsGeneral-purpose chatbot or question generatorThe keyed answer, the distractors, and whether the item asks for application or only recall
Drafting a case and its variantsGenerative AI or a case authoring toolThat vitals, labs and medications agree with each other and with the diagnosis, and that the case serves the course objective
Conversation practice with a patientConversational simulator or virtual patientThe patient profile and the rubric: what the student should find out, and what earns credit
Feedback on each attemptRubric-based scoring built into the simulatorA weekly sample of sessions read against the transcript, with scores corrected where they miss
Updating old slides and handoutsAuthoring tool that converts existing documentsThat the guidelines cited are current, and that nothing appears that wasn’t in the source
Commenting on written care plansGeneral-purpose chatbotThe final grade, and your institution’s rules on student data before anything gets pasted in

Two jobs stay off the list on purpose. Deciding what counts as competent is a faculty judgment, and so is the debrief after a simulation, where the student hears why a decision mattered from someone who has made it on a real unit.

Drafting cases

A first version in minutes, reviewed element by element

Row two is where most of the hours go, because a usable case needs a scenario, a patient with a history, findings that fit together and a plan for scoring it. It’s also the row where a generated draft needs the closest reading.

The craft of writing the case itself, from the objectives to the patient’s script and the scoring, has its own piece: how to write a nursing case study for class. The policy side (what students may use AI for, and how assessment changes once they can) is in AI in medical education: a faculty’s playbook.

What goes wrong

Generated material that nobody read twice

The clearest numbers come from practice questions, because that’s where most of the studies are. When second-year medical students were given USMLE-style practice quizzes written with ChatGPT, independent experts found item-writing flaws in 49% of the questions and factual or conceptual errors in 22%. The same experts rated 59 of 65 (91%) a reasonable starting point for revision, which is the honest summary: a usable draft, with the revision still to do.

A blinded comparison in Medical Teacher found GPT-4 questions broadly comparable to those written by experts, and still more often unfit for use without major revision. Some keyed a wrong answer as correct, and the correct option tended to land in the same position.

Whole cases have less published data behind them, and the checks are the same ones: numbers that agree with each other, doses that match the drug, a history that doesn’t change between sections, and an ending that follows from what the student did.

AI can take on part of the review. In a study of 85 internal medicine questions, two of the models agreed closely with experts on objective flaws such as negatively worded stems and inconsistent numbers, and only weakly to moderately on whether a question matched its learning outcome. That last judgment is the one a course depends on, and it stays with faculty.

Per-activity review controls in the authoring tool: preview, required or optional, edit and delete

The first month

Starting with a single course

A workable start is small: one course, one row of the table, one reviewer with clinical expertise. The practice question bank makes a good first row, because its errors are easy to spot and cheap to fix.

Before anything reaches students, the reviewer answers each generated question without looking at the key. When their answer and the key disagree, the item goes back for revision. There’s time for that kind of review: in the Hong Kong study, generating the set took about a quarter of the human time, 24.5 person-hours against 96.

Keep a log of every correction the reviewer makes, with the type of error next to it. After a few weeks it shows which mistakes the tool repeats in your subject, and that list becomes the checklist the next reviewer starts from.

Back to blog

Frequently asked

Questions people ask about this.

  • How is AI used in nursing education?

    Mostly in three places: generating practice material such as questions and case drafts, simulating patients for conversation and clinical reasoning practice, and giving students feedback on each attempt against a rubric. In all of them, faculty still decide what gets assessed and review what the tools produce.

  • What AI tools can nurse educators use?

    General-purpose chatbots and question generators for drafting, case authoring tools that turn existing materials into cases or modules, and conversational simulators or virtual patients for practice. Each needs a different check, from the keyed answer of a question to the rubric of a simulation.

  • Can AI write NCLEX-style questions?

    It can draft them quickly, but the drafts need expert review. In one medical education study, 49% of ChatGPT-written practice questions had item-writing flaws and 22% had factual or conceptual errors, although experts judged 91% a reasonable starting point for revision.

  • Is it safe to use AI-generated clinical cases in class?

    Only after someone with clinical expertise has checked them. Generated cases can carry values that don’t agree with each other or with the diagnosis, and studies of generated questions show errors that human reviewers catch.

  • How much time can AI save nursing faculty?

    It depends on the task. In a study built around a high-stakes emergency medicine exam, a 100-question set generated with ChatGPT-4o took 24.5 person-hours against 96 for the human-written set, before the review the authors call essential.

Once a month

One piece of news from the field each month, and what we’re learning from it.

No sales pitch. Unsubscribe in one click.

For universities & faculties · Free guide

Free: the Education 4.0 Guide

Higher education in the post-ChatGPT era — the methodologies that work and how to use AI as the answer.

GET STARTED

See what the next generation of healthcare training looks like.

Whether you run a faculty, a residency program, a clinical training department, or a continuing education operation — a 20-minute demo, tailored to your context, is the fastest way to know if this fits.

We'll get back to you within one business day. No spam. Privacy Policy