Module 6: The Future
Lesson 2 of 6~8 min read

What Is Actually Coming

Emerging AI capabilities in plain English

Listen to this lesson

0:00
-:--

In July 2025, an AI model achieved gold-medal standard at the International Mathematical Olympiad — solving five of six problems that only 67 of 630 of the world’s best young mathematicians matched. You do not need to care about mathematics competitions. You need to know what that result says about where this technology is going.

Module 1 explained how language models work: pattern learners that predict text, fluent but fallible. That picture remains true, but the technology has not stood still, and some of what is emerging matters directly for medicine.

This lesson covers four developments in plain English: reasoning, multimodality, agents, and the diagnostic benchmarks that make headlines. For each, the same honest question — what is real, and what does it mean for you?

Reasoning: from quick answers to worked answers

The earliest chatbots answered in one pass — pattern-matched fluency, which is why they were confidently wrong so often. The significant shift since has been reasoning models: systems that work through problems step by step before answering, checking their own intermediate steps, sometimes for minutes at a time.

The mathematics olympiad result is the cleanest evidence that this is not marketing. Competition mathematics cannot be answered by fluent recall; it requires sustained, novel, multi-step reasoning. Gold-medal standard, under official contest conditions, was reached in 2025.

What it means for medicine: the class of error where the AI simply failed to think a problem through is shrinking. Complex, multi-constraint questions — the polypharmacy review, the conflicting-guidelines question — are increasingly within genuine reach.

Reasoning does not cure fabrication. A model can reason impeccably from a false premise it hallucinated two steps earlier — and a well-reasoned wrong answer is more persuasive than a sloppy one. Better thinking makes verification more important, not less, because the errors that remain are harder to spot.

Multimodal: AI that sees and hears

The second shift: models are no longer text-only. Current systems process images, audio, and documents together — which is why ambient scribes work, and why you can now photograph a drug chart or an ECG and ask a model about it.

The consumer versions of this will reach your patients before the clinical versions reach you. Patients already arrive having shown ChatGPT their rash. The deployment-curve question — a validated, regulated diagnostic use of images in UK primary care — is a different matter: dermoscopy tools and diabetic retinopathy screening are furthest along, each moving through proper medical device regulation at deployment-curve speed.

The practical near-term impact is humbler and already here: tools that listen (scribes), tools that read documents (Module 5’s paper tide), and patients with camera phones and questions.

Agents: AI that does, not just says

The third development is the one attracting the most investment: agentic AI — systems that do not just answer but act. Given a goal, an agent breaks it into steps, uses other software, checks its results, and continues until done: book the appointment, chase the result, complete the referral form, reconcile the spreadsheet.

In general business software this is arriving now. In the NHS, be precise: as of mid-2026 there is no national programme deploying agentic AI in general practice — what exists is commentary, vendor promises, and early pilots elsewhere in the system. Direction of travel, not current policy.

But the direction matters, because so much practice workload is exactly the multi-step administrative sequence agents are designed for. When this arrives, it will arrive through your existing suppliers — your clinical system, your document tools — and it will raise a governance question you are already equipped for: the more steps a system takes without a human looking, the stronger your checking framework needs to be. Module 5’s principles do not expire; they scale.

When a supplier says “agentic”, ask one question: what actions can it take without a human approving each one? The answer tells you the risk level, whatever the brochure says.

The diagnosis benchmarks, honestly read

You will have seen the headlines. The most striking: in 2025, Microsoft published a system that solved 85% of a set of diagnostically complex New England Journal of Medicine cases, against 20% for experienced physicians on the same task.

Real result — and here is the honest reading. The cases were rare, atypical, and selected for difficulty; the comparison physicians were denied textbooks, colleagues, and their usual tools; the benchmark rewards eventual diagnosis of the exotic, which is not what general practice mostly is; and the system is a research demonstration, not a regulated device. Capability curve, not deployment curve.

What it genuinely signals: the raw diagnostic reasoning capability of these systems now exceeds what most people — including most doctors — assume. The gap between that capability and your consulting room is regulation, validation, integration, and accountability. That gap is measured in years. It is not measured in forever.

Hold both truths. Today, these systems are not safe or legal replacements for clinical judgement, and nothing in this course changes on that. Over years, diagnostic support of real power will move through regulation towards practice — and the doctors who understand it will shape how it is used. That is the position this course has been preparing you for.

What is genuinely unknown

Balance requires saying what nobody knows. Whether the capability curve keeps climbing at this rate or flattens. Whether hallucination can be engineered away or is intrinsic to the approach. Whether health systems can absorb these tools faster than they absorbed previous technology — the NHS took two decades over paperless. And whether the economics of AI companies survive contact with reality at current prices.

Any of those uncertainties could reshape the timeline. None of them changes what you should do, which is the subject of the remaining lessons: read the NHS’s actual plans, understand what cannot change about your role, prepare adaptably, and stay informed at sustainable cost.

Key Takeaway

Four real developments: reasoning models that work through problems (shrinking one class of error, making the remaining errors more persuasive), multimodal AI that sees and hears (reaching patients before clinicians), agents that act rather than answer (direction of travel, no NHS general practice deployment yet), and diagnostic benchmarks whose striking results are capability-curve demonstrations, not deployment-curve reality. Powerful, genuinely improving, and still years from your consulting room in regulated form.