SECTIONS
DOCTOR'S COACH ORIGINALS ●

AI in Dermatology: Which Tasks Could Change First?

Separate administrative assistance, image handling and triage—and ask what evidence makes each proposed workflow credible.

Healthcare technology workstation
Doctor's Coach editorial artwork / Doctor's Coach

A convincing image demo is not a clinic workflow

A model marks a patch of skin, produces a confident label and appears to compress a consultation into seconds. What the demonstration may leave out is who took the photograph, what information was missing and what happened when the output was wrong.

The useful question for dermatology is therefore smaller than “Can AI diagnose?” It is which defined tasks might benefit from assistance, what evidence supports that use and where a person remains responsible. The order below is an editorial assessment of candidates for evaluation, not a prediction about universal adoption or permission to use a product clinically.

Start with the work around the clinical decision

Drafting a non-confidential staff agenda or reorganising public teaching notes offers a bounded experiment. The input can be fictional or public, and someone can inspect the output before using it. These tasks do not demonstrate clinical accuracy, but they may reveal whether a tool is worth the time needed to check it.

Clinical documentation is a different proposition. It introduces questions about patient information, institutional approval, recording, storage and review. Do not move from a successful administrative experiment to uploading consultations without a separate assessment. Convenience does not settle whether the data should enter the system.

Image organisation and image interpretation are different jobs

A proposed tool might sort images, flag an incomplete capture or help compare records. Another might suggest a diagnosis or influence priority. Describe the exact output before assessing risk; “AI photography support” can hide very different responsibilities.

Even an organisational error can matter if it attaches material to the wrong record or hides an image a clinician expects to see. Ask how identity, ordering and missing information are checked. A system should have a clear way to stop and surface uncertainty rather than quietly fill gaps.

What the diversity evidence actually says

Daneshjou and colleagues examined dermatology models using the Diverse Dermatology Images dataset. In the directly accessible 2022 author manuscript, the tested models performed worse on darker skin tones and uncommon conditions. This is evidence about those evaluations, not proof that every current system has the same performance.

For an Indian service, the practical implication is to ask whether evaluation resembles the people, conditions and capture methods in that service. An overall accuracy figure can hide uneven performance. Ask for subgroup results, the reference standard and a description of the data excluded from testing before accepting a broad reliability claim.

Triage changes consequences before it changes diagnosis

If a system influences who receives earlier attention, its errors can alter access even when a clinician makes the final diagnosis. A proposal for triage should explain false reassurance and unnecessary escalation, not only successful classifications. It also needs a human route for cases the system cannot handle.

Consider a hypothetical clinic receiving photographs from different phones. A tool assessed on carefully captured images may not have been tested under those conditions. That mismatch is a question for the evaluation team, not something to dismiss because the product worked on a curated demonstration set.

Demand a bounded evaluation, not a leap of faith

A responsible proposal names the task, authorised data, intended users, review process and stopping conditions. It should establish whether the tool meets applicable institutional and regulatory requirements. This article is not a deployment protocol; involving appropriate clinical, technical and privacy expertise is part of deciding whether evaluation is justified.

Measure the whole workflow. How long does checking take? How often must an output be corrected? Does assistance change which errors people notice? A faster first draft can be a slower final result, particularly if the interface makes uncertainty difficult to inspect.

The reviewer needs room to disagree

WHO’s 2024 guidance on generative AI describes risks from inaccurate outputs and automation bias. These concerns are relevant when a generated explanation sounds more certain than its evidence. A polished paragraph does not become reliable because it accompanies a picture or uses specialist vocabulary.

Keep machine suggestions distinguishable from verified observations. Ask whether users can see what information the system relied on, report problems and continue safely when it is unavailable. The professional review step should be real work with time and responsibility assigned, not a disclaimer placed at the bottom of a screen.

Expect change task by task

The first useful change in a practice may be less dramatic than autonomous diagnosis: better organisation of permitted material or a more efficient reviewed draft. That is still worth evaluating if it solves a genuine problem without creating a larger one.

A dermatologist does not need to choose between enthusiasm and rejection. Ask a narrow question, require relevant evidence and keep the patient context larger than the model’s output. The technology earns a role through a checked workflow, not through the confidence of its presentation.

Sources & references

NEXT STORY

Gynaecology as a Career: Lifestyle, Earnings, Pressure & Reality