top of page

The FDA Is Asking How to Govern Generative AI in Medicine. My Op-Ed on What AI Skin Analysis Should Admit About Itself

  • Writer: Dr. Lazuk
    Dr. Lazuk
  • 20 minutes ago
  • 6 min read

On August 18, the FDA opened a public comment period on how it should regulate generative-AI-enabled medical devices, open through October 19. It is easy to read a request for feedback as a non-event — no rule has changed, nothing is enforceable yet. I read it differently. A regulator asking the question in public, in a discussion paper rather than a proposed rule, is itself a signal: the agency knows the current framework was built for a different generation of software, and it is trying to get ahead of a category that has already outrun its own vocabulary. This is the right moment for anyone offering an AI skin-analysis tool, including my own platform, to be precise about what the tool actually does, before precision becomes mandatory rather than voluntary.

A discussion paper is not an endorsement, and it is not a threat

I want to be direct about what this is not. It does not validate any consumer skin-analysis app as a medical device. It does not convert an algorithm's output into a diagnosis. And it does not mean regulation is imminent in a form anyone can predict yet. Discussion papers of this kind typically precede a proposed rule by a year or more, and the eventual framework will likely look different from anything currently being speculated about, including anything I write here. What the paper does mean is that the agency is actively thinking about where generative AI sits relative to four functionally distinct roles: measurement, decision support, diagnosis, and autonomous treatment recommendation. Those four categories get treated as interchangeable far too often in consumer marketing, and that conflation is the actual problem this comment period is trying to name.

It is worth sitting with why the FDA is doing this now rather than five years ago, when consumer AI skin apps first appeared, or five years from now, once the category has fully matured. Generative AI specifically — as opposed to earlier discriminative or classification-only models — introduces a new failure mode: a system that can produce fluent, confident-sounding output even when it is extrapolating well beyond its training data. A classification model that says pore size is moderate fails safely when it is wrong; it just produces a slightly inaccurate number. A generative model that produces a paragraph of confident-sounding treatment reasoning can fail in a way that reads as authoritative even when it is fabricated. That is a different risk profile, and it is the one regulators are now trying to name in public before it becomes normalized.

Where I think the AI skin-analysis category has gotten sloppy

Consumer skin-analysis tools, including AI-driven ones, have proliferated faster than the language used to describe them has matured. A tool that measures pore size or estimates a wrinkle score is doing something categorically different from a tool that recommends a treatment plan, and both are different again from a clinician reviewing that same data before making a medical decision. Too much of the industry blurs these into a single reassuring word: analysis. That word does real work for marketing and very little work for accuracy. It borrows the credibility of clinical assessment while carrying none of its accountability structure — no license at risk, no malpractice exposure, no board that can revoke anything if the output is wrong.

I would break the current landscape into roughly three tiers, because I think the industry itself resists this breakdown for commercial reasons. Tier one is pure measurement: pixel-level or model-based estimation of a physical quantity — redness index, pore density, wrinkle depth, pigment distribution. This is the least risky category, and honestly the one where generative AI adds the least value over more conventional computer vision. Tier two is pattern-based decision support: the tool compares a person's measurements to population data or historical outcomes and surfaces options, without asserting a single correct answer. This is more useful and more defensible, provided the tool is explicit about its confidence and its limitations. Tier three is what I would call assertive recommendation: the tool tells a person what they should do, in language that reads as a conclusion rather than an input to a conversation. This is where I think most of the current harm risk sits, and it is also where the commercial incentive to sound authoritative is strongest, because a clear instruction converts better than an open-ended prompt to see a clinician.

I built SkinDoctor.ai's AI skin intelligence platform to sit primarily in tiers one and two, and I have resisted pressure — including, at times, my own instinct to make the product feel more impressive in a demo — to push it into tier three. A 100-KPI skin analysis and a nine-section report are genuinely useful precisely because they stay in the business of measurement and pattern comparison. The moment a consumer tool starts producing paragraphs that read like a treatment plan, it has quietly promoted itself into a role no one validated it for, and the person reading it has no easy way to tell that the promotion happened.

The specific ways generative output misleads without lying

I think it is important to be precise about the mechanism here, because saying AI is sometimes wrong is not a useful enough critique to act on. The more specific problem is that generative models produce fluent output regardless of their actual confidence, and fluency reads as authority to most people most of the time. A system can generate a paragraph explaining why a particular treatment is recommended, complete with plausible-sounding physiological reasoning, without that reasoning having been validated against the individual case in front of it. The paragraph is not lying in the sense of asserting something the model knows to be false — it is doing something arguably more concerning: generating a plausible-sounding narrative that was never checked against ground truth at all.

This is different from a bad clinical decision made by a human, because a human clinician's reasoning process, however flawed, is at least in principle auditable and subject to professional standards. A generative model's output is fluent by design, and fluency has no necessary relationship to correctness. I think this is the single most important thing for regulators, and for platforms like mine, to hold onto: the danger is not that the AI will say something obviously wrong. It is that it will say something plausible-sounding and wrong, in a register indistinguishable from something plausible-sounding and right, and the person reading it has no reliable way to tell the difference without already possessing the expertise the tool was supposed to substitute for.

What I would want out of an eventual framework

A workable governance approach should require every consumer-facing AI health tool to state, in plain language and near the output itself, not buried in terms of service, which of the tiers it occupies: measurement, decision support, or recommendation. It should treat tier-three assertive recommendation as the category that actually needs the strictest oversight, up to and including a requirement that any language resembling a treatment recommendation be routed through, or explicitly co-signed by, a licensed clinician before it reaches the consumer. It should not treat a wrinkle-scoring algorithm the same way it treats software producing unsupervised treatment narratives. Right now, marketing copy across this category makes that distinction disappear on purpose, because vagueness sells better than precision, and an unregulated tier-three tool looks identical to a well-governed tier-two tool from the outside.

I also think any framework should preserve room for physician-supervised AI tools to keep improving, and should not default to treating every generative feature as equally hazardous. A tool that uses a language model to make a nine-section skin report more readable is not the same risk category as a tool that uses a language model to generate a treatment plan. Collapsing that distinction in the name of caution would slow down genuinely useful measurement technology without meaningfully reducing the actual harm vector, which is unsupervised assertive recommendation, not language-model-assisted formatting or summarization.

What good disclosure would actually look like

I am skeptical of disclosure requirements that amount to a single sentence in a privacy policy nobody reads. If I were designing the requirement, I would want it attached to the output itself, in the same interface the user is looking at, phrased in terms a layperson can act on: this number is a measurement; this comparison is a pattern match against a population, not your specific case; this paragraph was generated by a language model and has not been reviewed by a clinician for your individual situation. That last line matters more than any of the others, because it is the one most platforms currently avoid saying out loud.

I would also want a framework that distinguishes between AI output presented as a starting point for a conversation with a clinician, and AI output presented as a conclusion. The first is genuinely valuable — it can make a consultation more efficient and more informed. The second replaces something it was never validated to replace. Most of the harm I worry about does not come from bad measurement; it comes from a measurement tool quietly relabeling itself as a conclusion tool somewhere between the engineering team and the marketing page.

My bottom line

The FDA asking this question in public is an invitation to get precise before precision is mandatory. I would rather this industry, including my own platform, define the categories honestly now than have a regulator define them for us later, after a harm has already occurred and become the case study that shapes the eventual rule. Comments are open through October 19. That is a short window for an industry that has spent years being comfortable with vague claims, and I intend to submit comments of my own arguing for the tiered disclosure approach I have described here.

— Dr. Iryna Lazuk, MD, Lazuk Esthetics, Alpharetta, Georgia

Comments


bottom of page