7-minute read
Clinical testing can be valuable. But the phrase “clinically validated” does not tell us whether a product produced a meaningful benefit—or whether the study was designed well enough to support the marketing headline.
The Power of Two Words: Clinically Validated
“Clinically validated” sounds reassuring. It suggests that a claim has moved beyond marketing language and entered the more disciplined world of measurement, protocols and statistics.
Sometimes it has. Sometimes the phrase tells us only that a product was tested on people under defined conditions. That is useful—but it is not the same as proving that the product will deliver a noticeable result for a broad range of customers.
I recently received a launch email for an anti-ageing product. The company reported percentage improvements in wrinkles, firmness and skin texture after eight weeks. The small print said that 24 participants took part. Expert visual grading assessed wrinkles and texture, while an instrument measured firmness.
The company will remain unnamed because this is not a public dissection of one brand. The email simply provides a useful example of a much wider practice: placing an authoritative headline above a very brief study summary and allowing the reader to fill in the scientific gaps.
First, the Percentages Are Not Necessarily Weak
It would be tempting to dismiss the reported changes because the study was small. That would be too simple—and not scientifically fair.
A reduction of around one-third in a wrinkle-grading score could be clinically interesting. A smaller instrumental change in firmness might also be meaningful if the measurement was reliable, consistent and clearly different from natural variation. A study of 24 people can detect a substantial within-person change when responses are consistent.
The correct conclusion is not, “These results are poor.” The correct conclusion is, “The summary does not give us enough information to judge them.”
That distinction matters. Scientific scepticism is not automatic disbelief. It is the refusal to grant more certainty than the evidence permits.
What Does the Percentage Actually Describe?
A percentage looks precise, but precision of presentation is not the same as completeness of evidence.
Suppose an expert wrinkle score falls from 3.2 to 2.2 on a particular scale. That is roughly a 31% relative reduction. Whether the change is obvious to the person looking in the mirror depends on what the scale measures, how its steps are defined, how variable the readings are and whether the evaluator knew which images were taken after treatment.
The same problem applies to statements such as “28% visible improvement.” Does 28% describe the average change in a visual score? Does it mean 28% of participants improved? Was it compared with baseline, a placebo patch or untreated skin? The wording must make the denominator and comparator unmistakable.
Without that context, the consumer receives a number but not its meaning.
Below is a hypothetical example of how trial data might be recorded.

A Cutometer Is a Tool, Not a Verdict
The Cutometer is a well-established instrument used in cosmetic and dermatological research. It applies controlled suction to the skin and records how the skin deforms and recovers. Depending on the selected parameter, the output can help describe firmness, elasticity or viscoelastic behaviour.
That is valuable. Instrumental measurements can detect changes that are difficult to assess consistently by eye.
However, the presence of sophisticated equipment does not, by itself, validate an entire claim. The result still depends on the study design: the anatomical site, probe size, suction settings, acclimatisation, room temperature and humidity, operator technique, timing of measurements and the parameter selected for analysis.
Most importantly, an instrument cannot supply a missing control group or remove expectation bias. It measures what it measures. The study design determines what we are entitled to conclude from it.

*Image from Courage +Khazaka Electronic Website.
Expert Visual Grading Can Be Legitimate
Visual assessment is sometimes treated as inferior simply because it involves human judgement. That is not necessarily true.
Trained assessors can use standardised, validated scales under controlled lighting and positioning. Images can be randomised so the assessor does not know whether a photograph was taken before or after treatment. More than one grader can be used, and agreement between graders can be assessed.
But if the marketing summary says only “clinical expert visual grading,” the customer cannot tell whether these safeguards were used. The method may be rigorous. It may also be vulnerable to bias. The missing detail is the problem.
Is a Study of 24 People Too Small?
There is no universal minimum number that separates a valid cosmetic study from an invalid one. Appropriate sample size depends on the expected effect, variability of the measurements, design, statistical test and intended claim.
Small studies are common in cosmetic research because they are faster and less expensive. They can be useful for detecting signals, checking tolerability and guiding further development. A split-face or within-person design can also increase statistical efficiency because each participant acts as their own comparator.
Yet small samples usually provide less precise estimates and a weaker basis for generalising to the wider population. Twenty-four participants cannot represent every age, skin type, ethnicity, climate, hormonal stage or baseline condition. One or two unusual responses can also have a larger influence on an average.
If a brand wants to make a broad, confident promise, the evidence should be proportionate to that promise. A pilot-sized study may justify “observed in a small eight-week study.” It is a thinner foundation for language that sounds universal and conclusive.
The Questions Hidden Behind a Clinical Claim
Before deciding how persuasive a result is, I would want to know:
· Was the study randomised, blinded and controlled?
· Was there a placebo, an untreated site or an appropriate comparison product?
· Were all 24 participants included in the final analysis, and did anyone withdraw?
· What were the participants’ ages, skin types and baseline concerns?
· What exactly was measured, at which anatomical site and with which scale or instrumental parameter?
· Were the study conditions and skincare routines controlled?
· Were the changes statistically significant, and what were the confidence intervals?
· How widely did individual responses vary?
· Was the difference large enough to matter visually or practically—not merely detectable by an instrument?
· Was the protocol registered or the full report independently reviewed?
· Who funded, conducted and analysed the study?
A marketing email cannot contain an entire clinical report. That is reasonable. But a responsible evidence summary can state the design, comparator, sample, duration, primary endpoint and average result, then link to fuller methods. Transparency does not require drowning customers in statistics. It requires giving them enough information to understand what the headline can support.
Statistical Significance Is Not the Same as Visible Significance
A result can be statistically significant yet too small to notice in daily life. Conversely, a change that matters to participants may fail to reach statistical significance in a study that is too small or too variable.
That is why strong reporting should include more than a p-value. The effect size tells us how large the change was. A confidence interval shows the range of values reasonably compatible with the data. The distribution of individual responses tells us whether most people improved modestly or whether a few strong responders lifted the average.
For skincare, the most useful evidence often combines methods: instrumental measurements, blinded clinical grading, standardised photography and participant-reported experience. Each answers a different question. Agreement across several appropriate measures is more persuasive than one dramatic number standing alone.
Relative Change Can Make Modest Differences Look Large
Percentages are often relative to baseline. Relative changes can be legitimate, but they can sound more dramatic than the underlying absolute change.
Imagine a score changing from 1.0 to 0.7. That is a 30% relative reduction, but only a 0.3-point absolute difference. Whether that matters depends on the scale, measurement error and the smallest change people can actually see or value.
This is why “up to” claims deserve particular caution. They usually highlight the best individual response rather than the average result. Likewise, “100% of participants improved” says nothing about how much they improved unless the magnitude and definition of improvement are also supplied.
Controls Help Answer the Most Important Question
A before-and-after study can show that measurements changed during product use. It cannot automatically show that the featured ingredient caused the change.
A suitable comparator helps separate treatment effects from other influences. Depending on the product, that comparator might be an untreated area, the patch or formulation without the featured active, or another appropriate control. For a microarray patch, the delivery system itself is especially relevant: hyaluronic acid micropoints, occlusion and repeated application may affect appearance or hydration independently of the proprietary molecule.
Controls are not scientific decoration. They help answer the causal question hidden inside the marketing claim: what happened because of this product, beyond what would have happened anyway or because of its vehicle?
A Practical Translation Guide
|
Marketing phrase |
What to ask |
|
Clinically tested |
Tested how, on whom, for how long and against what? |
|
Clinically proven / validated |
Does the complete evidence justify certainty, or was one small study performed? |
|
Dermatologist tested |
Was tolerability assessed, or was efficacy also measured? |
|
Instrumentally measured |
Which instrument and parameter were used, and was the change meaningful? |
|
Visible improvement |
Who judged visibility, using what scale, and were they blinded? |
|
Up to X% improvement |
What was the average result and how many people achieved the maximum? |
|
All participants improved |
How was improvement defined, and what was the range of change? |
What the Available Summary Can—and Cannot—Tell Us
The email I received tells us that a human study was performed over eight weeks, that expert visual grading was used for some outcomes, and that an established instrument was used for firmness. Those are positive pieces of information.
It also says that all 24 participants showed measurable improvements across several outcomes. That is interesting, but it is not enough to establish why the changes occurred. Without a stated comparator, natural fluctuation, hydration, the patch vehicle, changes in routine, measurement familiarity and expectation effects cannot be separated confidently from the effect of the featured active.
Therefore, it would be wrong to declare the product ineffective. It would be equally wrong to treat the marketing headline as the final scientific word. The evidence summary supports curiosity—not certainty.
One Study Is Evidence, Not the Entire Evidence Base
A single positive study can be encouraging, particularly for an innovative product. Repetition is what makes a finding more dependable. Do similar results appear in another group, under another investigator, or in a study with a stronger comparator?
Independent replication is uncommon in commercial cosmetics because it is costly and competitors have little incentive to test another company’s finished product. That limitation should make the wording more careful, not make the first study worthless.
The most credible claim is therefore calibrated to the stage of the evidence: promising early result, replicated finding, or well-established effect. Science becomes more trustworthy when the language grows in confidence only as the evidence does.
The Azurlis™ Position: Evidence Without Theatre
Independent clinical studies are expensive. Large, well-controlled trials are beyond the present budget of many small skincare companies, including Azurlis™. That financial reality should be stated plainly—not disguised beneath borrowed authority.
The absence of a finished-product clinical trial does not mean formulation becomes guesswork. Ingredients can be selected using published evidence, supplier data, physicochemical compatibility, safety information and an understanding of how the complete formulation behaves. Products can be stability tested, preserved appropriately where required, assessed for tolerability and refined through careful use and feedback.
But ingredient evidence is not identical to clinical proof of the finished product. We should not quietly swap one for the other.
If Azurlis™ eventually conducts a consumer or clinical study, I would want the public claim to match the design precisely. I would rather say “a preliminary study observed…” than stretch limited data into a universal promise. A smaller claim that is fully supported has more integrity than a magnificent claim wearing a lab coat two sizes too large.
Good Skincare Is More Than a Percentage
Skin changes with age, hormones, UV exposure, climate, sleep, stress, nutrition, medication, health, genetics and consistency of use. No cosmetic product can isolate itself from that biological crowd and guarantee the same result for everyone.
Clinical testing can help us estimate what happened under particular conditions. It cannot turn individual biology into a vending machine: insert product, receive exact percentage.
Consumers deserve the data, but they also deserve the context. Ask what was compared, how the result was measured, how many people were studied, how variable the responses were and whether the change would matter outside the testing room.
“Clinically validated” should begin the conversation—not end it.
References
Meulyzer, C., et al. Skin involvement is measurable in Dupuytren's disease: Reliability of standardized skin elasticity measurements with Cutomoeter ® MPA 580. Hand Therapy. 2026, 31(3), 213-225.
Lee, E. et al. Artificial Intelligence Based Skin Analysis Models for Predicting Visual Grades and Device Measured Physiological Values From Facial Images. Skin Research and Technology. 2026, 32(9): 1-17.
*Courage + Khazala Electronic Website - Cutometer® Dual MPA 580