How do you prove a product is safe when there are millions of versions of it? That is the core problem facing hyper-personalized beauty brands, and a new technical paper from custom haircare company Prose, published in Cosmetics & Toiletries, lays out one answer.
The paper describes a two-part framework: a risk-based method for clearing safety and regulatory requirements across a massive formulation space, and a live A/B testing system, borrowed from software engineering, that tests ingredient changes on real customers at scale. The headline number is 82,447 consumer survey responses. The more interesting story is what that number can and cannot tell us.
The Personalization Problem
Personalized brands use online consultations, algorithms, and modular formulation to generate products tailored to individual hair or skin profiles. Prose says its model can produce combinations numbering in the millions.
That creates a regulatory headache. Under the EU Cosmetics Regulation (EC No. 1223/2009), every formula placed on the market requires a full safety and quality assessment. Testing millions of variants one by one is not realistic.
Testing the Boundaries, Not Every Formula
Prose’s solution is what it calls a Formulation Architecture. For each product category, the team maps every eligible ingredient and sets concentration limits. Toxicologists and formulators then select three types of formulas to test:
- Texture formulations: establish baseline product behavior.
- Realistic formulations: reflect the most common customer combinations.
- Critical formulations: the riskiest combinations, such as protein- or sugar-rich formulas that could support microbial growth.
These go through standard industry testing: preservative challenge testing under ISO 11930:2019, stability testing at 40°C for three months and for 12 months at multiple temperatures, in vitro eye irritation and 48-hour patch testing, and efficacy testing. For a personalized conditioner, all critical formulations met acceptance criteria, and an automated combing simulator showed reduced hair breakage across hair textures.
The logic is familiar to toxicologists: if the worst-case combination is safe, the range inside that envelope can be approved. This is the more scientifically conventional half of the paper.
A/B Testing at Home
The second half is where the beauty-tech angle comes in. Prose replaced the conditioning silicone in its conditioners with a new cationic silicone crosspolymer and tested the change directly through its e-commerce platform.
Customers were randomly assigned, without being told which group they were in, to receive either the original formula (70%) or the reformulated one (30%). Three weeks after delivery, they rated their results on a 1 to 5 scale.
The results favored the new formula:
- Overall satisfaction: 74% vs. 72.2% for the control.
- Dissatisfaction: 7.5% vs. 8.1%.
- Dry, under-conditioned hair: reported by 26.3% vs. 29.8%.
- Soft, manageable hair: 67.7% vs. 65.2%.
Both Chi-square testing and a Bayesian analysis supported the differences as statistically significant.
Reading the Numbers Carefully
These findings deserve context. With more than 82,000 responses, even very small differences will reach statistical significance. A 1.8 percentage point gain in satisfaction is real in a statistical sense, but whether an individual customer would notice it is a separate question.
A few other limitations are worth flagging:
- Self-reported outcomes: satisfaction surveys measure perception, not objective hair condition.
- Who responds: customers who answer surveys may differ from those who do not.
- Short follow-up: three weeks captures early impressions, not long-term effects. The authors acknowledge this and suggest six- to 12-month tracking.
- Industry authorship: the paper was written by Prose’s own team and published in a trade journal, not a peer-reviewed one.
The paper also does not break down the consumer results by hair type or texture, which matters for any brand built on personalization. A formula change that helps on average could still underperform for specific groups, including textured and coily hair, which is often underrepresented in product testing.
The Bigger Picture
What makes this paper notable is less the conditioner result than the model behind it. Direct-to-consumer brands own the full loop: the consultation data, the formulation, the delivery, and the feedback. That lets them run randomized experiments at a scale traditional consumer panels cannot match.
Done rigorously, this could make product development far more responsive to real-world use. The open question is whether the industry will pair large-scale consumer data with objective measurements and subgroup analysis, so that “personalized” reflects proven performance for each person rather than a slightly better average.