A quick color or number feels reassuring — but these apps mostly measure hazard, not real-world risk, and that difference changes everything. Here’s an honest look at what they get right, where they mislead, and how to use them without the anxiety.

EWG’s Skin Deep and the Yuka app both score cosmetics for “hazard,” giving you an easy number or color to glance at — useful for building awareness and encouraging label-reading, but limited in an important way. They largely rate hazard (whether an ingredient could ever cause harm) rather than risk (whether it actually will at the amount used), and they often ignore concentration. So safe, well-studied ingredients can get flagged, and a good score doesn’t guarantee a product suits your skin. Treat them as a starting point that prompts questions — not a final verdict — and always pair them with the actual ingredient list.
Three names come up most often. EWG’s Skin Deep is an online database from the Environmental Working Group that rates cosmetic products and ingredients on a hazard scale of 1 to 10. EWG Verified is a separate certification mark brands can carry if they meet EWG’s criteria. And Yuka is a phone app that scans a barcode and hands you a score out of 100, a color, and a list of flagged ingredients.
All three exist to solve a real frustration: ingredient lists are intimidating, and most people don’t have the background to decode them. A single number or a green-yellow-red light promises to cut through that in a second. It’s an understandable and genuinely appealing idea.
Skin Deep is a database. EWG Verified is a paid certification. Yuka is a scanning app. They’re related in spirit but different in how they work — and it’s worth not lumping them together.
At a high level, these tools assign each ingredient a hazard rating based on what’s known (or assumed) about it — drawing on scientific literature, regulatory lists, and various databases — then roll those ingredient ratings up into an overall product score. Yuka layers in other factors and its own weighting, and offers alternative suggestions on top.
The appeal is obvious: you don’t need to know what a given ingredient is, because the app has “done the research” and reduced it to a number. But that convenience hides a lot of judgment calls — which sources to trust, how to weigh conflicting studies, and crucially, what to do when the data is thin.
Those judgment calls are where reasonable experts disagree, and where the scores start to diverge from how a cosmetic chemist or toxicologist would actually assess a product. A number looks objective, but a number is only as good as the assumptions baked into it.
This is the single most important concept in the whole conversation, and it’s worth slowing down for. Hazard is the potential for something to cause harm under some circumstance. Risk is the actual likelihood of harm given how much you’re exposed to and how you’re exposed. They are not the same thing.
Classic example: water is “hazardous” — you can drown in it or even die from drinking far too much — but the risk from a normal glass of water is essentially zero. Salt, oxygen, and countless everyday substances are the same. The harm depends entirely on dose and context.
Most scoring tools primarily rate hazard, not risk. So an ingredient that could cause an issue at high concentration or in an unusual exposure may get flagged even when the amount in your lotion is trivial and well within safe use. The scary label and the real-world reality can be miles apart.
Following directly from that: the amount of an ingredient in a product usually matters enormously, and it’s exactly what a simple hazard score tends to overlook. A preservative at a fraction of a percent behaves very differently from that same preservative at ten times the concentration.
Cosmetic ingredients are generally used within established safe-use levels, and regulators set limits for many of them. A scoring app that flags an ingredient by name, without knowing or weighing how much is present, can’t distinguish between a formula using a trace amount responsibly and one using a lot. To the app, the ingredient is simply “present.”
This is the toxicology principle you’ve probably heard: the dose makes the poison. It applies to natural and synthetic ingredients alike, and it’s the reason a name-based flag can be so misleading. Concentration is often the whole story, and the score frequently doesn’t see it.
Another wrinkle: for many ingredients, complete safety data simply doesn’t exist, and the apps have to decide what to do about that. Often they take a precautionary stance — when in doubt, rate it as more concerning. That’s a defensible choice, but it has consequences.
It means a poorly-studied ingredient and a genuinely worrying one can end up looking similar in the score, and that “we don’t have much data” can read to a shopper as “this is dangerous.” Those are very different statements. Absence of evidence isn’t evidence of harm.
When you see a concerning score, it can mean “this ingredient is well-established as risky,” or it can mean “this ingredient is under-studied and rated cautiously.” The number alone won’t tell you which — and that distinction matters a lot.
Put the last few points together and you get a genuinely important takeaway: a bad score doesn’t necessarily mean a product is unsafe, and a good score doesn’t guarantee a product is right for you. Both can happen, and both do.
A well-studied, widely-used ingredient with an excellent safety record can pick up a middling score because of hazard-based or precautionary rating. Meanwhile, a product can score beautifully and still break you out, irritate your skin, or contain a fragrance you personally react to. The score didn’t know your skin.
None of this means the apps are worthless — it means the score is one data point with real blind spots. Reading it as a definitive safety grade is where people go wrong, and where a lot of unnecessary worry comes from.
It’s worth understanding the incentives, stated plainly and without conspiracy. EWG Verified is a paid certification program: brands apply and license the mark when they meet the criteria. That doesn’t make the mark meaningless, but it does mean the badge reflects participation in a paid program, not an independent ranking of every product on the shelf.
Scanning apps have their own models too — typically some mix of premium subscriptions and, in various markets, revenue tied to the alternative products they recommend. Again, that’s not inherently sinister; lots of useful free tools are funded somehow. It’s just context you deserve when a tool both scores your product and suggests what to buy instead.
Be a little more skeptical when a tool that flags your current product then points you toward a specific replacement. That’s the moment to ask how the recommendation is generated and whether the tool benefits from it.
For all the caveats, these tools do real good, and it would be unfair to pretend otherwise. Most importantly, they’ve made millions of people curious about what’s in their products — and a shopper who reads labels is better off than one who never looks.
Used as an awareness tool and a prompt to dig deeper, an app like this can be a real asset. The trouble only starts when the number stops being a starting point and becomes the final word.
The flip side is just as real. The same simplicity that makes these tools accessible is what makes them blunt instruments, and a few shortcomings show up again and again.
Perhaps the biggest cost is emotional: these tools can turn shopping into a stressful hunt for a perfect number, and can make people fearful of ordinary, safe products. Skincare should reduce stress, not manufacture it. A tool that leaves you anxious in the aisle isn’t serving you well.
So what’s the balanced approach? Keep the apps in your toolkit, but put them in their proper place — one input, not the judge and jury. Here’s a sane way to use them.
Do that, and the apps become genuinely helpful: a quick flag that prompts a two-minute look rather than a verdict that triggers a panic. That’s the sweet spot — informed, curious, and calm, rather than scared.
If you want to shop well, a handful of questions serve you better than any single number. They shift you from “is this scary?” to “is this right for me?” — which is the question that actually matters.
Try these: Is the ingredient list short and recognizable? Does it contain anything I personally know I react to, like a specific fragrance? Is there a clear reason each ingredient is there? How has my skin responded to similar products? And for anything genuinely medical, what would a dermatologist say — rather than an app?
These questions bring in the context a score can’t: your skin, your preferences, the concentration, and the purpose of each ingredient. A short, well-chosen ingredient list you understand will almost always serve you better than a high score on a product you don’t. Our guide to the personal-care swaps with the biggest impact takes this practical, low-stress approach.
As a brand whose products people might scan, here’s where we land. We’re genuinely in favor of transparency and label-reading — that’s the whole reason we keep our ingredient lists short and recognizable. If a tool gets someone reading labels, we’re for it.
But we don’t design products to chase a particular app’s score, and we’d gently encourage you not to shop by one number either. Scores shift as methodologies change, they can flag perfectly good ingredients, and they can’t know your skin. We’d rather you understand why each of our few ingredients is there than trust a color code.
We welcome the scrutiny that comes with these tools — short ingredient lists are easy to scan and easy to verify. The best version of this whole trend is a more curious, better-informed shopper, and that’s something we’re happy to be measured against. Our breakdowns of beeswax and vitamin E exist for exactly that reason.
EWG’s Skin Deep and the Yuka app are useful awareness tools with a real blind spot: they mostly measure hazard, not risk, and often ignore the concentration that determines whether an ingredient actually matters. A concerning score can mean “well-established risk” or merely “under-studied and rated cautiously,” and a great score can’t promise a product suits you.
Use them the way you’d use any quick tool — to prompt a closer look, not to deliver a final verdict. Pair the score with the ingredient list, a grasp of hazard versus risk, an eye on the incentives, and your own experience. That combination beats any single number.
Do that, and you get the best of these apps — more awareness, more label-reading, less mystery — without the anxiety and the false certainty. Informed and calm is the goal, and it’s very achievable.

Megan co-founded Bear Basics and leads design. As a mom, she writes our gentlest, most practical guides — on ingredients, sourcing, and simple family skincare — with an emphasis on honest, well-sourced, and simple. Read the full story →