Methodology
How every grade, every count and every percentage on this site was produced — what was searched, what was not, and which of these rules exist because this project broke them first.
This site is written by a non-clinician, and it says so on every page. Its authority is not borrowed from a credential. It comes from method and traceability: each claim carries the identifier of the record it came from, and each record is one a reader can open without asking our permission. That is the whole offer. The rest of this page is the method, in enough detail to be argued with.
Which products are in, and which are out
A comparison is only as honest as its universe, so ours is published as a rule a reader could re-run rather than as a description of one. A SKU is in if all three of the following hold and none of the disqualifiers applies.
- Form. Its primary marketed identity is live microorganisms — a name containing probiotic, synbiotic, biotic or an organism genus; or a retailer probiotic filing with a CFU count on the front of pack; or an FDA-approved live biotherapeutic. The test is subtractive: delete the organisms, and if the product still has a coherent name and category, it fails. That rule is what keeps greens powders, multivitamins, protein powders, fibre supplements and grocery ferments out of a men’s probiotic audit.
- Market. A US buyer could complete a purchase from a US-facing channel on the re-derivation date — or could have, and the product has since been delisted or discontinued, in which case it stays in as a dated negative finding carrying the evidence of its removal. Dropping a discontinued product would quietly delete the finding that a brand withdrew its men’s SKU while keeping every women’s counterpart.
- Evidence. The SKU resolves to a specific URL fetched in that pass, at one of three declared tiers: a brand or retailer page that returned HTTP 200; a marketplace page that returned a populated product title; or a storefront that returned a bot challenge — never a 404 — with existence corroborated by a second independent source. No SKU enters on a name alone. A product named by an editorial ranking page with no fetched URL is out as unverifiable, and that is recorded as a finding about the ranking page.
Applied that way on 31 July 2026, across four consolidated source passes and 203 raw rows, the universe is 167 SKUs. That number superseded an earlier 36-SKU universe; all 36 survived the re-derivation, none was dropped. 166 of those 167 rows have a published page. The one that does not is the Amazon listing of Culturelle Daily Health 8-in-1 for Men, which the teardown proved is the Culturelle Men’s Daily Health SKU on a second storefront — so it is one page, not a coverage gap, and the register says so rather than leaving the arithmetic to look tidy.
| Figure | Value | What it is |
|---|---|---|
| The universe | 167 | Rows the universe rule admits, after dedup and exclusion, as re-derived on 31 July 2026. It is a coverage figure, not a denominator: no percentage on this site divides by it. |
| Pages built | 166 | Universe rows with a published page in this build. The 1 row without one is the Amazon listing of Culturelle Daily Health 8-in-1 for Men, which the teardown proved is the Culturelle Men’s Daily Health SKU on a second storefront. |
| D₂₇ | 27 | The frozen founding cohort. The products torn down first, and the denominator for every product-level percentage on this site. It does not grow. |
| D₂₉₄ | 294 | Every organism named on the label of those 27 products. The denominator for every strain-level percentage. It does not grow either. |
| D₂₄ | 24 | Cost-per-day statistics only. Three of the 27 founding products have no verifiable US price, so the denominator shrank rather than three prices being estimated. |
Two tiers, and what each page is therefore worth
The 166 pages are not uniform, and a reader is entitled to know which kind of work is behind a row before weighing it against another. Tier 1 is 107 SKUs — those marketed to men, those named as a numbered pick by one of the harvested ranking pages, those disclosing a strain designator that has published human trial literature under that exact designator, and those with the sales visibility that means most men are already holding them. Tier 2 is the remaining 59: the long tail, still audited, but only on the facts needed to score it.
The depth label is printed on every product page, and it does not always follow the tier — three SKUs were reached at their assigned tier and then yielded no readable label at all.
| Depth | Products | What was actually done |
|---|---|---|
| Forensic teardown | 105 | The full pass: panel transcribed organism by organism, every strain designator queried against PubMed as the exact string printed, every certification chased to its issuer’s register, the literature behind each marketing claim opened and read. All Tier 1. |
| Label audit | 58 | The lighter pass: strain list and designators, per-strain quantity, CFU basis, price and cost per day, certification checked at the register, and whether a government label record exists. No narrative teardown. All Tier 2. |
| Label not recovered | 3 | The SKU was reached and no label surface could be recovered from any route. These pages are marked, not scored: no strain list, no CFU basis grade, no cost per day. The failure is the finding. |
Separately, 62 of the 166 pages carry a rendered verification flag — a statement, on the page, that some specific claim on it did not survive re-checking, naming what did not survive and what still stands. One of those is in the founding cohort; the other 61 are addenda. The schema will not let a page carry a silent caveat: an entry whose data did not survive verification fails the build unless it also says what failed.
What is excluded, and under which rule
This page previously carried a list of nine products “identified in the universe and never torn down”. That list is gone because it is no longer true: all nine now have published teardown pages, and the draft was asserting something the collection itself disproved.
What remains is the real exclusion list, and it is a list of rules, not of products we did not get to. Every candidate below was met during discovery and kept out for a stated reason. Publishing it is the only way a reader can tell the difference between “not in this audit” and “not on the market”.
| Rule | What it keeps out, and which candidates it kept out |
|---|---|
| Category error (fails rule A) | The product’s primary marketed identity is not live microorganisms. The test: delete the organisms and ask whether the product still has a coherent name and category. Athletic Greens AG1 is out on that test and is retained as a search-results finding instead, because two ranking pages call it a probiotic for men. So are SmartyPants Multivitamin for Men, Gundry MD PrebioThrive, Benefiber, Metamucil Fiber Gummies, Pendulum Polyphenol Booster and Gut Fuel, and Tiny Health, which sells microbiome test kits and no probiotic at all. |
| Audience out of scope (fails rule D1) | Women-only vaginal or urogenital positioning, or pediatric. Listed in the register rather than dropped, so a later pass does not re-discover them as coverage gaps: Perelel, RepHresh Pro-B, AZO Complete Feminine Balance, vH essentials, OLLY Happy Hoo-Ha, the Culturelle, Align, Member’s Mark and up&up women’s SKUs; and BioGaia Protectis Baby Drops, Visbiome GI Care Kids, Culturelle Kids, SmartyPants Kids, OLLY Kids. |
| Duplicate under the dedup key (fails rule D3) | Two listings sharing brand, labelled strength, dose form and unit count are one row, whatever the retailer calls them. Fifteen listings folded into a row that already existed — including Garden of Life’s Amazon men’s SKU with 11,588 ratings and Transparent Labs’ single product under three different editorial names. Where the key cannot be resolved because neither listing discloses a strength, the pair is carried as one row and flagged, never as two and never silently merged. |
| Unverifiable — no working URL (fails rule C) | No URL was fetched at any of the three declared evidence tiers, so the SKU enters on a name alone or not at all, and it does not enter. Nucific Bio X4, Bio-Kult US, Evivo, TruBiotics, Biotics Research BioDoph-7 Plus, Douglas Laboratories Multi-Probiotic 4000, and eight affiliate-funnel products named by one ranking page. None of these is asserted to be discontinued: a 404 on a guessed path is evidence about the URL, not about the product. |
| Brand defunct — no auditable label anywhere | Jetson. Re-confirmed 31 July 2026: the brand domain now redirects to an unrelated parked target, two further domains return no A record, and the Amazon listing 404s. There is no panel to tear down, so it is recorded as a brand-level dated negative finding rather than as a SKU. |
| Not purchasable in the US (fails rule B) | Carried as named comparators rather than dropped, because two of the best-evidenced organisms in this category are ones a US buyer cannot legitimately buy — which is itself a finding. Mutaflor (E. coli Nissle 1917) and Enterogermina (B. clausii): recorded as “no US retail channel found”, never as “no FDA record”, which would be a different and unsupported claim. |
Why the percentages divide by 27 and not by 166
This is the subtlest thing on the site, and it is the one a hostile reader is most likely to call a trick, so it is stated here at length rather than in a footnote.
Why freeze it. A denominator that grows every week means no percentage on the site can be checked against any earlier version of itself. Add one product and 51.9%, 33.3%, 7.4%, 14.6% and 55.6% all move at once, quietly, and every citation anyone ever made to any of them breaks without a single sentence on the page changing. A number that cannot be checked against the version somebody quoted is not a checkable number, and checkability is the entire offer here. So: freeze the cohort, grow by addendum, and date both.
The cost, stated plainly, because it is real. The frozen cohort is 27 of 166. That means every percentage on this site describes the founding sample and not the whole market, and not even the whole of this site’s own coverage. “Fifteen of 27 panels are machine-unreadable” is a fact about those 27 products as of the freeze date; it is not a claim that 55.6% of the US market is machine-unreadable, and it may not be read as one. The rule that follows is that a percentage on this site names its denominator in figures — “of 27”, “of 294” — in the same sentence, so a reader who never arrives at this page still carries the scope along with the number. That rule is not yet perfectly applied: at least one sentence on this site still says “in our set” where it should print the figures, and it is listed below as an open gap rather than described as fixed.
What would change it. Only a new dated cohort, published as one, with the old one left standing beside it. The founding percentages are never recomputed in place. If a second cohort is ever frozen, it gets its own symbol, its own date and its own denominator, and the two are reported side by side — never merged, never averaged.
The teardown: how a bottle becomes a row
The same nine steps run on every product, in this order. Steps 2 to 5 are the ones that make a row checkable rather than assertable, so they are stated in more detail than the rest.
- Take nothing from anyone, and buy nothing either. No samples, no review units, no early access, no brand contact before publication — and no purchases: no product in this audit was bought, opened, weighed or cultured. This is a documentary audit, which is why anyone with a browser can reproduce it, and why nothing here can tell you whether a bottle’s contents match its label.
- The NIH Dietary Supplement Label Database is the label source of record, and the record id is published on the row. Every product row that rests on a label record carries its DSLD identifier — for example DSLD 321381 (opens in a new tab) — so a reader can open the same record we read, rather than take our transcription on trust. Superseded label versions are pulled as well as current ones. That is not thoroughness for its own sake: it is the only reason this site can document seven label regressions, including a silent strain substitution under an unchanged product name and CFU count, and a brand that deleted a through-expiry potency commitment rather than improving it.
- Cross-check the UPC against the live product page. Mandatory, every time. A government record proves what a label version said; it does not prove that version is the one on the shelf. The UPC is the join key, and where a record is flagged off-market the row says so instead of claiming to describe the SKU currently sold. This step is also where the sharpest pricing finding in the audit comes from: an archived brand page advertising $47.99 and $0.53/serving side by side, against a label record for the same UPC confirming 90 capsules, 30 servings, three capsules per serving — so the advertised “per-serving” price is the per-capsule price and the real cost is $1.60/day, understated three-fold.
- Where a brand hard-blocks automated retrieval, fall back to the Internet Archive and say so. Three brands in the founding cohort return 403 or a bot challenge to automated requests for their own product pages, which means a shopper’s price-comparison tool or allergen checker cannot read them either. Where that happened, the dated Internet Archive capture is the artefact the row rests on, and the row identifies it as an archived capture rather than a live read. One product’s “no Supplement Facts panel of any kind” finding was confirmed against an archived capture as well as both live pages before it was allowed to publish.
- Where the panel exists only as a JPEG, transcribe it from the brand’s own image and publish the asset URL. Thirteen of the 27 founding-cohort panels are raster images, so someone has to read them by eye. A hand transcription is a claim like any other, so it ships with the address of the exact image file it was read from, and the transcription reproduces the artwork including its mistakes — one panel carries the brand’s own misprint, “Daily Value (DV) not esablished”, and it is transcribed with the typo intact because correcting it silently would be editing the evidence.
- Record the CFU figure and its temporal basis verbatim — at manufacture, through expiry, or none stated — or record that the label gives no basis at all.
- Record the per-organism quantity, or record that the panel gives only a blend total and the per-strain dose is therefore unknowable.
- Resolve every strain designator, or state that it resolves to nothing. The method for that is the next section.
- Compute price per day from the label’s own dosing instruction, never from the marketing; then compare the label’s dose against the dose used in the trial the marketing evokes.
The gate
No product page can be built without a dated artefact behind it. The content schema enforces it at build time: an entry with no label-database record and no dated capture fails the build rather than shipping with a gap. This is a structural rule, not a good intention, because good intentions are exactly what fail under publishing pressure.
What the method runs into
Most of what this audit found is not scandal. It is that the category is unreadable. Fifteen of the 27 founding-cohort panels cannot be parsed by a machine, which means a screen-reader user cannot access the ingredient list, and neither can a price-comparison tool, an allergen checker, a drug-interaction checker or an AI shopping assistant. That is 55.6% of the founding cohort, which is a statement about those 27 products and not about the market.
| Status | Products | Note |
|---|---|---|
| Panel retrievable as structured text anywhere | 12 | For 11 of those 12, the only machine-readable copy is the US government’s label record — not anything the brand published. |
| Panel published as a raster image only | 13 | Transcribed by hand from the brand’s own artwork. |
| Ingredient content published as text, but not as a Supplement Facts panel — and not retrievable by a machine | 1 | Seed DS-01 Daily Synbiotic. The 24 organisms, the four blend allocations and the AFU totals are text on the product page. There is no “Supplement Facts” heading, no serving size, no %DV column and no Other Ingredients list; seed.com returns a Cloudflare challenge to every automated request, including for its sitemap, and Seed has no government label record. So the text exists and no tool can reach it. |
| No Supplement Facts panel published in any form | 1 | Onnit Total GUT HEALTH. Not as text, not as an image, not in the government database, and not in the Internet Archive’s capture of either product page. |
The government label database, and what we can and cannot say about it
This table used to carry one more figure: “Ten of the 27 (37.0%) are not present in the NIH label database at all.” That sentence made a claim this project is not in a position to make, and it is withdrawn. What the collection can support is narrower and is now what the site says.
12 of the 27 founding-cohort teardowns (44.4%) rest on no NIH label-database record at all. The label on each of those pages was read off a panel photograph, a live brand-page capture or an Internet Archive capture instead. That is a fact about this audit’s evidence, and it is counted from the artefacts attached to each teardown, so it moves when they do.
It is not the same statement as “these products are absent from the database”, and the difference is the whole reason the figure was wrong. Of those 12, ten teardowns record a database query that came back empty for the SKU — Ancient Nutrition SBO Probiotics Men’s Once Daily and Codeage Men’s Daily Probiotic are the two that record no query at all, so nothing is established about them either way. And even for the ten, “no record found” is not “no record exists”: two of those brands are in the database under older or different SKUs — Onnit’s 2016 packet components, and a 2014 off-market Renew Life record — so what is missing is a government copy of the label being sold now, which is exactly the gap that makes a label regression invisible.
From a strain code to the evidence behind it
A strain designator is a scientific claim printed on a bottle, so it is checked like one. Every designator on every panel is queried against PubMed as the exact string printed on the label, through the public search API, so the query is a string anyone can re-run unchanged. Both the query and the number it returned are published. That is the whole point: an unpublished zero is an assertion, and a published query that returns zero is a finding.
Counts alone are not enough either, because the failure mode in this category is not a missing record — it is a record that matches the string and has nothing to do with the product. So every hit is opened and read before it is counted as evidence.
| Query, as run | Records | What they actually are |
|---|---|---|
"Bb-18"[tiab] | 0 | No record of any kind — not an RCT, not an animal study, not an in-vitro paper. The designator that replaced Bb-03 on a national brand’s label, with no notice. |
"Bb-03"[tiab] | 1 | A paper on UV-B tolerance in Beauveria isolates. An entomopathogenic-fungus isolate code that happens to collide with a probiotic designator. |
"Bl-05"[tiab] | 4 | B. licheniformis in a carp, duck-egg-white mayonnaise, an encapsulation study, an in-vitro growth study. Four hits, zero relevance. |
"St-21"[tiab] AND randomized controlled trial[pt] | 20 | All twenty are false positives: ST-21 is the acupuncture point Liangmen. The fielded query "Streptococcus thermophilus St-21" returns 2 — a case report and a mouse C. difficile study. No human RCT. |
"plantarum 299v" AND randomized controlled trial[pt] | 34 | A designator that genuinely resolves. This is what a strain code is supposed to look like when you follow it. |
"HN019" AND "irritable bowel" | 0 | Re-verified 29 July 2026. The strain marketed hardest for regularity has no IBS trials at all. |
P3-OM | 0 | And "P3OM" also returns 0, while "Lactobacillus plantarum OM" returns 16 and "plantarum OM" AND probiotic returns 7 — every one of which resolves to something else, including pig solid-state fermentation and EPR-ENDOR physics. ClinicalTrials.gov returns 0 studies. |
The three traps this method is built around
- The false positive. Short alphanumeric codes collide with everything — acupuncture points, fungal isolates, fish-feed studies. A raw count is worthless until every record behind it has been read.
- The one-character trap. One product contains L. plantarum 6595, and the 1996 founding paper states the identity in as many words: strains 299 (= DSM 6595) and 299v (= DSM 9843) are different strains PMID 8779562 (opens in a new tab) . The consumer literature belongs to 299v. The label is technically accurate and practically misleading, and only reading the source catches it.
- The untraceable house code. A brand-invented code looks exactly like an ATCC or DSM accession to a shopper who has learned “look for the strain code”. One panel carries 31 codes in a house format of which 30 return zero PubMed records; another carries 24. A third case is worse: ten designators built mechanically from the species name — genus initial, two digits, species initial — appeared on federally filed labels in 2021 and 2022, all returning zero hits, and had been deleted by 2024.
The 2020 renaming, applied to both name forms
In 2020 the genus Lactobacillus — 261 species as at March 2020 — was reorganised into 25 genera: 23 novel genera, plus the emended genus Lactobacillus, plus Paralactobacillus PMID 32293557 (opens in a new tab) . Labels did not follow. Six years later the valid name Lacticaseibacillus appears on 14 of roughly 214,780 labels in the NIH database, against 6,124 for the superseded Lactobacillus — 0.23% as many.
So every literature search here is run under both name forms, old and new. A search under only the current name misses the trial; a search under only the label’s name misses the reclassified record. Searching one form and reporting the result as complete is a manufactured absence.
Certification: the issuer’s register, never the badge
A certification mark on a product page is a claim by the brand. A listing in the certifier’s own public register is a fact about the product. This site only ever reports the second. Verifiable here has one meaning and it is narrow: the certifier’s own database was queried and that specific product was found in it. On that definition, 2 of the 27 founding-cohort products (7.4%) hold a certification we could confirm at the issuer. As with every percentage here, the denominator is the frozen founding cohort and not the 166 products this site now carries.
The two are Legion, confirmed at Labdoor with score 100.0, lot 0385A6, tested 14 July 2026; and Sports Research, confirmed as Non-GMO Project Verified through the Project’s own open API — brand id 9714, product 2-25768-3023, status Verified.
When the register cannot be reached
This happens constantly, and it is a finding about the verification infrastructure rather than about any brand. Confirmed at capture and again on re-check: nsf.org returns 403, the B Lab directory returns 403 behind a bot challenge, legacy Non-GMO Project product-finder paths return 404, and vegan.org’s certified-products search returns 403 to automated requests.
The grading scale
Grades are assigned strain-first, because that is the unit the evidence exists in — and the unit the regulator uses. Example 36 of the FTC’s Health Products Compliance Guidance (December 2022) considers an advertiser relying on two well-controlled trials of a different strain, in capsule form, in a Japanese population, to support a claim for its drink, and names “significant differences that would affect whether the findings could reasonably be expected to translate.” Three independently disqualifying mismatches: strain, format, population. Example 34 adds the population rule directly — advertisers should not rely on research in a specific test population for claims aimed at the general population without first making sure the extrapolation is sound. Reporting the sex breakdown of every trial on this site is that test, applied to ourselves.
- Grade A — replicated, at least one independent trial Two or more RCTs of the same named strain at a stated dose, consistent direction, at least one without strain-owner authorship, pooled or independently replicated, on a clinical endpoint.
- Grade B — one adequately powered strain-level trial One adequately powered RCT of a named strain that met a pre-registered primary endpoint, or strain-level pooling across multiple RCTs, certainty no lower than “low”, no fatal design defect.
- Grade C — one positive trial with a material design defect One positive RCT of a named strain with a material design defect — manufacturer-run and unreplicated, a secondary endpoint, a co-intervention that cannot be dismantled, or a surrogate endpoint.
- Grade D — class-level pooled evidence only Class-level pooled evidence only. Real, but it transfers to no product on a shelf, because the pool contains dozens of different organisms.
- Grade F — nothing qualifies Nothing qualifies. Searched, and either empty or refuted.
Two rules that apply to every grade on this site
- No outcome anywhere in this corpus reaches Grade A at strain level. Not one strain has two independent RCTs with a non-owner replication on a clinical endpoint, in the outcome it is sold for. That is the finding, and it is the headline.
- Nothing here is graded on male-specific evidence, because none exists. Across every strain and every outcome, zero trials have shown a benefit larger in men than in women — and the two sex-comparative analyses that do exist both run towards women.
Two further rules govern how a graded claim is written. Funding and author affiliation are disclosed on the claim, in the sentence, not in a reference list at the bottom. And every claim carries its evidence type visibly — systematic review, RCT, observational, animal, in vitro, regulatory document, marketing assertion, or searched-and-found-nothing — because “a study showed” is the sentence that does most of the lying in this category.
What this site does not know
This section is written against our own interest, and it is the most important one on the page.
How we know that gap is real: it already broke one of our own absolutes
An early draft of this project contained the sentence “There is no human randomised controlled trial of any probiotic strain with erectile function as a primary endpoint. Not one. Anywhere.” It was refuted by a 240-man randomised trial published in the Journal of Sexual Medicine, which measured the International Index of Erectile Function and reported improvement across all domains PMID 42059569 (opens in a new tab) [RCT] — registered on the Iranian Registry of Clinical Trials, which ClinicalTrials.gov does not index. One PubMed query finds it.
The documented failure mode of a single query
A count from one query is not a fact about the literature; it is a fact about that query. Three ways it goes wrong, each of which happened here and each of which is logged:
- A clean zero that was twenty false positives. A designator query was recorded as returning 0 when it returns 20 — all of them irrelevant, because the code collides with an acupuncture point. The conclusion was right and the reported number was wrong, which is the more dangerous combination, because it survives being checked casually and fails being checked properly.
- A scoped result reported as an unscoped one. The MeSH-scoped query
probiotics[mh] AND "erectile dysfunction"[mh]returns 0 and that is exactly right. The claim built on it — that there is no indexed literature at all at that intersection — is not, because the free-text search returns 12 records. The scoped statement publishes; the unscoped one does not. - A count nobody can reproduce. A search reported as 115 results was not
reproducible; a fielded version,
probiotic*[tiab] AND prostatitis[tiab], returns 26. If the query string is not published, the count is not evidence, and the fix is to publish the string or drop the number.
Things this site wants and does not have
Re-verification and corrections
Every page carries a last-verified date, and absence statements are re-run at each one, because absence is the most perishable claim on this site: a database gains records, and a sentence that was true in July can be false in October without anyone touching the page. Label records are re-checked on the same schedule — the audit has already documented seven cases of a brand removing a disclosure rather than improving it, so decay is the expected state, not the exception.
The cadence itself will be published as a measured number once a full re-check has been run and timed, and if the measured rate of change turns out to be faster than the cadence we can fund, the promise gets reduced in public rather than quietly missed.
How this site is checked before publication
Every finding on this site went through a separate adversarial pass whose only job was to break it — re-resolving each identifier, checking each verbatim quote against the source string, and re-running every absence claim. Where it found a claim was too strong, the weaker and correct version is what publishes, which is why several sentences here are less dramatic than the ones you will read elsewhere.
and it holds a correction to the same standard — the original wording verbatim, a date, and no removals. Corrections from anyone, including a brand whose product is audited here, are published on identical terms.