Across 104 B2B software homepages, our rules found 600 checkable claims: numbers and multipliers, customer counts, and superlatives such as "#1" or "leading". 24.3% of them had any visible source on the page next to them, whether a case-study link, a footnote or a named report. 88.5% of homepages carried at least one checkable claim with nothing beside it, and only 3 sourced every one. 92.3% mentioned AI. On a hand check of 15 homepages, 87.9% of the rules' flags were correct.
How the study was run
The sample
We started from Y Combinator's public company directory, using the open yc-oss/api mirror of it as published on 10 October 2026. We kept companies whose industry is listed as B2B, whose status is Active, and which list a website and a team size. That left 2,197 companies. We split them into four team-size bands (1 to 10, 11 to 50, 51 to 200, and 201 or more people, as listed in the directory) and drew 45 companies at random from each band with a fixed seed, so anyone with the same snapshot can redraw the same list. Sampling the same number from each band means large companies are over-represented compared with the directory, where most B2B companies are tiny. We did that on purpose, so the size comparison has enough companies in every band, and every overall number in this study should be read as "across this sample", not "across YC".
We worked through the shuffled lists and stopped once we passed 100 analysable homepages. Of the 180 companies drawn, we attempted 157. 37 could not be fetched: in most cases our fetch tool only opens addresses that appear in a search result, and a search on the company's own domain did not return its homepage. Two had moved to a new domain, one domain was parked for sale and one now hosts unrelated content. That left 120 fetched homepages. We then excluded 16: 12 that sell a service rather than software (such as staffing, freight forwarding, a law firm, an accounting firm, managed security, trade finance and AI research or data labs), 1 "coming soon" placeholder, 2 pages whose fetch returned only a title and a line of metadata, and 1 page written entirely in Chinese, which our English rules cannot read. 104 homepages remain: 32 from the smallest band, 25 from 11 to 50 people, 23 from 51 to 200 and 24 from 201 or more.
Reading the homepages
Each homepage was fetched once, on 10 October 2026, with a web-fetch tool that returns text rather than raw HTML. That tool does not return a whole page verbatim. So instead of asking for the full text, we asked it the same structured question for every site: quote, word for word, the headline and subheadline; every phrase with a number, a ranking word, a customer count, a "trusted by" line, a review score or a promised outcome; every testimonial with its attribution exactly as shown; the names in any logo strip; and the literal text of any footnote, source line or case-study link, with the claim it sits next to. Numbers that appear only inside product screenshots were tagged separately and left out of the counts. We stored each answer as a text file and kept our analysis scripts in a separate folder from that fetched data. When saving, we dropped the occasional sentence of commentary the tool added, and moved a handful of figures it listed as claims but that were plainly sample data inside a product mockup (for example an example dashboard's revenue line) into the screenshot tag.
This matters for how you read the results. The rules below run on those quoted extracts, not on the raw page. A model did the first pass of picking out claim-like phrases, so a claim the extract missed is invisible to us, and a source link the extract did not mention counts as absent.
The rules
All classification is done by a short Python script with fixed regular expressions, so every flag can be reproduced. Each quoted claim line can carry more than one label.
- Number or multiplier: a percentage, a multiplier such as "3x" or "10 times", or a number next to an outcome word such as saved, faster, reduced, increased, hours or ROI. Prices, plan limits and free-trial terms are skipped.
- Customer or user count: a number next to customers, companies, teams, users, developers, brands and similar words, or "thousands of customers".
- Superlative or ranking: "#1", best, leading or leader, top, fastest, "the only", world-class, "world's best", largest, highest, unmatched, best-in-class.
- Vague outcome: a promised result with no number, such as "boost revenue", "accelerate growth", "slash denials" or "transform your workflow".
- Review score or analyst: G2, Gartner, Forrester, Capterra, IDC, a star rating or a score out of 5.
- "Trusted by" line: trusted by, loved by, used by, backed by, relied on by.
- Sourced: a number, count or superlative counts as sourced if the claim names its source in the same line (a named report, study, benchmark, survey or ranking body such as Gartner or G2), or if a source, footnote or case-study link on the page sits next to it and shares a number or several words with it. Anything else counts as unsourced.
- Testimonials: a quote counts as fully attributed if it carries a first and last name plus a role or company; partially attributed if it has only a first name and initial, a handle, a job title or a company; and anonymous if nothing is shown.
- AI wording: any mention of AI, agents or LLMs, whether the headline mentions them, and whether the page uses "AI-powered", "AI-driven", "AI-native", "AI-first" or "agentic".
"Unsourced" here means only that nothing on the homepage points to evidence. The evidence may exist on another page. That is still the gap a buyer sees: a homepage asks readers to believe a number, and in most cases gives them no way to check it from where they are standing. The full dataset, with every flag count and up to three short unsourced snippets per company, is a CSV you can download.
What we found
1. Nearly nine in ten homepages make a claim with nothing beside it
The rules found 600 checkable claims across the 104 homepages: 367 numbers and multipliers, 54 customer or user counts and 192 superlatives (a line can be more than one). 146 of them, 24.3%, had any visible source. Of those, 117 were backed by a link or footnote and 29 named their source in the same sentence.
92 of the 104 homepages (88.5%) carried at least one unsourced checkable claim. The typical page had 3 (median; mean 4.4), and 36 pages had five or more. Only 3 homepages made checkable claims and sourced all of them. 9 made no checkable claims at all; these were mostly short pages with a headline, a product description and a demo button.
Show the numbers as a table
| Group | Value |
|---|---|
| Mentions AI anywhere | 92.3% |
| At least one unsourced checkable claim | 88.5% |
| Uses a number or multiplier | 76.9% |
| Uses a superlative or ranking | 73.1% |
| Promises an outcome with no number | 69.2% |
| Shows a customer logo strip | 66.3% |
| Says "trusted by" or "backed by" | 58.7% |
| Shows customer quotes | 57.7% |
| States a customer or user count | 35.6% |
| Cites a review score or analyst | 21.2% |
2. Numbers are everywhere, and seven in ten have nothing behind them on the page
80 homepages (76.9%) used at least one number or multiplier, and 71 (68.3%) used at least one with no source next to it. Of the 367 number claims, 112 (30.5%) were sourced. The most common shapes were a multiplier such as "5x faster" (54 lines on 35 homepages), a percentage increase or lift (53 lines on 23 homepages), a percentage reduction (29 lines on 16) and a time or money saving with a figure (29 lines on 20).
The numbers that were sourced had one thing in common: they were attached to a named customer. A card reading "Ramp saved $8M and 20,000+ hours" with a link to that customer's story counts as sourced. A bare "80% Reduction in Labor Hours" in a row of stat tiles does not. The unsourced figures were mostly of the second kind: a row of three or four large numbers ("25% higher close rates", "62% more engagement", "300% higher click-through rates" on one sales-video homepage) with no customer, sample or time period attached. On one infrastructure startup's page, "~50x faster" sat in the hero area with no benchmark named. A restaurant-software homepage led its stat row with "7x Customer Retention". In each case the number may well be true; the page simply does not say compared with what, measured how, or for whom.
A related pattern: 6 homepages used animated counters that rendered as "0%" or "0x" in the fetched text, because the real figure is drawn in by a script after the page loads. A buyer with a slow connection or a screen reader can see the same thing. We counted these as number claims, because the page is making them, but we could not read the values.
Show the numbers as a table
| Group | Value |
|---|---|
| Numbers and multipliers (367) | 30.5% |
| Customer or user counts (54) | 16.7% |
| Superlatives and rankings (192) | 14.6% |
| All checkable claims (600) | 24.3% |
3. Superlatives are the least supported claim on the page
76 homepages (73.1%) called the product the best, the leading, the fastest, the first or the only something. Of 192 superlative lines, 28 (14.6%) were sourced, the lowest rate of any claim type. "Leading" or "leader" appeared in 57 lines on 36 homepages, "best" on 29, "world's" on 18, "only" on 14 and "#1" on 10. 14 homepages (13.5%) put a superlative in the headline or subheadline itself, for example "The #1 converting AI Employee for local businesses" or "The best AI software engineer".
The superlatives that were sourced were nearly all analyst or review badges: a Gartner Magic Quadrant "Leader" with a report link, a G2 "Leader" badge, a named benchmark next to "#1 for turning data questions into SQL". Self-awarded rankings ("#1 ERP Automation Platform", "#1 AI Platform for Retail Shelf Intelligence", "the industry leader in HR software") had no source on the page. 22 homepages (21.2%) cited a review score or an analyst at all.
4. Logo strips rarely link to anything
69 homepages (66.3%) showed a strip of customer logos, with a median of 9 logos. 33 of those (47.8%) had no case-study or customer-story link anywhere on the homepage. 61 homepages (58.7%) introduced logos or investors with a "trusted by", "loved by" or "backed by" line, the single most common phrase in the sample. A logo is a claim that a company uses the product. On roughly half the pages that show them, the logo is the only evidence offered.
Explicit customer counts were less common: 37 homepages (35.6%) stated how many customers, teams or users they have ("Over 60,000 businesses trust Podium", "Join over 600,000 companies using Apollo"). 9 of 54 count claims (16.7%) had a source beside them, usually a link to a customers page.
5. Most testimonials are attributed, but a fifth are not fully
60 homepages (57.7%) quoted customers, 257 quotes in all. 184 (71.6%) carried a full name plus a role or company. 58 (22.6%) were partial: a first name and initial ("Marcus D., Pro broker"), a social handle, a job title on its own ("Deputy CISO, Fortune 500 Enterprise") or a company name with no person. 15 (5.8%) had no attribution at all. 30 homepages had at least one quote that was not fully attributed. This was the best-supported element we measured, which is worth noting: founders clearly know a named quote is worth more than an anonymous one.
6. Seven in ten promise an outcome they never quantify
72 homepages (69.2%) carried at least one outcome promise with no number: "Generate more pipeline and close more deals", "Accelerate growth without increasing risk", "Slash first-pass denials", "Drive revenue, power list growth and increase margins". The rules found 169 such lines. A vague promise is not false, and some are fine as a headline. The problem is that it cannot be checked at all, so it does no persuasive work with a sceptical buyer.
7. Almost every homepage mentions AI, and six in ten lead with it
96 of 104 homepages (92.3%) mentioned AI, agents or LLMs. 63 (60.6%) put it in the headline or subheadline, and 43 (41.3%) used "AI-powered", "AI-driven", "AI-native", "AI-first" or "agentic". That includes companies selling payroll, freight, retail shelf cameras and grant management. The sample spans YC batches from 2009 to 2026, so it is not only a sample of this year's AI startups. Still, read it as a description of this sample rather than of B2B software as a whole. "AI-powered" is not a claim we score as unsupported, but it is the most common phrase on these pages that tells the buyer nothing they can verify.
8. Bigger companies make more claims and source more of them
Claims grow with company size. The smallest companies averaged 3.6 checkable claims per homepage; companies of 201 or more people averaged 8.7. The share with a source rose too, from 18.1% to 29.8%, mostly because larger companies have case studies to link to: 54.2% of their homepages had a case-study link, against 12.5% of the smallest. Superlatives rose with size as well, appearing on 95.8% of the largest companies' homepages. Even in the best-sourced band, most checkable claims had nothing beside them. With 23 to 32 homepages per band, treat these as directions, not precise rates.
Show the numbers as a table
| Group | Value |
|---|---|
| 1-10 people (32 sites) | 3.6 |
| 11-50 people (25 sites) | 4.6 |
| 51-200 people (23 sites) | 7 |
| 201+ people (24 sites) | 8.7 |
| Team size | Homepages | Checkable claims per page | Share of checkable claims with a source | Pages with a superlative | Pages with logos | Pages with a case-study link | AI in the headline |
|---|---|---|---|---|---|---|---|
| 1-10 | 32 | 3.6 | 18.1% | 62.5% | 40.6% | 12.5% | 62.5% |
| 11-50 | 25 | 4.6 | 22.8% | 64% | 60% | 44% | 56% |
| 51-200 | 23 | 7 | 22.8% | 73.9% | 91.3% | 60.9% | 65.2% |
| 201+ | 24 | 8.7 | 29.8% | 95.8% | 83.3% | 54.2% | 58.3% |
How often the rules were wrong
Pattern rules make mistakes, so we checked them by hand. We took every seventh analysed homepage in alphabetical order, giving 15 homepages and 118 flagged claim lines, and judged each label against the extracted text. Overall 87.9% of labels were correct (15 false positives out of 124).
| Rule | Items flagged in the spot-check | False positives | Precision |
|---|---|---|---|
| Numbers and multipliers | 65 | 7 | 89.2% |
| Superlatives and rankings | 29 | 4 | 86.2% |
| Vague outcome claims | 18 | 3 | 83.3% |
| "Trusted by" lines | 7 | 0 | 100% |
| Customer or user counts | 3 | 1 | 66.7% |
| Review score or analyst | 2 | 0 | 100% |
| All rules together | 124 | 15 | 87.9% |
The false positives were specific and repeatable. The number rule flagged plan or segment labels ("For Growth 101-1000"), a "Detection on Day 1" tagline, a hypothetical ("a 2% error rate means hundreds of broken edits per day"), an uptime SLA and a money-back guarantee, none of which is a performance claim. The superlative rule flagged a menu item ("Top Influencers") and phrases that describe customers rather than the product ("the world's largest companies", "the world's leading design practices"). The vague-outcome rule fired on the word "grow" in "Built for Growing Distributors" and "fast-growing companies". The count rule's one false positive was a number inside a product example.
We also checked the sourced or unsourced verdict on the 85 correctly labelled number, count and superlative claims. It was right 90.6% of the time. Of 62 "unsourced" verdicts, 7 were wrong: three rankings named their leaderboard in the sentence ("#1 on Design Arena", "Top 25 provider on OpenRouter") and the rule did not recognise those names, and four claims sat next to case-study links whose titles contained none of the words the rule looks for. One claim was marked sourced only because it shared words with an unrelated customer story. Both kinds of error lean the same way: the true share of sourced claims is probably a few points higher than 24.3%. We report the numbers from the rules as run, and did not tune the rules on the spot-check sample, because a precision figure measured on the data you tuned on is not a precision figure.
Limitations
- One directory, deliberately stratified. Every company is from YC's B2B list, which is not a random draw of B2B software companies; YC companies may write their homepages differently from bootstrapped or non-US companies. Sampling equally by team size over-represents large companies. Nothing here describes all B2B SaaS.
- We read extracts, not raw pages. The fetch tool returned quoted extracts chosen by a model, so claims or source links it left out are invisible to us. It could not see where links pointed, so "sourced" means a source is shown, not that the source is good or says what the claim says.
- Homepage only. A claim with no source on the homepage may be fully documented elsewhere on the site. We measured what a visitor sees on arrival.
- Some companies dropped out. 37 of 157 attempted homepages could not be fetched, mostly because of how our fetch tool works rather than anything about the company, and 16 fetched pages were excluded. Companies with simpler sites may be over-represented among those that remained.
- One snapshot. Every page was fetched on 10 October 2026. Homepages change often.
- English rules, imperfect precision. The rules read English only and were about 87.9% precise on a 15-site hand check. Recall, the share of real claims the rules missed, was not measured.
- Unsourced is not untrue. We did not try to verify any claim. A company with an unsourced "50x faster" may have the benchmark; it just is not on the page.
What this means for founder content
The same habits show up wherever founders write, not just on homepages. A LinkedIn post that says "we cut onboarding time by 60%" with no customer, no baseline and no time period is the homepage stat row in a different place, and readers who have learned to skim past unsourced numbers on websites skim past them in the feed too. The fix is small and the same in both places: attach each number to a named customer or a named source, say what it was measured against, and drop the superlative unless someone else awarded it.
Our free post checker at klype.ai/check flags unsourced statistics, vague result claims and the other patterns in this study before you publish. Our earlier study, 500 AI-written LinkedIn posts run through the checker, found that 13% of AI-drafted posts quoted a percentage with no source, and that 65 of the 66 posts that used a percentage gave none.
Dataset and schema
Download: b2b-saas-homepage-claims-2026.csv (120 rows, UTF-8, one row per fetched homepage, including the 16 excluded ones with the reason). The file holds the rule counts for each homepage and up to three unsourced claim snippets, each shorter than 15 words. It does not reproduce page copy.
| Column | Meaning |
|---|---|
| company, url, fetch_date | Company name from the YC directory, the address fetched, and the fetch date |
| sample_source | Where the company was drawn from |
| yc_team_size_band, yc_subindustry | Team-size band and subindustry as listed in the YC directory |
| included, exclusion_reason | 1 if the homepage is in the analysis; otherwise why not |
| n_claim_lines | Distinct claim-like lines quoted from the homepage |
| n_stat, n_stat_unsourced | Number and multiplier claims, and how many had no visible source |
| n_count, n_count_unsourced | Customer or user count claims, and how many had no visible source |
| n_superlative, n_superlative_unsourced | Superlative and ranking claims, and how many had no visible source |
| n_vague_outcome, n_third_party_rating, n_trusted_by | Outcome promises with no number, review score or analyst mentions, and "trusted by" lines |
| hero_superlative | 1 if the headline or subheadline contains a superlative |
| has_logo_strip, n_logos, has_case_study_link, has_any_source | Logo strip present and its size, any case-study link, any source text at all |
| n_testimonials and the three attribution columns | Quotes shown, split into full name with role or company, partial, and anonymous |
| ai_mention, ai_in_hero, ai_powered_phrase | AI wording anywhere, in the headline, and "AI-powered" style phrases |
| animated_counters_unreadable | 1 if counters rendered as zero in the fetched text |
| unsourced_claim_snippet_1 to 3 | Up to three unsourced claims, shortened to under 15 words |
Common questions
What share of B2B SaaS homepage claims have a source?
In this sample of 104 homepages from YC's B2B directory, 24.3% of number, customer-count and superlative claims had any visible source on the homepage, such as a case-study link, a footnote or a named report. Superlatives were the least sourced at 14.6%.
Does this mean the companies' claims are false?
No. We did not test any claim. "Unsourced" means nothing on the homepage points to evidence. The evidence may exist elsewhere on the site or in sales material.
Why are the companies not named in the article?
The point is the pattern, not any one company. Homepages change weekly, and singling out a few would misrepresent a habit that 88.5% of the sample shares. The downloadable dataset lists every company with its counts for anyone who wants to reproduce the analysis.
Can I reuse the dataset?
Yes. The CSV holds the rule counts and short snippets for every fetched homepage. Please cite this page if you publish anything based on it.