Blog / Writing and voice

We generated 500 LinkedIn posts with an AI model and ran them through the checker

One model, 25 common prompts, four prompt styles, five runs each. Every post scored with the same rules as our free post checker. Here is what the data shows, what it does not show, and the full dataset.

Across 500 posts from one AI model, the mean checker score was 79.9 out of 100 and 6.6% scored under 50. Asking the model to "make it sound human" raised the mean from 62.4 for a plain request to 94.2. The most common flags were hashtag blocks (50%, every plain and "engaging" post), closing questions and flat rhythm, and 13% of posts quoted a percentage with no source.

How the study was run

One model wrote every post: the AI assistant that also ran this analysis and drafted this page. We gave it 25 common LinkedIn prompts (leadership lessons, product launch, failure, hiring, remote work, a hot take on AI, a burnout story, fundraising, layoffs, a customer story, a pricing change, first year as a founder, conference takeaways, a book recommendation, a career change, a team win, a productivity routine, a mistake we made, an industry trend, mentorship, a sales lesson, culture, work-life, a tool we use and a year in review). Each prompt was asked in four styles: a plain request, "make it engaging", "make it sound human" and "write it in my voice" with no voice samples supplied. Each combination was run five times. That is 25 x 4 x 5 = 500 posts, averaging 138 words.

The model was told to answer the way a default assistant answers those prompts, not to write deliberately good or bad posts. It received no facts about any author or company, so every name, number and anecdote in the set was invented by the model.

Every post was then scored with the same rules as the free LinkedIn post checker, ported line for line from the page's JavaScript to Node and checked against the page on sample drafts. A post starts at 100 and loses 9 points per unit of weight for each pattern found: 3 units for an invented-looking story, 2 for a cliche opener, an unsourced figure or a closing question, and 1 for the rest. The run is dated 9 October 2026. The full dataset, with every post, score and flag, is a CSV you can download.

What we found

1. Most posts were flagged for something, but few scored badly

The mean score was 79.9 out of 100 and the median 82. 28.8% of posts had no flags at all, 44.2% scored under 80 and 6.6% scored under 50. The average post carried 1.4 flags; one post carried five. The low scores were concentrated in two of the four styles, which is the main story of this study.

Share of the 500 posts flagged for each patternPercent of all posts where the checker rule fired. Patterns that never fired are listed in the text.
Share of the 500 posts flagged for each patternPercent of all posts where the checker rule fired. Patterns that never fired are listed in the text.0%30%60%Hashtag block: 50%Hashtag block50%Ends on a question to the reader: 28.2%Ends on a question to the reader28.2%Flat rhythm: 23.2%Flat rhythm23.2%Reads like an invented story: 17%Reads like an invented story17%Figures with no source: 13%Figures with no source13%Uniform paragraphs: 5.4%Uniform paragraphs5.4%Cliche phrasing: 3.4%Cliche phrasing3.4%Three-item lists: 1.8%Three-item lists1.8%Enumerated lessons: 1.8%Enumerated lessons1.8%Summary ending: 0.8%Summary ending0.8%Vague result claim: 0.2%Vague result claim0.2%
Show the numbers as a table
GroupValue
Hashtag block50%
Ends on a question to the reader28.2%
Flat rhythm23.2%
Reads like an invented story17%
Figures with no source13%
Uniform paragraphs5.4%
Cliche phrasing3.4%
Three-item lists1.8%
Enumerated lessons1.8%
Summary ending0.8%
Vague result claim0.2%

2. Half the posts ended in a hashtag block, and that comes down to how this model writes

The most common flag, at 50%, was a block of three or more hashtags at the end. This number needs reading carefully. It is exactly the plain and "engaging" posts: all 125 of each ended with a hashtag line, about four tags long, and none of the 250 "human" or "voice" posts had a single hashtag. That is a property of this one model's default output, not a finding about AI-written posts in general. When asked plainly or for engagement it added hashtags every time; when asked to sound human or personal it never did. Strip the hashtag line from every post and the overall mean rises from 79.9 to 84.4, and the share under 50 falls from 6.6% to 2.8%.

3. "Make it sound human" helped more than anything else

Plain requests averaged 62.4. "Make it sound human" averaged 94.2, with no posts under 50 and 69.6% of posts flag-free, against none of the plain posts. Some of that gain is the missing hashtags, but not most of it: with the hashtag line removed, plain posts still averaged 71.4. "Write it in my voice", with no voice given, came second at 87.5. "Make it engaging" sat between plain and voice at 75.5.

Mean checker score by prompt styleMean score out of 100, 125 posts per style. Higher means fewer flagged patterns.
Mean checker score by prompt styleMean score out of 100, 125 posts per style. Higher means fewer flagged patterns.050100Plain request: 62.4Plain request62.4Make it engaging: 75.5Make it engaging75.5Make it sound human: 94.2Make it sound human94.2Write it in my voice: 87.5Write it in my voice87.5
Show the numbers as a table
GroupValue
Plain request62.4
Make it engaging75.5
Make it sound human94.2
Write it in my voice87.5
StyleMean scoreUnder 50No flagsMean flagsMean score without the hashtag lineSentence-length SD (words)
Plain request62.420%0%2.771.44.1
"Make it engaging"75.55.6%0%284.53.5
"Make it sound human"94.20%69.6%0.494.25.6
"Write it in my voice"87.50.8%45.6%0.787.55.2

4. "Make it engaging" produced the flattest rhythm and the most invented figures

Every one of the 125 "engaging" posts contained emoji, and they had the most uniform sentences in the set: a mean sentence-length standard deviation of 3.5 words, against 5.6 for "sound human". The flat-rhythm rule fired on 55.2% of engaging posts and 4.8% of human ones. Engaging posts also carried the most percentages: 26.4% contained a percentage with no named source, against 1.6% of human posts. Across all 500, 13% of posts contained an unsourced percentage. Of the 66 posts that used a percentage at all, 65 gave no source.

5. More than half the posts ended on a question

Counting the text directly, 51.8% of posts ended on a question to the reader: 93.6% of plain posts, 91.2% of engaging posts, 20.8% of voice posts and 1.6% of human posts. The checker's own rule flagged only 28.2%. The difference is a gap in the checker: in 114 of the 118 missed cases the question was followed by an emoji such as a pointing finger, which stopped the rule from seeing the question mark at the end of the line. That gap is now fixed in the free checker. Nearly all of those were "engaging" posts, which is why that style shows 0% on the rule and 91.2% in the text.

Share of posts whose last line is a question, by styleCounted directly from the text after removing hashtags and trailing emoji, 125 posts per style.
Share of posts whose last line is a question, by styleCounted directly from the text after removing hashtags and trailing emoji, 125 posts per style.0%50%100%Plain request: 93.6%Plain request93.6%Make it engaging: 91.2%Make it engaging91.2%Make it sound human: 1.6%Make it sound human1.6%Write it in my voice: 20.8%Write it in my voice20.8%

The checker rule itself flagged 28.2% of posts. The gap is explained in finding 5.

Show the numbers as a table
GroupValue
Plain request93.6%
Make it engaging91.2%
Make it sound human1.6%
Write it in my voice20.8%

6. About one post in six told a story that never happened

The invented-story rule, which looks for two or more sentences describing who did what and when, fired on 17% of posts. It fired most on "write it in my voice" (23.2%) and least on "sound human" (12%). By topic, customer stories were flagged 55% of the time and mentorship posts 40%. Because the model had no facts about any author, every specific scene in the set was made up, including many the rule did not flag: a manager who resigned, a customer who sent a whiteboard photo, a collapse at an airport gate.

7. Openers repeated within a style

Without a voice to copy, the model fell back on a few openings. "Let me tell you" opened 19 posts, all in the "my voice" style. "I want to" opened 16, all in the human and voice styles. A quarter of all first lines (25%) contained an emoji and 22% started with "I". The classic cliche openers the checker looks for, such as "In today's fast-paced world", were rare: 3.4% of posts.

8. Topic mattered less than style

Mean scores by topic ranged from 70.3 for customer stories to 89.7 for product launches. Topics that invite an anecdote (customer story, mistake we made, mentorship) scored lowest, mainly through the invented-story rule. Pricing-change posts had the most unsourced percentages (50%). With 20 posts per topic, these are directions, not rankings. The 31.8-point gap between the plain and "sound human" styles is far larger than any gap between topics.

Mean checker score by topicMean score out of 100, 20 posts per topic, lowest first. With 20 posts each, differences of a few points are noise.
Mean checker score by topicMean score out of 100, 20 posts per topic, lowest first. With 20 posts each, differences of a few points are noise.050100Customer story: 70.3Customer story70.3Pricing change: 74.8Pricing change74.8Leadership lessons: 75.7Leadership lessons75.7Fundraising: 75.7Fundraising75.7Failure: 76.2Failure76.2Conference takeaways: 76.2Conference takeaways76.2Mistake we made: 76.2Mistake we made76.2Mentorship: 76.6Mentorship76.6Year in review: 77Year in review77Sales lesson: 77.5Sales lesson77.5Burnout story: 78.8Burnout story78.8Team win: 79.3Team win79.3First year as a founder: 79.8First year as a founder79.8Culture: 79.8Culture79.8Hiring: 80.2Hiring80.2Career change: 81.5Career change81.5Industry trend: 81.5Industry trend81.5Layoffs: 82Layoffs82Productivity routine: 82Productivity routine82Hot take on AI: 83.8Hot take on AI83.8Remote work: 84.7Remote work84.7Book recommendation: 84.7Book recommendation84.7A tool we use: 85.6A tool we use85.6Work-life: 87.8Work-life87.8Product launch: 89.7Product launch89.7
Show the numbers as a table
GroupValue
Customer story70.3
Pricing change74.8
Leadership lessons75.7
Fundraising75.7
Failure76.2
Conference takeaways76.2
Mistake we made76.2
Mentorship76.6
Year in review77
Sales lesson77.5
Burnout story78.8
Team win79.3
First year as a founder79.8
Culture79.8
Hiring80.2
Career change81.5
Industry trend81.5
Layoffs82
Productivity routine82
Hot take on AI83.8
Remote work84.7
Book recommendation84.7
A tool we use85.6
Work-life87.8
Product launch89.7

Four checker patterns never fired: em dash overuse, "not just X, it's Y", markdown asterisks and the universal-lesson ending. Do not read the em dash result as typical. The session that wrote these posts was working under a house rule against em dashes, and the 500 posts contain none.

Limitations

  • One model, one session. Every post came from a single assistant. The five runs per cell are five separate drafts written in the same working session, not independent API calls with controlled random seeds. Other models, other versions and other settings will produce different numbers, so nothing here describes ChatGPT, Gemini or AI writing in general.
  • The writer knew the rules. The model read the checker's code before it wrote the posts. It was told to write as a default assistant would, but it may still have avoided some patterns it knew were being counted. If anything, that would push the flag rates down.
  • English only, LinkedIn-style prompts only. All prompts and posts are in English, and the 25 prompts were chosen by us.
  • Rule-based scoring. The checker matches patterns. It misses things, as the emoji-terminated questions showed before we fixed that rule, and it can flag a real story that happens to read like an invented one. A score is a count of patterns, not a measure of quality or a detector of authorship.
  • Small cells. Each style has 125 posts, but each topic has only 20, and each topic-and-style cell only 5.

What to do about it

If you use an assistant to draft LinkedIn posts, the data points to a short edit list. Delete the hashtag line if the draft has one. Delete the closing question, including any emoji after it. Remove every percentage you cannot name a source for, or add the source in the sentence. Replace any anecdote that did not happen to you with one that did, or make the point without it. Break up runs of same-length sentences with one very short line and one longer one. If you only change one thing in the prompt, the plain-versus-human gap suggests asking for the post to sound like a person rather than asking for engagement.

Then paste the result into the free checker. Each flag links to a page explaining the rule, for example hashtag blocks, closing questions, unsourced statistics and invented stories. For a longer walk-through, see how to tell if a LinkedIn post is AI-written.

Dataset and schema

Download: ai-linkedin-posts-study-2026.csv (500 rows, UTF-8, one row per post, comma-separated with quoted text fields).

ColumnMeaning
idRow number, 1 to 500
promptOne of the 25 topics
styleplain, engaging, human or voice
seedRun number within the prompt and style, 1 to 5
wordsWord count, split on whitespace
scoreChecker score, 0 to 100
n_flagsNumber of patterns flagged (excluding the short-draft note)
flagsNames of the flagged patterns, separated by semicolons
ends_questionWhether the checker's last line ends in a question mark (true or false)
pct_countNumber of percentages in the post
sentence_sdStandard deviation of sentence length in words, hashtag line excluded
sentence_meanMean sentence length in words
textThe full post as generated

Common questions

Does asking an AI to make a LinkedIn post sound human actually work?

In this study, yes. For this one model, "make it sound human" raised the mean checker score from 62.4 for a plain request to 94.2, and 69.6% of those posts had no flags. Part of the gain came from the model dropping hashtags, but plain posts still averaged 71.4 with the hashtag line removed.

Do these results apply to ChatGPT, Claude or Gemini?

Not directly. All 500 posts came from a single assistant in a single session, so the numbers describe that model's output under these prompts. Other models and settings will differ. Run your own drafts through the free checker to see how your tool behaves.

Why does the checker report fewer closing questions than the study counted?

The checker looks for a question mark at the end of the last line. When a question was followed by an emoji, the rule missed it; we fixed the free checker on 9 October 2026 after this study, and the numbers here are from the version before the fix. Counting the text directly, 51.8% of posts ended on a question; the rule flagged 28.2%.

Can I download and reuse the dataset?

Yes. The CSV linked on this page holds every post with its prompt, style, run number, score, flags and rhythm measures. Please cite this page if you publish anything based on it.