Across 500 posts from one AI model, the mean checker score was 79.9 out of 100 and 6.6% scored under 50. Asking the model to "make it sound human" raised the mean from 62.4 for a plain request to 94.2. The most common flags were hashtag blocks (50%, every plain and "engaging" post), closing questions and flat rhythm, and 13% of posts quoted a percentage with no source.
How the study was run
One model wrote every post: the AI assistant that also ran this analysis and drafted this page. We gave it 25 common LinkedIn prompts (leadership lessons, product launch, failure, hiring, remote work, a hot take on AI, a burnout story, fundraising, layoffs, a customer story, a pricing change, first year as a founder, conference takeaways, a book recommendation, a career change, a team win, a productivity routine, a mistake we made, an industry trend, mentorship, a sales lesson, culture, work-life, a tool we use and a year in review). Each prompt was asked in four styles: a plain request, "make it engaging", "make it sound human" and "write it in my voice" with no voice samples supplied. Each combination was run five times. That is 25 x 4 x 5 = 500 posts, averaging 138 words.
The model was told to answer the way a default assistant answers those prompts, not to write deliberately good or bad posts. It received no facts about any author or company, so every name, number and anecdote in the set was invented by the model.
Every post was then scored with the same rules as the free LinkedIn post checker, ported line for line from the page's JavaScript to Node and checked against the page on sample drafts. A post starts at 100 and loses 9 points per unit of weight for each pattern found: 3 units for an invented-looking story, 2 for a cliche opener, an unsourced figure or a closing question, and 1 for the rest. The run is dated 9 October 2026. The full dataset, with every post, score and flag, is a CSV you can download.
What we found
1. Most posts were flagged for something, but few scored badly
The mean score was 79.9 out of 100 and the median 82. 28.8% of posts had no flags at all, 44.2% scored under 80 and 6.6% scored under 50. The average post carried 1.4 flags; one post carried five. The low scores were concentrated in two of the four styles, which is the main story of this study.
Show the numbers as a table
| Group | Value |
|---|---|
| Hashtag block | 50% |
| Ends on a question to the reader | 28.2% |
| Flat rhythm | 23.2% |
| Reads like an invented story | 17% |
| Figures with no source | 13% |
| Uniform paragraphs | 5.4% |
| Cliche phrasing | 3.4% |
| Three-item lists | 1.8% |
| Enumerated lessons | 1.8% |
| Summary ending | 0.8% |
| Vague result claim | 0.2% |
2. Half the posts ended in a hashtag block, and that comes down to how this model writes
The most common flag, at 50%, was a block of three or more hashtags at the end. This number needs reading carefully. It is exactly the plain and "engaging" posts: all 125 of each ended with a hashtag line, about four tags long, and none of the 250 "human" or "voice" posts had a single hashtag. That is a property of this one model's default output, not a finding about AI-written posts in general. When asked plainly or for engagement it added hashtags every time; when asked to sound human or personal it never did. Strip the hashtag line from every post and the overall mean rises from 79.9 to 84.4, and the share under 50 falls from 6.6% to 2.8%.
3. "Make it sound human" helped more than anything else
Plain requests averaged 62.4. "Make it sound human" averaged 94.2, with no posts under 50 and 69.6% of posts flag-free, against none of the plain posts. Some of that gain is the missing hashtags, but not most of it: with the hashtag line removed, plain posts still averaged 71.4. "Write it in my voice", with no voice given, came second at 87.5. "Make it engaging" sat between plain and voice at 75.5.
Show the numbers as a table
| Group | Value |
|---|---|
| Plain request | 62.4 |
| Make it engaging | 75.5 |
| Make it sound human | 94.2 |
| Write it in my voice | 87.5 |
| Style | Mean score | Under 50 | No flags | Mean flags | Mean score without the hashtag line | Sentence-length SD (words) |
|---|---|---|---|---|---|---|
| Plain request | 62.4 | 20% | 0% | 2.7 | 71.4 | 4.1 |
| "Make it engaging" | 75.5 | 5.6% | 0% | 2 | 84.5 | 3.5 |
| "Make it sound human" | 94.2 | 0% | 69.6% | 0.4 | 94.2 | 5.6 |
| "Write it in my voice" | 87.5 | 0.8% | 45.6% | 0.7 | 87.5 | 5.2 |
4. "Make it engaging" produced the flattest rhythm and the most invented figures
Every one of the 125 "engaging" posts contained emoji, and they had the most uniform sentences in the set: a mean sentence-length standard deviation of 3.5 words, against 5.6 for "sound human". The flat-rhythm rule fired on 55.2% of engaging posts and 4.8% of human ones. Engaging posts also carried the most percentages: 26.4% contained a percentage with no named source, against 1.6% of human posts. Across all 500, 13% of posts contained an unsourced percentage. Of the 66 posts that used a percentage at all, 65 gave no source.
5. More than half the posts ended on a question
Counting the text directly, 51.8% of posts ended on a question to the reader: 93.6% of plain posts, 91.2% of engaging posts, 20.8% of voice posts and 1.6% of human posts. The checker's own rule flagged only 28.2%. The difference is a gap in the checker: in 114 of the 118 missed cases the question was followed by an emoji such as a pointing finger, which stopped the rule from seeing the question mark at the end of the line. That gap is now fixed in the free checker. Nearly all of those were "engaging" posts, which is why that style shows 0% on the rule and 91.2% in the text.
The checker rule itself flagged 28.2% of posts. The gap is explained in finding 5.
Show the numbers as a table
| Group | Value |
|---|---|
| Plain request | 93.6% |
| Make it engaging | 91.2% |
| Make it sound human | 1.6% |
| Write it in my voice | 20.8% |
6. About one post in six told a story that never happened
The invented-story rule, which looks for two or more sentences describing who did what and when, fired on 17% of posts. It fired most on "write it in my voice" (23.2%) and least on "sound human" (12%). By topic, customer stories were flagged 55% of the time and mentorship posts 40%. Because the model had no facts about any author, every specific scene in the set was made up, including many the rule did not flag: a manager who resigned, a customer who sent a whiteboard photo, a collapse at an airport gate.
7. Openers repeated within a style
Without a voice to copy, the model fell back on a few openings. "Let me tell you" opened 19 posts, all in the "my voice" style. "I want to" opened 16, all in the human and voice styles. A quarter of all first lines (25%) contained an emoji and 22% started with "I". The classic cliche openers the checker looks for, such as "In today's fast-paced world", were rare: 3.4% of posts.
8. Topic mattered less than style
Mean scores by topic ranged from 70.3 for customer stories to 89.7 for product launches. Topics that invite an anecdote (customer story, mistake we made, mentorship) scored lowest, mainly through the invented-story rule. Pricing-change posts had the most unsourced percentages (50%). With 20 posts per topic, these are directions, not rankings. The 31.8-point gap between the plain and "sound human" styles is far larger than any gap between topics.
Show the numbers as a table
| Group | Value |
|---|---|
| Customer story | 70.3 |
| Pricing change | 74.8 |
| Leadership lessons | 75.7 |
| Fundraising | 75.7 |
| Failure | 76.2 |
| Conference takeaways | 76.2 |
| Mistake we made | 76.2 |
| Mentorship | 76.6 |
| Year in review | 77 |
| Sales lesson | 77.5 |
| Burnout story | 78.8 |
| Team win | 79.3 |
| First year as a founder | 79.8 |
| Culture | 79.8 |
| Hiring | 80.2 |
| Career change | 81.5 |
| Industry trend | 81.5 |
| Layoffs | 82 |
| Productivity routine | 82 |
| Hot take on AI | 83.8 |
| Remote work | 84.7 |
| Book recommendation | 84.7 |
| A tool we use | 85.6 |
| Work-life | 87.8 |
| Product launch | 89.7 |
Four checker patterns never fired: em dash overuse, "not just X, it's Y", markdown asterisks and the universal-lesson ending. Do not read the em dash result as typical. The session that wrote these posts was working under a house rule against em dashes, and the 500 posts contain none.
Limitations
- One model, one session. Every post came from a single assistant. The five runs per cell are five separate drafts written in the same working session, not independent API calls with controlled random seeds. Other models, other versions and other settings will produce different numbers, so nothing here describes ChatGPT, Gemini or AI writing in general.
- The writer knew the rules. The model read the checker's code before it wrote the posts. It was told to write as a default assistant would, but it may still have avoided some patterns it knew were being counted. If anything, that would push the flag rates down.
- English only, LinkedIn-style prompts only. All prompts and posts are in English, and the 25 prompts were chosen by us.
- Rule-based scoring. The checker matches patterns. It misses things, as the emoji-terminated questions showed before we fixed that rule, and it can flag a real story that happens to read like an invented one. A score is a count of patterns, not a measure of quality or a detector of authorship.
- Small cells. Each style has 125 posts, but each topic has only 20, and each topic-and-style cell only 5.
What to do about it
If you use an assistant to draft LinkedIn posts, the data points to a short edit list. Delete the hashtag line if the draft has one. Delete the closing question, including any emoji after it. Remove every percentage you cannot name a source for, or add the source in the sentence. Replace any anecdote that did not happen to you with one that did, or make the point without it. Break up runs of same-length sentences with one very short line and one longer one. If you only change one thing in the prompt, the plain-versus-human gap suggests asking for the post to sound like a person rather than asking for engagement.
Then paste the result into the free checker. Each flag links to a page explaining the rule, for example hashtag blocks, closing questions, unsourced statistics and invented stories. For a longer walk-through, see how to tell if a LinkedIn post is AI-written.
Dataset and schema
Download: ai-linkedin-posts-study-2026.csv (500 rows, UTF-8, one row per post, comma-separated with quoted text fields).
| Column | Meaning |
|---|---|
| id | Row number, 1 to 500 |
| prompt | One of the 25 topics |
| style | plain, engaging, human or voice |
| seed | Run number within the prompt and style, 1 to 5 |
| words | Word count, split on whitespace |
| score | Checker score, 0 to 100 |
| n_flags | Number of patterns flagged (excluding the short-draft note) |
| flags | Names of the flagged patterns, separated by semicolons |
| ends_question | Whether the checker's last line ends in a question mark (true or false) |
| pct_count | Number of percentages in the post |
| sentence_sd | Standard deviation of sentence length in words, hashtag line excluded |
| sentence_mean | Mean sentence length in words |
| text | The full post as generated |
Common questions
Does asking an AI to make a LinkedIn post sound human actually work?
In this study, yes. For this one model, "make it sound human" raised the mean checker score from 62.4 for a plain request to 94.2, and 69.6% of those posts had no flags. Part of the gain came from the model dropping hashtags, but plain posts still averaged 71.4 with the hashtag line removed.
Do these results apply to ChatGPT, Claude or Gemini?
Not directly. All 500 posts came from a single assistant in a single session, so the numbers describe that model's output under these prompts. Other models and settings will differ. Run your own drafts through the free checker to see how your tool behaves.
Why does the checker report fewer closing questions than the study counted?
The checker looks for a question mark at the end of the last line. When a question was followed by an emoji, the rule missed it; we fixed the free checker on 9 October 2026 after this study, and the numbers here are from the version before the fix. Counting the text directly, 51.8% of posts ended on a question; the rule flagged 28.2%.
Can I download and reuse the dataset?
Yes. The CSV linked on this page holds every post with its prompt, style, run number, score, flags and rhythm measures. Please cite this page if you publish anything based on it.