Leading AI models have overtaken amateur writers in a new creative writing benchmark, while professional authors retain a substantial advantage.
Vulsar AI’s Creative Writing Bench V1 compares 24 large language models with human writing across 475 prompts. Its evaluator uses a reward model trained on human preference data to predict which stories readers would favour.
GPT 6 Astra leads the overall leaderboard with an 87.8% predicted win rate. The combined human reference follows at 86.6%, placing Astra ahead by 1.2 percentage points.
Read More: Baitussalam Student Scores a Perfect 200 in Cambridge O Level Mathematics
That result does not establish that AI has surpassed professional writers. The benchmark separates professional and amateur writing, revealing a much wider gap at the higher skill level.
GPT 5.6 Sol records an overall score of 77.6%. Muse Spark 1.3 follows at 70.6%, narrowly ahead of Claude Fable 5.1 at 70.3%.
The professional leaderboard tells a different story. Human professionals score 99.9%, followed by Grok 4.6 at 78.5% and Astra at 76.2%.
In the amateur category, Astra scores 93.9%, Sol reaches 85.2%, and human writers record 79.7%.
Vulsar uses different prompt sets for the professional and amateur categories. Their combined results form the overall ranking, making the category breakdown essential to interpreting the headline figures.
Performance drops sharply beyond the leading models. Claude Opus 5, Kimi K3 and Grok 4.6 fall within the roughly 50% to 65% band overall.
Further down, Qwen3.8-27B scores 23.2%, while DeepSeek V4.1 Flash reaches 19.8%. Gemma 4 26B records 10.9%.
Read More: China Launches World’s First Floating Artificial Island for Deep-Sea Research
The spread suggests that competitive creative writing remains concentrated among a relatively small group of advanced systems. Results from those leaders do not describe the capabilities of every available AI model.
Human writers average 2,592 tokens per prompt, compared with Astra’s 1,537. Claude Fable 5.1 averages 2,114 tokens.
Those figures show a clear difference in output length. They do not, by themselves, explain differences in narrative quality or establish that longer stories perform better.
The supplied report also highlights concerns about smaller models repeating phrases and losing coherence during multi-chapter tasks. It describes human strengths in sustaining complicated narratives and changing direction while writing.
However, Vulsar’s published task asks for one short story per prompt. Its results therefore cannot directly establish performance across novels or extended multi-chapter projects.
The scores represent evaluator predictions, rather than direct reader votes on every benchmark response. An 87.8% overall score also does not mean Astra beats humans in 87.8% of direct comparisons.
Read More: Here Is How Your Research Could Build a $100 Million Company
Vulsar averages comparisons across eligible opponents and prompts. It excludes refusals and empty responses from scoring.
The company also cautions that high scores do not verify originality against existing writing or training material.
The findings support a narrower conclusion: leading models compete strongly with amateur writing under this evaluation, while professionals remain ahead.
