AI Beats Human Writers in New Benchmark
A new Creative Writing benchmark from Vulsar AI suggests that the most advanced AI models can now outperform amateur human writers in long-form writing, but professional writers still maintain a clear advantage.
The benchmark used 475 prompts and compared 24 large language models with human writers. Vulsar evaluated the submissions using a reward model fine-tuned on human preference data to estimate how often readers would prefer each response.
GPT 6 Astra Takes the Lead
GPT 6 Astra ranked first on the overall leaderboard with an 87.8% predicted win rate.
The combined human-writer baseline followed closely at 86.6%, meaning Astra narrowly surpassed the human score in the overall comparison.
However, Vulsar’s separate evaluation of professional writers showed a noticeably stronger result than the amateur human baseline. This means the benchmark does not show AI surpassing professional creative writers.
GPT 5.6 Sol ranked behind Astra with a 77.6% predicted win rate, followed by Claude Fable 5.1 at 70.3%.
| Model / Group | Predicted Win Rate |
|---|---|
| GPT 6 Astra | 87.8% |
| Human Writers | 86.6% |
| GPT 5.6 Sol | 77.6% |
| Claude Fable 5.1 | 70.3% |
| Qwen3.8-27B | 23.2% |
| DeepSeek V4.1 Flash | 19.8% |
| Gemma 4 26B | 10.9% |
Performance Drops Quickly Beyond Frontier Models
The results showed a substantial gap once the benchmark moved away from the strongest frontier models.
Models including Claude Opus 5, Kimi K3 and Grok 4.6 generally scored in the 50% to 65% range.
Smaller models performed considerably worse. Qwen3.8-27B recorded 23.2%, DeepSeek V4.1 Flash scored 19.8%, while Gemma 4 26B finished at 10.9%.
The results suggest that strong creative writing performance remains concentrated among a relatively small number of frontier AI systems.
Humans Still Write Much More
Human writers also generated considerably longer responses.
They produced an average of 2,592 tokens per prompt, compared with 1,537 tokens for GPT 6 Astra and 2,114 tokens for Claude Fable 5.1.
Discussions around the benchmark also highlighted a weakness in smaller models during longer, multi-chapter tasks. These models can lose coherence and begin repeating phrases as the story continues.
Human writers remain stronger at maintaining complicated narratives over long stretches of text and adapting their style or direction while writing.
Professional Writers Remain Ahead
The benchmark shows that the strongest AI models are beginning to compete with, and in some cases surpass, amateur human writers according to Vulsar’s predicted preference metric.
However, the results do not show AI overtaking professional writers.
For now, Vulsar’s testing suggests the gap between frontier AI and casual human writing has narrowed considerably, while skilled professional creative writing remains a harder benchmark for current models to match.
The post AI Beats Human Writers in New Benchmark appeared first on ProPakistani.



