StoryScope: Investigating idiosyncrasies in AI fiction
Abstract
AI-generated stories can be distinguished from human-written fiction based on discourse-level narrative features like character agency and chronological structure, achieving high accuracy without relying on stylistic cues.
As AI-generated fiction becomes increasingly prevalent, questions of authorship and originality are becoming central to how written work is evaluated. While most existing work in this space focuses on identifying surface-level signatures of AI writing, we ask instead whether AI-generated stories can be distinguished from human ones without relying on stylistic signals, focusing on discourse-level narrative choices such as character agency and chronological discontinuity. We propose StoryScope, a pipeline that automatically induces a fine-grained, interpretable feature space of discourse-level narrative features across 10 dimensions. We apply StoryScope to a parallel corpus of 10,272 writing prompts, each written by a human author and five LLMs, yielding 61,608 stories, each ~5,000 words, and 304 extracted features per story. Narrative features alone achieve 93.2% macro-F1 for human vs. AI detection and 68.4% macro-F1 for six-way authorship attribution, retaining over 97% of the performance of models that include stylistic cues. A compact set of 30 core narrative features captures much of this signal: AI stories over-explain themes and favor tidy, single-track plots while human stories frame protagonist' choices as more morally ambiguous and have increased temporal complexity. Per-model fingerprint features enable six-way attribution: for example, Claude produces notably flat event escalation, GPT over-indexes on dream sequences, and Gemini defaults to external character description. We find that AI-generated stories cluster in a shared region of narrative space, while human-authored stories exhibit greater diversity. More broadly, these results suggest that differences in underlying narrative construction, not just writing style, can be used to separate human-written original works from AI-generated fiction.
Community
Openrouter has a fusion mode that blends the best responses from all the models. How does the writing from that compare?
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- SlopShape: Identifying AI-Generated Commercial Web Content (2026)
- CraftAlign: Feature-Grounded Evaluation and Revision Guidance for AI Stories (2026)
- How LLMs Build Fictional Worlds: Setting and Narrative Space in AI-Generated Creative Storytelling (2026)
- Detecting and Guiding LLM-Generated Korean Poetry with Interpretable Form-level Features (2026)
- LITERARYBIGFIVE: Author-Personalized Text Generation in a Unified Interpretable Space (2026)
- The Limits of Automatic Evaluation of Creativity in Large Language Models (2026)
- Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 14
zerofata/G4-MeroMero-v2-31B-GGUF
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper