An important part of creative writing is generating original ideas and interpreting familiar concepts in new and unexpected ways, while maintaining high-quality output. But without enough output diversity, even a model whose first generation feels highly creative soon becomes repetitive: it starts to feel like AI slop. Reinforcement learning (RL) is widely used to improve the capabilities of large language models, but these improvements can come at the cost of creativity and output diversity.
We developed a novel RL training framework to improve quality and diversity together, creating Pagestorm Diversity from Pagestorm, our previously released book-writing model. Our approach combines a quality reward model with a diversity reward focused on choices beyond the prompt's requirements, encouraging distinctive responses that remain faithful to the user's request.
Training focuses on book previews because they establish the creative direction that later stages develop. A broad prompt can branch into romance, political intrigue, or adventure at this stage, making these early choices especially influential.
RL fine-tuning often rewards each response independently for how well it completes a task. In a standard task-focused setup, the environment provides a reward for how well the response fulfills the request, and training encourages responses that score well. This evaluates whether an individual answer is successful; it does not directly reward a range of different successful answers. Repeatedly producing a similar, high-scoring response can still satisfy that objective.
To address this challenge, we add a diversity reward that evaluates responses in relation to one another. We generate a group of responses to the same prompt and evaluate both how well each response fulfills the request and how its creative choices compare with those in the group. Quality is assessed individually, while diversity is a property of the relationship between responses.
The quality reward model primarily measures alignment with the user's prompt: does the response follow the instructions and include the requested elements? It also checks general writing requirements, including readability and grammatical correctness. This gives training a clear signal for producing responses that fulfill the request and read well.
The diversity reward adds an incentive to explore different valid interpretations of the same request. Several responses can follow the prompt equally well while sharing almost the same premise, characters, or conflict. We want the model to develop distinct creative directions across those responses, giving the user a broader range of ideas to choose from.
The challenge is deciding what should vary. A simple request such as "I want a fantasy story" leaves room for a magical romance, a struggle for the throne, or a journey through an unfamiliar world. More detailed prompts narrow that freedom: specified characters, settings, or events should remain consistent while the model explores the choices left open.
To identify those choices, an LLM reads the prompt and each generated response, separating prompt-required content from content the model chose itself. Only the model-chosen content enters diversity scoring. Shared requirements are excluded, so the model is not penalized for repeating the elements the user asked every response to contain.
A semantic encoder maps that model-chosen content into embeddings, representations of its meaning. We compare the embeddings using cosine similarity, so diversity reflects differences in ideas rather than just differences in wording. Two responses that rephrase the same premise remain close in this semantic space.
We then estimate local density within the group: how closely related is each response to the others? A response surrounded by similar ideas receives a lower diversity reward, while one in a less crowded semantic neighborhood receives a higher reward. Training therefore encourages the model to move beyond ideas it repeatedly generates for that prompt.
Pushing diversity indefinitely can lead to nonsensical outputs. We therefore use a similarity cutoff to set an effective density target: the creative ideas should be sufficiently spread out, without continually pushing them further apart. Once two responses are different enough, that comparison provides no additional incentive to separate them. This gives training a point at which the diversity is sufficient, while the quality reward continues to encourage clear writing that follows the prompt.
Finally, we combine the quality and diversity rewards for each response and use Group Relative Policy Optimization (GRPO) to optimize the model with this combined signal. After each update, we generate the next group and repeat the process, encouraging responses that fulfill the request while exploring a wider range of creative possibilities.
We wanted strong generalization and stable training dynamics, so we chose model latent adapters. This approach adds newly initialized parameters to a frozen Pagestorm base model, giving RL dedicated parameters to train while keeping all original model weights untouched. Following the design of adapter-based fine-tuning, these added parameters form two small MLP modules within each Transformer layer. In our setup, the modules adapt the model through sparse latent edits to its internal representations. Keeping these changes focused within the added modules allows us to adjust the model's behavior while supporting generalization beyond the directly trained stage.
With this model latent adapter setup, the benefits extended beyond book previews. We observed quality and diversity improvements in later generation stages too, despite training only the preview component. The gains were smaller than at the directly trained stage, but the improvements carried into the rest of the book-generation process.
The evaluation below shows that Pagestorm Diversity improves quality and diversity at the same time, without sacrificing one for the other. Under the same generation settings, it follows the prompt more closely while exploring a broader range of creative directions.
We evaluated both models on the same 360 prompts used in our original Pagestorm paper, generating 16 book previews per prompt. Temperature and all other sampling parameters were fixed across models. Quality was measured with the training quality reward model. Diversity was measured with the training diversity signal, using semantic local-density estimation of model-chosen content within each group of 16 responses. The plot reports the mean scores across this evaluation.
To see how these differences appear in the writing itself, we also compare complete book previews from the base Pagestorm model and Pagestorm Diversity, generated from the same dragon-and-princess fantasy prompt under identical sampling settings. We visualize each model's per-token entropy across its outputs, with darker shading indicating higher entropy. All six generations from each model are included below.
Together, the evaluation and examples show what we are building with Pagestorm Diversity: stronger writing and more distinctive starting points, giving the model a wider range of creative possibilities to develop into a book.
Acknowledgements
Research supported with Cloud TPUs from Google's TPU Research Cloud (TRC).
Citation
@misc{zierstek2026pagestormdiversity,
title = {Training {Pagestorm} for Quality and Diversity},
author = {Zierstek, Jan and Batelic, Matteo and Sch{\"o}nenberger, Tim},
year = {2026},
month = oct,
howpublished = {Pageshift research blog},
url = {https://pageshift.ai/research/pagestorm-diversity}
}