Skip to content
All open roles

Data

Data Engineer Storytelling

Build and validate pipelines that turn books into high-quality structured narrative data.

About us

Pageshift is a research lab committed to pushing the frontier of AI storytelling and creativity. We are envisioning a world in which most entertainment is personalized and AI-generated. Our goal is to build the underlying story engine that powers it all. To do this, we are not afraid to explore new ways and create novel categories of model capability.

The role

You will process books and other written stories into structured datasets. The work combines LLM-driven extraction with systematic sampling and manual review, and extends to pipelines that evaluate whether model generations use storytelling concepts correctly.

What you’ll do

  • Develop and validate LLM-driven data pipelines for books and other written stories
  • Build LLM-based validation pipelines for narrative features
  • Audit extracted data through systematic sampling and manual review

What we’re looking for

  • Passion for storytelling
  • Care about data and data quality
  • A willingness to inspect a substantial amount of data manually
  • Understanding of narrative flow, world-building and character development
  • Basic LLM prompting knowledge, or a willingness to learn
  • Basic Python and data-processing knowledge, or a willingness to learn

Nice to have

  • Experience building datasets
  • Experience with theory-led analysis of stories
  • A deep understanding of current LLM limitations in story and narrative comprehension