The role
You will process books and other written stories into structured datasets. The work combines LLM-driven extraction with systematic sampling and manual review, and extends to pipelines that evaluate whether model generations use storytelling concepts correctly.
What you’ll do
- Develop and validate LLM-driven data pipelines for books and other written stories
- Build LLM-based validation pipelines for narrative features
- Audit extracted data through systematic sampling and manual review
What we’re looking for
- Passion for storytelling
- Care about data and data quality
- A willingness to inspect a substantial amount of data manually
- Understanding of narrative flow, world-building and character development
- Basic LLM prompting knowledge, or a willingness to learn
- Basic Python and data-processing knowledge, or a willingness to learn
Nice to have
- Experience building datasets
- Experience with theory-led analysis of stories
- A deep understanding of current LLM limitations in story and narrative comprehension