The role
You will work across supervised fine-tuning and reinforcement learning, implementing custom training loops and adapting language-model architectures for long-context workloads. The role includes hands-on debugging, profiling and experimentation in distributed environments.
What you’ll do
- Implement and maintain a cluster-scale codebase for SFT and RL training
- Build custom training loops
- Modify existing LLM architectures
- Design, run and evaluate focused experiments
- Identify performance bottlenecks and distributed-scaling issues
What we’re looking for
- Passion for entertainment and storytelling
- A willingness to work on difficult problems rather than easy or hype-driven ones
- A good understanding of ML and LLM fundamentals, including Transformers, attention, tokenization and GPT training objectives
- Experience training or fine-tuning language models
- Experience with JAX or PyTorch
- Working knowledge of current model research
Nice to have
- A relevant project you can demonstrate; API-prompting projects alone are not sufficient
- Experience implementing or working with distributed training systems