Axiom Futures AI Safety Course Week 4 notesJuly 10, 2024collecting human feedback, fitting a reward model, and optimizing the policy with RL.