Reward Modeling — Roadmap Step & Resources
Training a model to predict which outputs humans would prefer
Open on CachedInfo