Reward Modeling — Roadmap Step & Resources

Training a model to predict which outputs humans would prefer

Open on CachedInfo