Skip to content
Learn AI by building
← World of AI

Generative AI

RLHF

Reinforcement learning from human feedback — how a raw model becomes a usable assistant.

People rank alternative responses, a reward model is trained to predict those rankings, and the language model is then optimised against that reward.

It is what converts a next-token predictor into something that answers the question asked. It also inherits the taste and the blind spots of whoever produced the rankings, which is why annotation guidelines are a serious artefact.

Also in Generative AI

JOIN NOW

Begin the first module

It is free, it is the real curriculum, and if it is not for you, you have lost nothing but an evening.

Join any time · Build AI skills at your pace