← Dictionary
Dictionary · AI

What is RLHF?

Reinforcement Learning from Human Feedback

It is an improvement process that trains artificial intelligence with human feedback.

Overview

RLHF is when the answers given by artificial intelligence are scored by humans, making the model safer and more helpful. This process ensures that the model is not only knowledgeable but also behaves in accordance with human preferences.

Analogy: It's like teaching an intern the job. The intern prepares a report, and you train him by saying 'this part is very good, but change that tone'. Over time, the intern learns what you like.

How it works

First, the model produces different responses. People rank these answers from best to worst. With this feedback, a reward model is trained and the main AI model is fine-tuned to get high scores from this reward model.

Where it is used

It is applied at the final stage so that chatbots such as ChatGPT can speak naturally and safely like humans.

Commonly confused with

It is not just training, it is the process of aligning the model's behavior.

Frequently asked questions

Why is it necessary?

Because models trained only with internet data can sometimes give crude or inaccurate information.

Are people rating it?

Yes, scoring is usually done by trained experts or broad audiences.

Related terms

This explanation was written in plain language for TreScout and machine-translated from the Turkish original · the Turkish version prevails. If something looks wrong or missing, write to hello@trescout.com. Read in Turkish →