Chapter 17
Product, UX, and Model Character
Why post-training is not only about correctness or safety, but also about shaping the assistant’s persona and fit to a product.
A model is experienced as a character
Users do not encounter weights or loss functions. They encounter a style of interaction that feels more or less helpful, warm, cautious, or annoying.
This is why character training is a product problem as much as a research problem.
Post-training decides not only what the model can do, but how it feels to use.
Persona is increasingly measurable and steerable
Work on persona vectors and assistant-axis style interventions suggests that stylistic behavior can be represented and manipulated more explicitly than many people expected.
That creates new power for product shaping, but also new responsibility.
If you can steer persona, then persona becomes part of your alignment surface.
The product loop closes the circle
The real destination of RLHF is not a leaderboard win. It is a model that behaves well in a product with real users and real feedback loops.
So the entire pipeline should be read backward from UX: what behaviors matter, how will users experience them, and what data will reflect that?
This is the chapter that turns technical alignment back into product design.
Review
5 quick checks
Questions with revealed answers.
1 Why does model character matter in post-training?
Because users experience the assistant through tone, demeanor, and interaction style, not only through task success.
2 What do persona-vector style methods suggest?
That some aspects of model behavior and persona can be represented and steered more explicitly than simple prompt wording alone.
3 Why is persona an alignment issue and not only a UX issue?
Because style changes what advice is given, how safety boundaries are expressed, and how users interpret the model’s behavior.
4 How should product concerns influence the whole RLHF pipeline?
They should define which behaviors matter, which data to collect, and which evaluations reflect real user experience.
5 What is the final takeaway of this chapter?
RLHF is ultimately about shaping model behavior in lived product contexts, not only optimizing abstract training objectives.