Skip to content
AI News HubLIVE
In-site rewrite6 min read

The Post-training Process OpenAI Used for ChatGPT

Summary

This is the third post in a four-part series about post-training. If you missed them, read part 1 and part 2. The final post, on implementing your own pipeline, will be coming October 7. Now that you understand the gist of reinforcement learning and supervised fine-tuning, it’s time to explore what post-training has actually accomplished […]

SourceO'Reilly AI & ML RadarAuthor: Sharon Zhou
The Post-training Process OpenAI Used for ChatGPT
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

This is the third post in a four-part series about post-training. If you missed them, read part 1 and part 2. The final post, on implementing your own pipeline, will be coming October 7. Now that you understand the gist of reinforcement learning and supervised fine-tuning, it’s time to explore what post-training has actually accomplished in the frontier models you know and love. Remember our prompt “Why do people like golden retrievers?” GPT-3 would often answer nonsensically. But that all changed in November 2022, with the launch of ChatGPT. Now “Why do people like golden retrievers?” actually returned a reasonable response like “Because they are affectionate, patient, and make excellent family pets” no matter who was typing (with no weird formatting tricks to consider). Anyone who could send a text message could get a response back on any topic. I’ll cover some of those behavior changes below, then take you through ChatGPT’s training pipeline as described in OpenAI’s InstructGPT paper. Conversational and helpful The most visible impact of post-training is models that can chat with you and hold a relatively long conversation. This sounds simple, but it’s not. Being conversational means more than responding to a question with an answer. The model needs to recognize when a question is ambiguous and ask for clarification or make the right assumptions in a quick response. It should adjust its tone and detail level to the context, for example being brief for a quick factual question, but thorough for a learning-oriented one. The model should be coherent across multiturn conversations without losing the thread. It also needs to handle messy real-world inputs: You attach a giant PDF and ask it to find one specific clause, and it should either find it or tell you it can’t, not hallucinate an answer. Safety and alignment Post-training is also the primary mechanism for making models safe. Safety in this context means a few things: Refusing to generate harmful content (like instructions for creating weapons, when asked) Avoiding biased or discriminatory outputs Not making up information when unsure (hallucination reduction) Respecting user privacy However, you can define safety rules in whatever way you want and teach the model to abide by them, within the limits of what your reward signals can capture. If you think cats are unsafe, because you’re a dog person, you can teach the model that in post-training—as long as you can properly encode that into a reward signal. Safety is often in tension with helpfulness. On one extreme, a model that’s too conservative will refuse reasonable requests, something that has frustrated many users. On the other extreme, a model that’s too permissive will comply with harmful ones. Navigating this trade-off is difficult. Ultimately, it comes down to determining where to draw the line, which is (as of today) a human decision within labs. Post-training is the tool to implement wherever the line is drawn. Tool use and function calling Tool use is one of the most practically important capabilities enabled by post-training. Tools include search engines, APIs, calculators, databases, and code interpreters. Being able to hit a search engine alone allows the model to not hallucinate, given its own knowledge cutoff. Tools are extremely useful ways for models to interact with the world, and are fundamental components in building agents. Tool use is a set of new behaviors. The model needs to recognize when a user’s request would benefit from an external tool. It needs to know which tools are available to it, and not hallucinate a tool. It needs to formulate a correct API call with the right parameters. It needs to interpret the results that come back and incorporate them into a natural language response. It needs to do all of this seamlessly, without the user needing to know the details of the underlying tool. This is taught almost entirely through SFT, at least initially. The training data includes many examples of conversations where the model correctly decides to invoke a tool that it has access to, constructs the right call, and processes the result. RL can further improve tool use by rewarding the model for correct tool invocations and penalizing unnecessary or incorrect ones. As an example of tool use, let’s say you’re building a veterinary appointment scheduling assistant. A user asks: “My golden retriever has been limping since yesterday. Can I see Dr. Patel this afternoon?” A pretrained model might generate plausible but fictional appointment times. A post-trained model with tool use instead calls the clinic’s scheduling API, checks Dr. Patel’s availability, and responds: “Dr. Patel has an opening at 3:15pm today. I’ve tentatively held it for you. Should I confirm?” The model needed to decide if the user’s intent was urgent, select the right tool, construct the API call with the right veterinarian and time constraints, and present the result conversationally. Tool use has expanded through the Model Context Protocol (MCP), a lightweight standard for connecting models to external services like Gmail, GitHub, or a company’s internal databases. Rather than building custom integrations for each tool, MCP provides a standard interface that any API can plug into, and different frontier models have now included learning MCP in their post-training recipes. Agentic frameworks take this further by allowing models to chain multiple tool calls together to accomplish common multistep tasks more easily. Reasoning (“thinking”) Reasoning models, or models that are trained to “think” before they answer, are an exciting result of post-training. Rather than producing an immediate response, these models generate an internal chain of thought, working through the problem step-by-step, before arriving at a final answer. As a result, their answers are more often correct than nonreasoning models that might guess at an answer. This capability has an interesting relationship with pretraining and post-training. The raw ability to reason is latent in pretrained models; they’ve been trained on text that includes mathematical proofs, logical arguments, scientific analyses, and code with comments explaining the logic. But pretrained models don’t default to reasoning. They default to pattern-matching, which often produces plausible-looking but incorrect answers. Reasoning models dramatically outperform standard models on tasks that require multistep logic: mathematical problem-solving, complex coding, scientific analysis, and planning. The improvements are not incremental. On the 2024 AIME exam, GPT-4o was only able to get 12% of problems correct on average. OpenAI’s o1 reasoning model solved 74% off the bat, with a single attempt. With 1,000 attempts and a learned scoring function to rerank the attempts, it reached 93%, a result placing it among the top 500 students who took the AIME math exam in the US. More capable reasoning requires more compute, both during training and at inference time. Scaling laws meet post-training. Models that have learned to spend more inference (test-time) compute on reasoning tend to reach better answers and therefore exhibit higher intelligence. For some frontier reasoning models, the RL post-training phase uses as much compute as the entire pretraining phase. The cost is not only in post-training compute but also in inference (test-time) tokens and latency. Reasoning takes up a lot of tokens and can result in a longer time to get a response back to the user. But the type of request matters. For a quick factual question, you don’t need reasoning. For a complex technical problem, the extra latency is well worth it. This is something that model providers can modulate during post-training. The classic ChatGPT pipeline As I mentioned above, the first post-training pipeline that captured global attention was ChatGPT’s, and it drew on the pipeline described in the InstructGPT paper. While modern systems use more advanced approaches today, this classic pipeline remains the conceptual foundation for nearly all alignment methods. The pipeline has three stages, each building on the previous one: Supervised fine-tuning (SFT) on human demonstrations Training a reward model on human preference comparisons Reinforcement learning with human feedback (RLHF) to optimize the main model using the reward model Stage 1: SFT on demonstrations The first stage is straightforward and teaches the model to follow instructions and behave like an assistant. OpenAI contracted ~40 human labelers to label their data, and they were careful to filter for people who were good at identifying harmful outputs. The labelers had to write ideal responses to prompts. But what’s interesting is that the prompts came from two sources: (1) prompts submitted by real users through the OpenAI API and (2) prompts that labelers wrote themselves. The users had to write prompts too, because these were the days before ChatGPT. There weren’t that many real users with instruction-like prompts through the API to collect. The prompts were diverse and mostly in English. There’s also an extensive data cleaning pipeline to remove duplicates and remove sensitive PII (personally identifiable information). Importantly, they split the training, validation, and test sets by human labeler. This is to avoid data leakage that could happen within a single user’s data between training and validation/testing. The resulting SFT dataset had ~13,000 prompts, all with human-labeled responses. The base model was GPT-3 at the time, a pretrained model without any post-training. Using SFT, they trained GPT-3 for 16 epochs, which was effective for the final RLHF model. This was interesting, because for the SFT stage alone, the model overfit after just 1 epoch, but ultimately SFT was an intermediate stage so they picked the best checkpoint for the final RLHF model. They also mixed in 10% pretraining data during this phase, because it would help the next RL phase. At this point, this SFT model could already be pretty useful: It could have a conversation and follow instructions, which is leaps and bounds beyond the pretrained GPT-3 checkpoint. Stage 2: Preference data and reward modeling This next stage is training the reward model. The reward model needs to grade millions of responses during RL training. In the original method, OpenAI’s team mainly trained the reward model on responses from the SFT model. However, as the policy model is trained in the RL loop and generates new, and likely better, responses from its evolving checkpoints, the reward model needs to stay robust. As a result, they also continually updated the reward model using responses from new RL checkpoints over time. To train the reward model in InstructGPT’s RLHF pipeline, OpenAI needed pairwise comparisons of two model responses from one prompt, and a label for which one is better. For example, given “What’s 2+2?” and the responses are “4” and “Yes,” the label should say “4” is better than “Yes.” Note again that these are responses from the SFT model (and later, the RL-ed models during the RL training loop), not the pretrained base model. So labeling can only happen after you’ve SFT-ed your model. If you need to retrain that model, you likely need to relabel to make sure the reward model is trained on the right distribution of data pairs. Reward model training The reward model was small at 6B parameters, for both efficiency and stability, and included a head that outputted a scalar reward. They had tried multiple sizes, but found this was more stable than using the original 175B main model. It was also more compute efficient, as the reward model would take up extra compute, for both inference and training, on top of training the main model itself. More recently, reward models have become a lot larger, but note that they don’t have to be the same model or same size model as the main model. [truncated for AI cost control]

Key points and analysis

Article intelligence

EngineersAdvanced

Key points

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • This is the third post in a four-part series about post-training. If you missed them, read part 1 and part 2. The final post, on implementing your own pipeline, will be coming Oct…

Highlights and analysis are generated automatically and may contain errors. Check the original source.