Skip to content
AI360Xpert
Paper Breakdowns
Paper breakdown

Llama 2

The 2023 paper from Meta that introduced a commercially viable, chat-tuned foundation model, detailing the immense effort required for RLHF and safety alignment.

Paper: Llama 2: Open Foundation and Fine-Tuned Chat Models

Authors: Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Rui Hou, Inan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, Thomas Scialom · 2023

Read the paper
Llama 2 focused heavily on alignment, using multiple stages of RLHF with distinct Reward Models for 'Helpfulness' and 'Safety' to steer the chat model.
Llama 2 focused heavily on alignment, using multiple stages of RLHF with distinct Reward Models for 'Helpfulness' and 'Safety' to steer the chat model.

The Problem

While LLaMA 1 proved open-weights foundation models were viable, it had two major flaws. First, it was released under a non-commercial research license. Second, it was just a base model (a next-token predictor); it had not undergone the Reinforcement Learning from Human Feedback (RLHF) required to turn it into a helpful, conversational assistant (like ChatGPT). The open-source community's attempts at fine-tuning LLaMA 1 were impressive, but lacked the rigorous safety and alignment engineering of proprietary models.

The Idea

Meta released Llama 2 with a commercially permissible license. They increased the pre-training data to 2 Trillion tokens and doubled the context window. However, the most significant contribution of the 76-page paper was its exhaustive detailing of the alignment process. Meta developed Llama-2-Chat through rigorous Supervised Fine-Tuning (SFT) followed by multiple stages of iterative RLHF, specifically separating "Helpfulness" and "Safety" into distinct reward models to prevent them from interfering with each other.

How It Works

The alignment pipeline for Llama-2-Chat:

  1. Pre-training: A base model trained on 2T tokens. (Introduced Grouped-Query Attention for the 70B model to speed up inference).
  2. Supervised Fine-Tuning (SFT): The model is trained on tens of thousands of high-quality, human-written instruction/response pairs.
  3. Reward Modeling: Meta collected over 1 million binary comparisons from human annotators ("Which response is better?"). Crucially, they trained two separate Reward Models: one optimized to judge Helpfulness, and another optimized to judge Safety (refusing harmful requests).
  4. RLHF (PPO): The SFT model is trained using Proximal Policy Optimization to maximize the scores from the Reward Models. They also introduced "Ghost Attention" (GAtt) to help the model remember its system prompt (e.g., "act as a pirate") across a multi-turn conversation.

Why It Mattered

Llama 2 became the de facto standard for open-source LLMs in 2023. The commercial license meant startups could actually build products on it. The paper itself became a textbook on how to perform RLHF at scale, revealing the immense logistical and engineering effort required to align a model safely.

What Came After

Llama 2 dominated the open-source ecosystem until the release of models like Mistral 7B and eventually Llama 3. The separation of Helpfulness and Safety reward models became a standard practice in alignment engineering.