Llama 3 logo
AI Models

Llama 3

Meta’s latest open source new generation large model

Visit Official Website
PRODUCT PREVIEW

A look at Llama 3

Open source page
Screenshot of Llama 3 on its official website
Captured from the official website · 2026-10-05Product pages can change over time.
OVERVIEW

About Llama 3

What is Llama 3

Llama 3 is the latest generation of large language model (LLM) launched by Meta Company in open source. It includes models with two parameter sizes of 8B and 70B, marking another major progress in the field of open source artificial intelligence. As the third generation product of the Llama series, Llama 3 not only inherits the powerful functions of the previous generation model, but also provides a more efficient and reliable AI solution through a series of innovations and improvements. It is designed to support a wide range of application scenarios through advanced natural language processing technology, including but not limited to programming, problem solving, translation and dialogue generation.

Llama 3 series model

Llama 3 currently provides two models, namely 8B (8 billion parameters) and 70B (70 billion parameters) versions. These two models are designed to meet different levels of application needs, providing users with flexibility and freedom of choice.

  • Llama-3-8B: 8B parameter model, this is a relatively small but efficient model with 8 billion parameters. Designed for application scenarios that require fast inference and less computing resources, while maintaining high performance standards.
  • Llama-3-70B: 70B parameter model, this is a larger model with 70 billion parameters. It can handle more complex tasks, provide deeper language understanding and generation capabilities, and is suitable for applications with higher performance requirements.

In the future, Llama 3 will also launch a 400B parameter scale model, which is still being trained. Meta also stated that it will release a detailed research paper after completing the training of Llama 3.

Llama 3 official website entrance

  • Official project homepage: https://llama.meta.com/llama3/
  • GitHub model weights and code: https://github.com/meta-llama/llama3/
  • Hugging Face model: https://huggingface.co/collections/meta-llama/meta-llama-3-66214712577ca38149ebb2b6

Improvements in Llama 3

  • Parameter size: Llama 3 provides models with two parameter sizes of 8B and 70B. Compared with Llama 2, the increased number of parameters enables the model to capture and learn more complex language patterns.
  • Training data set: The training data set of Llama 3 is 7 times larger than that of Llama 2, containing more than 15 trillion tokens, including 4 times the code data, which makes Llama 3 even better at understanding and generating code.
  • Model architecture: Llama 3 adopts a more efficient word segmenter and Grouped Query Attention (GQA) technology to improve the model's reasoning efficiency and ability to process long texts.
  • Performance improvements: Through improved pre-training and post-training processes, Llama 3 has made progress in reducing false rejection rates, improving response alignment, and increasing model response diversity.
  • Security: New trust and security tools such as Llama Guard 2, as well as Code Shield and CyberSec Eval 2, have been introduced to enhance the security and reliability of the model.
  • Multi-language support: Llama 3 adds high-quality non-English data in more than 30 languages ​​to the pre-training data, laying the foundation for future multi-language capabilities.
  • Reasoning and code generation: Llama 3 has demonstrated greatly improved capabilities in reasoning, code generation and instruction following, making it more accurate and efficient in processing complex tasks.

Performance evaluation of Llama 3

According to Meta's official blog, the Llama 3 8B model after fine-tuning with instructions is better than the models with the same level of parameter scale (Gemma 7B, Mistral 7B) in MMLU, GPQA, HumanEval, GSM-8K, MATH and other data set benchmark tests, while the fine-tuned Llama 3 70B model is better than MLLU, HumanEval, GSM-8K It also outperforms similarly sized Gemini Pro 1.5 and Claude 3 Sonnet models in other benchmark tests.

In addition, Meta has developed a new set of high-quality human assessments containing 1,800 prompts covering 12 key use cases: Seeking Advice, Brainstorming, Categorization, Closed Q&A, Coding, Creative Writing, Extraction, Character/Persona, Open Q&A, Reasoning, Rewriting, and Summarizing. By comparing with competing models such as Claude Sonnet, Mistral Medium and GPT-3.5, human evaluators conducted a preference ranking based on this evaluation set. The results showed that Llama 3 performed very well in real-world scenarios, with a minimum winning rate of 52.9%.

Technical architecture of Llama 3

  • Decoder architecture: Llama 3 adopts a decoder-only architecture, which is a standard Transformer model architecture mainly used to handle natural language generation tasks.
  • Tokenizer and vocabulary: Llama 3 uses a tokenizer with 128K tokens, which allows the model to encode language more efficiently, significantly improving performance.
  • Grouped Query Attention (GQA): In order to improve inference efficiency, Llama 3 uses GQA technology in both the 8B and 70B models. This technique reduces the computational effort while maintaining model performance by grouping queries in the attention mechanism.
  • Long sequence processing: Llama 3 supports sequences of up to 8,192 tokens and uses masking technology to ensure that self-attention does not cross document boundaries, which is especially important for processing long texts.
  • Pre-training data set: Llama 3 is pre-trained on more than 15TB of tokens. This data set is not only huge in scale, but also of high quality, providing the model with rich language information.
  • Multilingual data: To support multilingual capabilities, Llama 3’s pre-training data set contains more than 5% non-English high-quality data, covering more than 30 languages.
  • Data filtering and quality control: Llama 3’s development team developed a series of data filtering pipelines, including heuristic filters, NSFW (not suitable for the workplace) filters, semantic deduplication methods, and text classifiers, to ensure the high quality of training data.
  • Scalability and parallelization: The training process of Llama 3 adopts data parallelization, model parallelization and pipeline parallelization. The application of these technologies allows the model to be efficiently trained on a large number of GPUs.
  • Instruction Fine-Tuning: Based on the pre-trained model, Llama 3 further improves the model's performance on specific tasks, such as dialogue and programming tasks, through instruction fine-tuning.

How to use Llama 3

Developer

Meta has open sourced its Llama 3 model on GitHub, Hugging Face, and Replicate. Developers can use tools such as torchune to customize and fine-tune Llama 3 to suit specific use cases and needs. Interested developers can view the official getting started guide and go to download and deploy.

  • Official model download: https://llama.meta.com/llama-downloads
  • GitHub address: https://github.com/meta-llama/llama3/
  • Hugging Face address: https://huggingface.co/meta-llama
  • Replicate address: https://replicate.com/meta

Ordinary users

Ordinary users who are not technical and want to experience Llama 3 can use it in the following ways:

  • Visit Meta’s latest Meta AI chat assistant to experience it (Note: Meta.AI is region-locked and can only be used in some countries)
  • Visit Chat with Llama provided by Replicate to experience https://llama3.replicate.dev/
  • Using Hugging Chat (https://huggingface.co/chat/), you can manually switch the model to Llama 3

Related tools

Explore related listings in the directory.