AMD Unveils Own Language Model Instella with Three Billion Parameters
- This model, boasting three billion parameters, is trained on AMD's proprietary Instinct MI300X GPUs and is available as open-source under a research license.
- AMD is making waves with the introduction of Instella, a family of fully open-source language models.
- The resulting language model encompasses a total of three billion parameters.
AMD Enters the AI Arena with instella, a 3 Billion Parameter Language Model
Table of Contents
- AMD Enters the AI Arena with instella, a 3 Billion Parameter Language Model
- AMD Instella: Your Questions Answered about AMD’s New AI Language Model
- What is AMD Instella?
- How many parameters does Instella have?
- What are language model parameters?
- What GPUs were used to train Instella?
- Where Can I Access Instella?
- What are the different ”Stages” of Instella?
- How does instella perform compared to other models?
- What is Instella’s architecture?
- What is the training pipeline based on?
- What is the Instella licence?
- What is a ResearchRAIL license?
- What are the ethical considerations with Instella?
- What is the importance of instella’s release?
- Key specifications of Instella-3B
Published: 2025-03-09
AMD has launched its own language model, Instella. This model, boasting three billion parameters, is trained on AMD’s proprietary Instinct MI300X GPUs and is available as open-source under a research license.
Introducing AMD Instella: A New Era for open-source AI
AMD is making waves with the introduction of Instella, a family of fully open-source language models. These models are accessible on both Github and Hugging Face. Instella comprises four models, each representing a distinct phase of the training process. Collectively, the models were trained using 4.15 trillion tokens, with the initial pretraining model, Instella-3B-Stage1, accounting for the largest share at 4.065 trillion tokens. The training was conducted on 128 Instinct MI300X GPUs.AMD states that this model demonstrates the company’s capability to leverage its own hardware for deploying scalable AI training models.
Instella’s Architecture and Performance
The resulting language model encompasses a total of three billion parameters. According to AMD, it delivers comparable or even superior performance compared to models like Llama-3.2-3B and Gemma-2-2B. the model incorporates 36 decoder layers, each equipped with 32 attention heads. These decoder layers facilitate the generation of output text, while the attention heads enable the model to focus on different components of that text. The training pipeline for Instella is based on OLMo.
Licensing and Ethical Considerations
AMD is offering Instella as open-source under a ResearchRAIL license. It’s vital to note that this license isn’t entirely open and unrestricted.It permits the model’s use for research purposes, subject to adherence to rules established by AMD.Specifically, the tool cannot be used for “harmful” applications, including fraud, discrimination, or the creation of malware.

The Significance of Instella
The release of Instella marks a critically important step for AMD in the AI landscape. By providing an open-source, high-performance language model, AMD is empowering researchers and developers to explore new possibilities in AI. The ethical considerations embedded in the licensing also highlight the importance of responsible AI growth.
Key Takeaways:
- AMD releases Instella, a 3 billion parameter language model.
- Trained on AMD Instinct MI300X GPUs.
- Available open-source under a ResearchRAIL license.
- aims for performance comparable to Llama-3.2-3B and Gemma-2-2B.
- Focus on ethical AI usage, prohibiting harmful applications.
AMD Instella: Your Questions Answered about AMD’s New AI Language Model
AMD has officially entered the AI language model arena with the release of Instella. this Q&A provides detailed answers to your most pressing questions about this new open-source model.
What is AMD Instella?
AMD Instella is a family of fully open-source large language models (LLMs) developed by AMD. The initial model, Instella-3B, contains 3 billion parameters and is designed to be accessible to researchers and developers. It is available on both Github and Hugging Face.
How many parameters does Instella have?
Instella-3B has three billion parameters.
What are language model parameters?
Language model parameters are variables that the model learns during training to make predictions. The number of parameters generally correlates with the model’s ability to understand and generate complex language.
What GPUs were used to train Instella?
Instella was trained on AMD’s Instinct MI300X GPUs, showcasing AMD’s capability to use its own hardware for complex AI tasks. Specifically, the training was conducted on 128 Instinct MI300X GPUs.
Where Can I Access Instella?
You can access Instella on the following platforms:
GitHub: https://github.com/AMD-AIG-AIMA/Instella
Hugging Face: https://huggingface.co/amd/Instella-3B
What are the different ”Stages” of Instella?
Instella comprises four models, each representing a distinct phase of the training process. 4.15 trillion tokens were used for training the models collectively, with the initial pretraining model, Instella-3B-Stage1, accounting for the largest share at 4.065 trillion tokens.
How does instella perform compared to other models?
AMD claims that Instella delivers comparable or even superior performance compared to other open-source models in its size range, such as Llama-3.2-3B and Gemma-2-2B. Though, comprehensive benchmarks are still emerging from the community.
What is Instella’s architecture?
The Instella model features:
3 Billion Parameters: This determines the model’s capacity for learning and generating text.
36 Decoder Layers: These layers are responsible for generating output text.
32 attention Heads: These enable the model to focus on different parts of the input text, improving context understanding.
What is the training pipeline based on?
The training pipeline for Instella is based on olmo (Open Language Model).
What is the Instella licence?
Instella is released under a ResearchRAIL license. This license allows open-source use for research purposes.
What is a ResearchRAIL license?
ResearchRAIL stands for Research Responsible AI License. It permits the use of the model for research subject to certain rules established by AMD.
What are the ethical considerations with Instella?
AMD has incorporated ethical considerations into Instella’s licensing. The license prohibits the use of the model for “harmful” applications.
Specifically, Instella cannot be used for:
Fraud
Discrimination
Creation of malware
* Any other harmful purpose
What is the importance of instella’s release?
Instella’s release signals AMD’s entry into the AI landscape. By providing an open-source, high-performance language model, AMD aims to foster innovation and exploration in AI research & advancement, as well as enable responsible AI.
Key specifications of Instella-3B
| Feature | Description |
| —————— | ————————————————————————————————- |
| Model Size | 3 Billion Parameters |
| Training Data | 4.15 trillion tokens |
| Architecture | 36 Decoder Layers, 32 Attention Heads |
| Training Hardware | 128 AMD Instinct MI300X GPUs |
| License | ResearchRAIL (for research purposes onyl, with restrictions on harmful applications) |
| Availability | GitHub, Hugging Face |
| Performance Goals | Comparable to Llama-3.2-3B and Gemma-2-2B |
