{"id":108381,"date":"2026-08-01T19:59:35","date_gmt":"2026-08-01T19:59:35","guid":{"rendered":"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/"},"modified":"2026-08-01T19:59:35","modified_gmt":"2026-08-01T19:59:35","slug":"amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus","status":"publish","type":"post","link":"https:\/\/youzum.net\/zh\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/","title":{"rendered":"AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs"},"content":{"rendered":"<p class=\"wp-block-paragraph\">AMD released <strong>Instella-MoE-16B-A3B<\/strong>, a fully open Mixture-of-Experts language model trained from scratch on Instinct MI300X and MI325X GPUs. The model holds 16B total parameters but activates only 2.8B per token. AMD is publishing weights from every training stage, along with data mixtures, training configs, and inference code. Two systems-level choices carry the release: Gated Multi-head Latent Attention and FarSkip-Collective connectivity. <\/p>\n<h2 class=\"wp-block-heading\"><strong>Is it deployable?<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">Partly. The weights ship under a <strong>ResearchRAIL license for academic and research purposes only<\/strong>, so this is not a drop-in commercial model. The <a href=\"https:\/\/github.com\/AMD-AGI\/Instella-MoE\">training codebase<\/a> is MIT licensed, and that is the more reusable asset here.<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Company level:<\/strong> AI research labs, university groups, and enterprise R&amp;D teams with data-center GPU capacity. Not a fit for lean startups wanting a hosted commercial endpoint.<\/li>\n<li><strong>Industries:<\/strong> semiconductor and cloud infrastructure, AI tooling vendors, and academic research.<\/li>\n<li><strong>Applications:<\/strong> reproducing an end-to-end MoE recipe, studying expert-parallel serving, evaluating 64K long-context behavior, and running RL post-training experiments.<\/li>\n<li><strong>Serving cost:<\/strong> 16B parameters in BF16 need roughly 32 GB of weight memory, so one high-memory accelerator suffices. AMD ships <a href=\"https:\/\/github.com\/sgl-project\/sglang\">SGLang<\/a> inference code.<\/li>\n<\/ul>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\"><img fetchpriority=\"high\" decoding=\"async\" width=\"2162\" height=\"1152\" data-attachment-id=\"81487\" data-permalink=\"https:\/\/www.marktechpost.com\/2026\/08\/01\/amd-instella-moe-16b-a3b-fully-open-mixture-of-experts-llm\/screenshot-2026-08-01-at-11-57-14-am-2\/\" data-orig-file=\"https:\/\/www.marktechpost.com\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1.png\" data-orig-size=\"2162,1152\" data-comments-opened=\"0\" data-image-meta='{\"aperture\":\"0\",\"credit\":\"\",\"camera\":\"\",\"caption\":\"\",\"created_timestamp\":\"0\",\"copyright\":\"\",\"focal_length\":\"0\",\"iso\":\"0\",\"shutter_speed\":\"0\",\"title\":\"\",\"orientation\":\"0\",\"alt\":\"\"}' data-image-title=\"Screenshot 2026-08-01 at 11.57.14\u202fAM\" data-image-description=\"\" data-image-caption=\"\" data-large-file=\"https:\/\/www.marktechpost.com\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-1024x546.png\" src=\"https:\/\/www.marktechpost.com\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1.png\" alt=\"\" class=\"wp-image-81487\" \/><figcaption class=\"wp-element-caption\">https:\/\/rocm.blogs.amd.com\/artificial-intelligence\/instella-moe\/README.html<\/figcaption><\/figure>\n<\/div>\n<h2 class=\"wp-block-heading\"><strong>Architecture<\/strong><\/h2>\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/rocm.blogs.amd.com\/artificial-intelligence\/instella-moe\/README.html\">Instella-MoE<\/a> is a decoder-only MoE with 27 layers, hidden size 2048, 16 attention heads, and a 128,896-token vocabulary. Each MoE layer uses 2 shared experts plus 6 routed experts selected from 64. That yields 2.8B active parameters against 16B total. A Multi-Token Prediction objective is used during pre-training and mid-training.<\/p>\n<p class=\"wp-block-paragraph\">There are two structural choices that are important to know. <strong>Gated MLA<\/strong> adds a lightweight learned output gate to Multi-head Latent Attention. A dedicated linear projection derives an input-conditioned gate, applied multiplicatively before the output projection. <strong><a href=\"https:\/\/github.com\/AMD-AGI\/FarSkip-Collective\">FarSkip-Collective<\/a><\/strong> passes outdated and partial activations into the MoE and attention layers, overlapping expert-parallel communication with computation. AMD reports a 12.7% pre-training speedup and up to a 39.2% reduction in time to first token when serving with expert parallelism.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Training pipeline<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">Pre-training covers 7.1T tokens from open corpora including Nemotron-CC-v2, MegaMath, FineMath, RefineCode, and TxT360. Mid-training uses Dolma3 Dolmino 100B across three data variants, merged by weight averaging. A long-context stage extends the window from 4K to 64K using YaRN, an increased RoPE theta, and document masking.<\/p>\n<p class=\"wp-block-paragraph\">Post-training runs SFT on Dolci-Think-SFT-7B plus Nemotron mixtures, ending on a feedback-driven 512K-example set targeting measured weaknesses. DPO follows, with router bias updates and the auxiliary load-balancing loss disabled to prevent degradation. RL runs in the Miles framework: 1,400 steps of instruction-following RLVR, then Multi-Teacher On-Policy Distillation to fold that gain back without losing math or code.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Results<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">The base checkpoint averages 76.7, the strongest among fully open models, ahead of Moonlight-16B-A3B (76.2), SmolLM3-3B-Base (70.5), OLMo-3-7B (70.1), and OLMoE-1B-7B (61.9). It trails Qwen3.5-4B-Base (79.5). It leads on WinoGrande (86.5) and scores 65.7 on HumanEval+. Long-context averages are 41.5 on HELMET and 79.4 on RULER.<\/p>\n<p class=\"wp-block-paragraph\">Post-training climbs from SFT (71.58) to DPO (72.67) to Think (73.22), above Olmo3-7B-Think (71.97), Gemma-4-E4B think (70.47), and Qwen3.5-4B (69.73). IFEval rises from 77.08 to 83.70.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Interactive explainer<\/strong><\/h2>\n<div>\n<\/div>\n<p class=\"wp-block-paragraph\">\n<h2 class=\"wp-block-heading\"><strong>Key Takeaways<\/strong><\/h2>\n<ul class=\"wp-block-list\">\n<li>16B total parameters, 2.8B active per token: 2 shared plus 6 of 64 routed experts.<\/li>\n<li>Gated MLA and FarSkip-Collective give a 12.7% training speedup and 39.2% lower TTFT.<\/li>\n<li>Trained end-to-end on AMD Instinct MI300X and MI325X with ROCm, Primus, and Miles.<\/li>\n<li>Base averages 76.7 and Think averages 73.22, both leading fully open peers.<\/li>\n<li>ResearchRAIL weights limit commercial use; the MIT-licensed training code does not.<\/li>\n<\/ul>\n<\/p><p class=\"wp-block-paragraph\">\n<hr class=\"wp-block-separator has-alpha-channel-opacity\" \/>\n<\/p><p class=\"wp-block-paragraph\">\n<\/p><p class=\"wp-block-paragraph\">Check out the<strong>\u00a0<a href=\"https:\/\/rocm.blogs.amd.com\/artificial-intelligence\/instella-moe\/README.html\">ROCm blog<\/a>, <a href=\"https:\/\/huggingface.co\/collections\/amd\/instella-moe\">Hugging Face collection<\/a> <\/strong>and<strong> <a href=\"https:\/\/github.com\/AMD-AGI\/Instella-MoE\">GitHub<\/a>.\u00a0<\/strong>Also,\u00a0feel free to follow us on\u00a0<strong><a href=\"https:\/\/x.com\/intent\/follow?screen_name=marktechpost\" target=\"_blank\" rel=\"noreferrer noopener\"><mark>Twitter<\/mark><\/a><\/strong>\u00a0and don\u2019t forget to join our\u00a0<strong><a href=\"https:\/\/www.reddit.com\/r\/machinelearningnews\/\" target=\"_blank\" rel=\"noreferrer noopener\">150k+ML SubReddit<\/a><\/strong>\u00a0and Subscribe to\u00a0<strong><a href=\"https:\/\/www.aidevsignals.com\/\" target=\"_blank\" rel=\"noreferrer noopener\">our Newsletter<\/a><\/strong>. Wait! are you on telegram?\u00a0<strong><a href=\"https:\/\/t.me\/machinelearningresearchnews\" target=\"_blank\" rel=\"noreferrer noopener\">now you can join us on telegram as well.<\/a><\/strong><\/p>\n<p class=\"wp-block-paragraph\">Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?\u00a0<strong><a href=\"https:\/\/forms.gle\/wbash1wF6efRj8G58\" target=\"_blank\" rel=\"noreferrer noopener\"><mark>Connect with us<\/mark><\/a><\/strong><\/p>\n<p class=\"wp-block-paragraph\"><strong>Sources:<\/strong> <a href=\"https:\/\/rocm.blogs.amd.com\/artificial-intelligence\/instella-moe\/README.html\">ROCm blog<\/a> \u00b7 <a href=\"https:\/\/huggingface.co\/collections\/amd\/instella-moe\">Hugging Face collection<\/a> \u00b7 <a href=\"https:\/\/github.com\/AMD-AGI\/Instella-MoE\">GitHub<\/a><\/p>\n<p>The post <a href=\"https:\/\/www.marktechpost.com\/2026\/08\/01\/amd-instella-moe-16b-a3b-fully-open-mixture-of-experts-llm\/\">AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs<\/a> appeared first on <a href=\"https:\/\/www.marktechpost.com\/\">MarkTechPost<\/a>.<\/p>","protected":false},"excerpt":{"rendered":"<p>AMD released Instella-MoE-16B-A3B, a fully open Mixture-of-Experts language model trained from scratch on Instinct MI300X and MI325X GPUs. The model holds 16B total parameters but activates only 2.8B per token. AMD is publishing weights from every training stage, along with data mixtures, training configs, and inference code. Two systems-level choices carry the release: Gated Multi-head Latent Attention and FarSkip-Collective connectivity. Is it deployable? Partly. The weights ship under a ResearchRAIL license for academic and research purposes only, so this is not a drop-in commercial model. The training codebase is MIT licensed, and that is the more reusable asset here. Company level: AI research labs, university groups, and enterprise R&amp;D teams with data-center GPU capacity. Not a fit for lean startups wanting a hosted commercial endpoint. Industries: semiconductor and cloud infrastructure, AI tooling vendors, and academic research. Applications: reproducing an end-to-end MoE recipe, studying expert-parallel serving, evaluating 64K long-context behavior, and running RL post-training experiments. Serving cost: 16B parameters in BF16 need roughly 32 GB of weight memory, so one high-memory accelerator suffices. AMD ships SGLang inference code. https:\/\/rocm.blogs.amd.com\/artificial-intelligence\/instella-moe\/README.html Architecture Instella-MoE is a decoder-only MoE with 27 layers, hidden size 2048, 16 attention heads, and a 128,896-token vocabulary. Each MoE layer uses 2 shared experts plus 6 routed experts selected from 64. That yields 2.8B active parameters against 16B total. A Multi-Token Prediction objective is used during pre-training and mid-training. There are two structural choices that are important to know. Gated MLA adds a lightweight learned output gate to Multi-head Latent Attention. A dedicated linear projection derives an input-conditioned gate, applied multiplicatively before the output projection. FarSkip-Collective passes outdated and partial activations into the MoE and attention layers, overlapping expert-parallel communication with computation. AMD reports a 12.7% pre-training speedup and up to a 39.2% reduction in time to first token when serving with expert parallelism. Training pipeline Pre-training covers 7.1T tokens from open corpora including Nemotron-CC-v2, MegaMath, FineMath, RefineCode, and TxT360. Mid-training uses Dolma3 Dolmino 100B across three data variants, merged by weight averaging. A long-context stage extends the window from 4K to 64K using YaRN, an increased RoPE theta, and document masking. Post-training runs SFT on Dolci-Think-SFT-7B plus Nemotron mixtures, ending on a feedback-driven 512K-example set targeting measured weaknesses. DPO follows, with router bias updates and the auxiliary load-balancing loss disabled to prevent degradation. RL runs in the Miles framework: 1,400 steps of instruction-following RLVR, then Multi-Teacher On-Policy Distillation to fold that gain back without losing math or code. Results The base checkpoint averages 76.7, the strongest among fully open models, ahead of Moonlight-16B-A3B (76.2), SmolLM3-3B-Base (70.5), OLMo-3-7B (70.1), and OLMoE-1B-7B (61.9). It trails Qwen3.5-4B-Base (79.5). It leads on WinoGrande (86.5) and scores 65.7 on HumanEval+. Long-context averages are 41.5 on HELMET and 79.4 on RULER. Post-training climbs from SFT (71.58) to DPO (72.67) to Think (73.22), above Olmo3-7B-Think (71.97), Gemma-4-E4B think (70.47), and Qwen3.5-4B (69.73). IFEval rises from 77.08 to 83.70. Interactive explainer Key Takeaways 16B total parameters, 2.8B active per token: 2 shared plus 6 of 64 routed experts. Gated MLA and FarSkip-Collective give a 12.7% training speedup and 39.2% lower TTFT. Trained end-to-end on AMD Instinct MI300X and MI325X with ROCm, Primus, and Miles. Base averages 76.7 and Think averages 73.22, both leading fully open peers. ResearchRAIL weights limit commercial use; the MIT-licensed training code does not. Check out the\u00a0ROCm blog, Hugging Face collection and GitHub.\u00a0Also,\u00a0feel free to follow us on\u00a0Twitter\u00a0and don\u2019t forget to join our\u00a0150k+ML SubReddit\u00a0and Subscribe to\u00a0our Newsletter. Wait! are you on telegram?\u00a0now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?\u00a0Connect with us Sources: ROCm blog \u00b7 Hugging Face collection \u00b7 GitHub The post AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs appeared first on MarkTechPost.<\/p>","protected":false},"author":2,"featured_media":108382,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"pmpro_default_level":"","site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"_pvb_checkbox_block_on_post":false,"footnotes":""},"categories":[52,5,7,1],"tags":[],"class_list":["post-108381","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-club","category-committee","category-news","category-uncategorized","pmpro-has-access"],"acf":[],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v25.3 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs - YouZum<\/title>\n<meta name=\"description\" content=\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/youzum.net\/zh\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/\" \/>\n<meta property=\"og:locale\" content=\"zh_CN\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs - YouZum\" \/>\n<meta property=\"og:description\" content=\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\" \/>\n<meta property=\"og:url\" content=\"https:\/\/youzum.net\/zh\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/\" \/>\n<meta property=\"og:site_name\" content=\"YouZum\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/DroneAssociationTH\/\" \/>\n<meta property=\"article:published_time\" content=\"2026-08-01T19:59:35+00:00\" \/>\n<meta name=\"author\" content=\"admin NU\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"\u4f5c\u8005\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin NU\" \/>\n\t<meta name=\"twitter:label2\" content=\"\u9884\u8ba1\u9605\u8bfb\u65f6\u95f4\" \/>\n\t<meta name=\"twitter:data2\" content=\"3 \u5206\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/#article\",\"isPartOf\":{\"@id\":\"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/\"},\"author\":{\"name\":\"admin NU\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c\"},\"headline\":\"AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs\",\"datePublished\":\"2026-08-01T19:59:35+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/\"},\"wordCount\":667,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\"},\"image\":{\"@id\":\"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/#primaryimage\"},\"thumbnailUrl\":\"https:\/\/youzum.net\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-WnMGN8.webp\",\"articleSection\":[\"AI\",\"Committee\",\"News\",\"Uncategorized\"],\"inLanguage\":\"zh-Hans\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/\",\"url\":\"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/\",\"name\":\"AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs - YouZum\",\"isPartOf\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/#primaryimage\"},\"image\":{\"@id\":\"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/#primaryimage\"},\"thumbnailUrl\":\"https:\/\/youzum.net\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-WnMGN8.webp\",\"datePublished\":\"2026-08-01T19:59:35+00:00\",\"description\":\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\",\"breadcrumb\":{\"@id\":\"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/#breadcrumb\"},\"inLanguage\":\"zh-Hans\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"zh-Hans\",\"@id\":\"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/#primaryimage\",\"url\":\"https:\/\/youzum.net\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-WnMGN8.webp\",\"contentUrl\":\"https:\/\/youzum.net\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-WnMGN8.webp\",\"width\":2162,\"height\":1152},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/youzum.net\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/yousum.gpucore.co\/#website\",\"url\":\"https:\/\/yousum.gpucore.co\/\",\"name\":\"YouSum\",\"description\":\"\",\"publisher\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/yousum.gpucore.co\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"zh-Hans\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\",\"name\":\"Drone Association Thailand\",\"url\":\"https:\/\/yousum.gpucore.co\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"zh-Hans\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png\",\"contentUrl\":\"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png\",\"width\":300,\"height\":300,\"caption\":\"Drone Association Thailand\"},\"image\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/\"},\"sameAs\":[\"https:\/\/www.facebook.com\/DroneAssociationTH\/\"]},{\"@type\":\"Person\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c\",\"name\":\"admin NU\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"zh-Hans\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png\",\"contentUrl\":\"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png\",\"caption\":\"admin NU\"},\"url\":\"https:\/\/youzum.net\/zh\/members\/adminnu\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs - YouZum","description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/youzum.net\/zh\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/","og_locale":"zh_CN","og_type":"article","og_title":"AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs - YouZum","og_description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","og_url":"https:\/\/youzum.net\/zh\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/","og_site_name":"YouZum","article_publisher":"https:\/\/www.facebook.com\/DroneAssociationTH\/","article_published_time":"2026-08-01T19:59:35+00:00","author":"admin NU","twitter_card":"summary_large_image","twitter_misc":{"\u4f5c\u8005":"admin NU","\u9884\u8ba1\u9605\u8bfb\u65f6\u95f4":"3 \u5206"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/#article","isPartOf":{"@id":"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/"},"author":{"name":"admin NU","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c"},"headline":"AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs","datePublished":"2026-08-01T19:59:35+00:00","mainEntityOfPage":{"@id":"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/"},"wordCount":667,"commentCount":0,"publisher":{"@id":"https:\/\/yousum.gpucore.co\/#organization"},"image":{"@id":"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/#primaryimage"},"thumbnailUrl":"https:\/\/youzum.net\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-WnMGN8.webp","articleSection":["AI","Committee","News","Uncategorized"],"inLanguage":"zh-Hans","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/","url":"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/","name":"AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs - YouZum","isPartOf":{"@id":"https:\/\/yousum.gpucore.co\/#website"},"primaryImageOfPage":{"@id":"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/#primaryimage"},"image":{"@id":"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/#primaryimage"},"thumbnailUrl":"https:\/\/youzum.net\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-WnMGN8.webp","datePublished":"2026-08-01T19:59:35+00:00","description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","breadcrumb":{"@id":"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/#breadcrumb"},"inLanguage":"zh-Hans","potentialAction":[{"@type":"ReadAction","target":["https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/"]}]},{"@type":"ImageObject","inLanguage":"zh-Hans","@id":"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/#primaryimage","url":"https:\/\/youzum.net\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-WnMGN8.webp","contentUrl":"https:\/\/youzum.net\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-WnMGN8.webp","width":2162,"height":1152},{"@type":"BreadcrumbList","@id":"https:\/\/youzum.net\/amd-releases-instella-moe-16b-a3b-a-fully-open-mixture-of-experts-llm-with-2-8b-active-parameters-trained-on-instinct-gpus\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/youzum.net\/"},{"@type":"ListItem","position":2,"name":"AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs"}]},{"@type":"WebSite","@id":"https:\/\/yousum.gpucore.co\/#website","url":"https:\/\/yousum.gpucore.co\/","name":"YouSum","description":"","publisher":{"@id":"https:\/\/yousum.gpucore.co\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/yousum.gpucore.co\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"zh-Hans"},{"@type":"Organization","@id":"https:\/\/yousum.gpucore.co\/#organization","name":"Drone Association Thailand","url":"https:\/\/yousum.gpucore.co\/","logo":{"@type":"ImageObject","inLanguage":"zh-Hans","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/","url":"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png","contentUrl":"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png","width":300,"height":300,"caption":"Drone Association Thailand"},"image":{"@id":"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/DroneAssociationTH\/"]},{"@type":"Person","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c","name":"admin NU","image":{"@type":"ImageObject","inLanguage":"zh-Hans","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/image\/","url":"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png","contentUrl":"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png","caption":"admin NU"},"url":"https:\/\/youzum.net\/zh\/members\/adminnu\/"}]}},"rttpg_featured_image_url":{"full":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-WnMGN8.webp",2162,1152,false],"landscape":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-WnMGN8.webp",2162,1152,false],"portraits":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-WnMGN8.webp",2162,1152,false],"thumbnail":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-WnMGN8-150x150.webp",150,150,true],"medium":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-WnMGN8-300x160.webp",300,160,true],"large":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-WnMGN8-1024x546.webp",1024,546,true],"1536x1536":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-WnMGN8-1536x818.webp",1536,818,true],"2048x2048":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-WnMGN8-2048x1091.webp",2048,1091,true],"trp-custom-language-flag":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-WnMGN8-18x10.webp",18,10,true],"woocommerce_thumbnail":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-WnMGN8-300x300.webp",300,300,true],"woocommerce_single":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-WnMGN8-600x320.webp",600,320,true],"woocommerce_gallery_thumbnail":["https:\/\/youzum.net\/wp-content\/uploads\/2026\/08\/Screenshot-2026-08-01-at-11.57.14-AM-1-WnMGN8-100x100.webp",100,100,true]},"rttpg_author":{"display_name":"admin NU","author_link":"https:\/\/youzum.net\/zh\/members\/adminnu\/"},"rttpg_comment":0,"rttpg_category":"<a href=\"https:\/\/youzum.net\/zh\/category\/ai-club\/\" rel=\"category tag\">AI<\/a> <a href=\"https:\/\/youzum.net\/zh\/category\/committee\/\" rel=\"category tag\">Committee<\/a> <a href=\"https:\/\/youzum.net\/zh\/category\/news\/\" rel=\"category tag\">News<\/a> <a href=\"https:\/\/youzum.net\/zh\/category\/uncategorized\/\" rel=\"category tag\">Uncategorized<\/a>","rttpg_excerpt":"AMD released Instella-MoE-16B-A3B, a fully open Mixture-of-Experts language model trained from scratch on Instinct MI300X and MI325X GPUs. The model holds 16B total parameters but activates only 2.8B per token. AMD is publishing weights from every training stage, along with data mixtures, training configs, and inference code. Two systems-level choices carry the release: Gated Multi-head&hellip;","_links":{"self":[{"href":"https:\/\/youzum.net\/zh\/wp-json\/wp\/v2\/posts\/108381","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/youzum.net\/zh\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/youzum.net\/zh\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/youzum.net\/zh\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/youzum.net\/zh\/wp-json\/wp\/v2\/comments?post=108381"}],"version-history":[{"count":0,"href":"https:\/\/youzum.net\/zh\/wp-json\/wp\/v2\/posts\/108381\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/youzum.net\/zh\/wp-json\/wp\/v2\/media\/108382"}],"wp:attachment":[{"href":"https:\/\/youzum.net\/zh\/wp-json\/wp\/v2\/media?parent=108381"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/youzum.net\/zh\/wp-json\/wp\/v2\/categories?post=108381"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/youzum.net\/zh\/wp-json\/wp\/v2\/tags?post=108381"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}