{"id":113985,"date":"2026-08-27T00:45:34","date_gmt":"2026-08-27T00:45:34","guid":{"rendered":"https:\/\/youzum.net\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/"},"modified":"2026-08-27T00:45:34","modified_gmt":"2026-08-27T00:45:34","slug":"z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context","status":"publish","type":"post","link":"https:\/\/youzum.net\/zh\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/","title":{"rendered":"Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context"},"content":{"rendered":"<p class=\"wp-block-paragraph\">Z.ai has released <a href=\"https:\/\/z.ai\/blog\/glm-5.3-flash\">GLM-5.3-Flash<\/a>, the first natively multimodal model in the GLM-5 series and the cheapest capable coding model the lab has shipped. It is a mixture-of-experts model with <strong>320B total parameters and 18B active per token<\/strong>, a <strong>1,048,576-token context window<\/strong>, and image and video input \u2014 released under an <strong>MIT license<\/strong> with weights on <a href=\"https:\/\/huggingface.co\/zai-org\/GLM-5.3-Flash\">Hugging Face<\/a>. According to Z.ai reports, it beats GLM-5.2 across benchmarks and real workloads at roughly one-tenth the price, while landing within half a point of <strong>Claude Opus 4.8<\/strong> on its internal coding benchmark. The model spent its first week running anonymously as <em>\u201cOx Alpha\u201d<\/em> on OpenCode and OpenRouter, served entirely on domestically produced Chinese AI chips.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Is it deployable?<\/strong><\/h2>\n<p class=\"wp-block-paragraph\"><strong>Yes<\/strong>, <strong>on two tracks.<\/strong> The weights are live on <a href=\"https:\/\/huggingface.co\/zai-org\/GLM-5.3-Flash\">Hugging Face<\/a> under an <strong>MIT license<\/strong>, and a hosted API is already priced and serving.<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Which companies can realistically self-host?<\/strong> Not everyone. The default FP8 checkpoint is roughly <a href=\"https:\/\/recipes.vllm.ai\/zai-org\/GLM-5.3-Flash\">306 GiB of weights before KV cache<\/a>, and the current vLLM path supports NVIDIA Hopper and newer only. That puts self-hosting in reach of mid-size and large orgs with at least an 8-GPU node (or a GB200 tray at TP4), plus AI-native startups renting GPU capacity. Everyone below that line consumes it as an API \u2014 where the economics, not the hardware, are the story.<\/li>\n<li><strong>Industries with immediate fit<\/strong>: software and devtools, IT\/BPO automation, financial services and insurance document operations, enterprise BI and back-office knowledge work, e-commerce and any team shipping UI at volume.<\/li>\n<li><strong>Applications:<\/strong> repo-scale coding agents, terminal and browser\/computer-use agents, million-token log and contract analysis, UI regression checking from screenshots, and spreadsheet\/deck\/dashboard reasoning that would otherwise need an OCR-to-text pipeline.<\/li>\n<\/ul>\n<div>\n<\/div>\n<p class=\"wp-block-paragraph\">\n<h2 class=\"wp-block-heading\"><strong>The architecture is where the efficiency comes from<\/strong><\/h2>\n<\/p><p class=\"wp-block-paragraph\">GLM-5.3-Flash starts from a newly trained base model on a <strong>30T-token multimodal corpus<\/strong>. <strong>Three changes are worth knowing:<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Hybrid attention<\/strong>: For the first time in the GLM series, Z.ai combines linear and sparse attention. Per the <a href=\"https:\/\/recipes.vllm.ai\/zai-org\/GLM-5.3-Flash\">vLLM recipe<\/a>, the 45-layer language model interleaves <strong>KDA linear-attention layers with NoPE sparse MLA layers<\/strong>, routes each token through <strong>8 of 288 experts<\/strong>, and ships <strong>native FP8 weights<\/strong> plus one MTP draft layer. Linear attention handles local dependency; sparse attention retrieves the globally relevant context.<\/li>\n<li><strong>IndexPool<\/strong>: At million-token context, retrieval itself becomes the bottleneck. IndexPool compresses groups of indexer key vectors through weighted pooling to hold down latency and memory. Z.ai reports <strong>~3\u00d7 less attention compute and a 4.4\u00d7 smaller KV cache<\/strong> versus GLM-5.3.<\/li>\n<li><strong>mHC<\/strong>: The model adopts <strong>Manifold-Constrained Hyper-Connections<\/strong> to improve scaling efficiency. Against GLM-4.5, at similar total parameter count, GLM-5.3-Flash roughly halves both activated parameters and layer count.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\"><strong>Benchmarks<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">Most numbers mentioned in the table below are Z.ai-reported and the harnesses differ per test \u2014 the model card\u2019s <a href=\"https:\/\/huggingface.co\/zai-org\/GLM-5.3-Flash\">footnotes<\/a> specify temperature, context limits and judge models per benchmark, so treat cross-model comparisons as setup-dependent.<\/p>\n<figure class=\"wp-block-table\">\n<table class=\"has-fixed-layout\">\n<thead>\n<tr>\n<th>Benchmark<\/th>\n<th>GLM-5.3-Flash<\/th>\n<th>Reference<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Terminal-Bench 2.1<\/td>\n<td><strong>84.3<\/strong><\/td>\n<td>Opus 4.8: 85.0 \u00b7 GPT-5.6 Terra: 87.4<\/td>\n<\/tr>\n<tr>\n<td>DeepSWE v1.1<\/td>\n<td><strong>63.4<\/strong><\/td>\n<td>GLM-5.2: 46.2<\/td>\n<\/tr>\n<tr>\n<td>AutomationBench<\/td>\n<td><strong>48.8<\/strong><\/td>\n<td>GLM-5.2: 26.2<\/td>\n<\/tr>\n<tr>\n<td>HLE<\/td>\n<td><strong>55.3<\/strong><\/td>\n<td>\u2014<\/td>\n<\/tr>\n<tr>\n<td>OfficeQA Pro<\/td>\n<td><strong>62.4<\/strong><\/td>\n<td>ahead of Opus 4.8<\/td>\n<\/tr>\n<tr>\n<td>Z.ai Code Bench v1.0 (max)<\/td>\n<td><strong>29.0<\/strong><\/td>\n<td>Opus 4.8: 29.5<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/figure>\n<p class=\"wp-block-paragraph\">Independently, <a href=\"https:\/\/artificialanalysis.ai\/models\/glm-5-3-flash\">Artificial Analysis<\/a> scores it <strong>57 on the Intelligence Index<\/strong>, with <strong>48.7 output tokens\/sec<\/strong> and <strong>1.52s TTFT<\/strong> on Z.ai\u2019s API \u2014 strong intelligence-per-dollar, but slow and verbose. Vision is the weak flank: it trails Gemini 3.7 Flash on BabyVision and MVbench.<\/p>\n<h2 class=\"wp-block-heading\"><strong>The serving story is the underreported part<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">Z.ai states the entire Ox Alpha preview ran on <strong>domestically produced Chinese AI chips<\/strong>, using a custom SGLang-based engine that disaggregates encoding, prefill and decoding, and reports a <strong>3\u00d7 end-to-end serving improvement<\/strong> across tens of thousands of accelerators.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Pricing and access<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">Standard API pricing is <strong>$0.15\/M input, $0.03\/M cached input, $0.50\/M output<\/strong>. Z.ai reports a score of 57 on Artificial Analysis Intelligence Index v4.1.1 at <strong>$0.045 per task<\/strong> on the discounted tier. The model is live for <strong>all <a href=\"https:\/\/zcode.z.ai\/en\">GLM Coding Plan<\/a> tiers<\/strong> \u2014 Lite ($18\/mo), Pro ($80), Max ($168) \u2014 at <strong>3\u00d7 the usable quota of GLM-5.3<\/strong>, and its multimodal capabilities surface in <strong>ZCode<\/strong> through Browser Use and Computer Use. Local serving is supported on <strong>SGLang, vLLM, TokenSpeed and KTransformers<\/strong>.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Key Takeaways<\/strong><\/h2>\n<ul class=\"wp-block-list\">\n<li>320B-A18B natively multimodal MoE, 1M context, MIT-licensed weights on Hugging Face.<\/li>\n<li>Hybrid KDA linear + NoPE sparse MLA attention: ~3\u00d7 less attention compute, 4.4\u00d7 smaller KV cache.<\/li>\n<li>84.3 Terminal-Bench 2.1 and 63.4 DeepSWE v1.1 \u2014 near Opus 4.8, well past GLM-5.2.<\/li>\n<li>$0.15\/$0.50 per M tokens; 3\u00d7 GLM-5.3 quota for every GLM Coding Plan tier.<\/li>\n<li>Self-hosting needs ~306 GiB FP8 weights on Hopper-or-newer; everyone else uses the API.<\/li>\n<\/ul>\n<p class=\"wp-block-paragraph\">\n<hr class=\"wp-block-separator has-alpha-channel-opacity\" \/>\n<\/p><p class=\"wp-block-paragraph\">Check out the\u00a0<strong><a href=\"https:\/\/huggingface.co\/zai-org\/GLM-5.3-Flash\" target=\"_blank\" rel=\"noreferrer noopener\">Model Weights<\/a><\/strong> and <a href=\"https:\/\/z.ai\/blog\/glm-5.3-flash\" target=\"_blank\" rel=\"noreferrer noopener\"><strong>Blog<\/strong><\/a><strong>.<\/strong>\u00a0Also,\u00a0feel free to follow us on\u00a0<strong><a href=\"https:\/\/x.com\/intent\/follow?screen_name=marktechpost\" target=\"_blank\" rel=\"noopener\"><mark>Twitter<\/mark><\/a><\/strong>\u00a0and don\u2019t forget to join our\u00a0<strong><a href=\"https:\/\/www.reddit.com\/r\/machinelearningnews\/\" target=\"_blank\" rel=\"noopener\">150k+ML SubReddit<\/a><\/strong>\u00a0and Subscribe to\u00a0<strong><a href=\"https:\/\/magic.beehiiv.com\/v1\/f5e63dd4-5653-4f09-83e2-321a8b1ba526?email=%7B%7Bemail%7D%7D\" target=\"_blank\" rel=\"noopener\">our Newsletter<\/a><\/strong>. Wait! are you on telegram?\u00a0<strong><a href=\"https:\/\/t.me\/machinelearningresearchnews\" target=\"_blank\" rel=\"noopener\">now you can join us on telegram as well.<\/a><\/strong><\/p>\n<p class=\"wp-block-paragraph\">Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?\u00a0<strong><a href=\"https:\/\/forms.gle\/wbash1wF6efRj8G58\" target=\"_blank\" rel=\"noopener\"><mark>Connect with us<\/mark><\/a><\/strong><\/p>\n<p>The post <a href=\"https:\/\/www.marktechpost.com\/2026\/08\/26\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/\">Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context<\/a> appeared first on <a href=\"https:\/\/www.marktechpost.com\/\">MarkTechPost<\/a>.<\/p>","protected":false},"excerpt":{"rendered":"<p>Z.ai has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series and the cheapest capable coding model the lab has shipped. It is a mixture-of-experts model with 320B total parameters and 18B active per token, a 1,048,576-token context window, and image and video input \u2014 released under an MIT license with weights on Hugging Face. According to Z.ai reports, it beats GLM-5.2 across benchmarks and real workloads at roughly one-tenth the price, while landing within half a point of Claude Opus 4.8 on its internal coding benchmark. The model spent its first week running anonymously as \u201cOx Alpha\u201d on OpenCode and OpenRouter, served entirely on domestically produced Chinese AI chips. Is it deployable? Yes, on two tracks. The weights are live on Hugging Face under an MIT license, and a hosted API is already priced and serving. Which companies can realistically self-host? Not everyone. The default FP8 checkpoint is roughly 306 GiB of weights before KV cache, and the current vLLM path supports NVIDIA Hopper and newer only. That puts self-hosting in reach of mid-size and large orgs with at least an 8-GPU node (or a GB200 tray at TP4), plus AI-native startups renting GPU capacity. Everyone below that line consumes it as an API \u2014 where the economics, not the hardware, are the story. Industries with immediate fit: software and devtools, IT\/BPO automation, financial services and insurance document operations, enterprise BI and back-office knowledge work, e-commerce and any team shipping UI at volume. Applications: repo-scale coding agents, terminal and browser\/computer-use agents, million-token log and contract analysis, UI regression checking from screenshots, and spreadsheet\/deck\/dashboard reasoning that would otherwise need an OCR-to-text pipeline. The architecture is where the efficiency comes from GLM-5.3-Flash starts from a newly trained base model on a 30T-token multimodal corpus. Three changes are worth knowing: Hybrid attention: For the first time in the GLM series, Z.ai combines linear and sparse attention. Per the vLLM recipe, the 45-layer language model interleaves KDA linear-attention layers with NoPE sparse MLA layers, routes each token through 8 of 288 experts, and ships native FP8 weights plus one MTP draft layer. Linear attention handles local dependency; sparse attention retrieves the globally relevant context. IndexPool: At million-token context, retrieval itself becomes the bottleneck. IndexPool compresses groups of indexer key vectors through weighted pooling to hold down latency and memory. Z.ai reports ~3\u00d7 less attention compute and a 4.4\u00d7 smaller KV cache versus GLM-5.3. mHC: The model adopts Manifold-Constrained Hyper-Connections to improve scaling efficiency. Against GLM-4.5, at similar total parameter count, GLM-5.3-Flash roughly halves both activated parameters and layer count. Benchmarks Most numbers mentioned in the table below are Z.ai-reported and the harnesses differ per test \u2014 the model card\u2019s footnotes specify temperature, context limits and judge models per benchmark, so treat cross-model comparisons as setup-dependent. Benchmark GLM-5.3-Flash Reference Terminal-Bench 2.1 84.3 Opus 4.8: 85.0 \u00b7 GPT-5.6 Terra: 87.4 DeepSWE v1.1 63.4 GLM-5.2: 46.2 AutomationBench 48.8 GLM-5.2: 26.2 HLE 55.3 \u2014 OfficeQA Pro 62.4 ahead of Opus 4.8 Z.ai Code Bench v1.0 (max) 29.0 Opus 4.8: 29.5 Independently, Artificial Analysis scores it 57 on the Intelligence Index, with 48.7 output tokens\/sec and 1.52s TTFT on Z.ai\u2019s API \u2014 strong intelligence-per-dollar, but slow and verbose. Vision is the weak flank: it trails Gemini 3.7 Flash on BabyVision and MVbench. The serving story is the underreported part Z.ai states the entire Ox Alpha preview ran on domestically produced Chinese AI chips, using a custom SGLang-based engine that disaggregates encoding, prefill and decoding, and reports a 3\u00d7 end-to-end serving improvement across tens of thousands of accelerators. Pricing and access Standard API pricing is $0.15\/M input, $0.03\/M cached input, $0.50\/M output. Z.ai reports a score of 57 on Artificial Analysis Intelligence Index v4.1.1 at $0.045 per task on the discounted tier. The model is live for all GLM Coding Plan tiers \u2014 Lite ($18\/mo), Pro ($80), Max ($168) \u2014 at 3\u00d7 the usable quota of GLM-5.3, and its multimodal capabilities surface in ZCode through Browser Use and Computer Use. Local serving is supported on SGLang, vLLM, TokenSpeed and KTransformers. Key Takeaways 320B-A18B natively multimodal MoE, 1M context, MIT-licensed weights on Hugging Face. Hybrid KDA linear + NoPE sparse MLA attention: ~3\u00d7 less attention compute, 4.4\u00d7 smaller KV cache. 84.3 Terminal-Bench 2.1 and 63.4 DeepSWE v1.1 \u2014 near Opus 4.8, well past GLM-5.2. $0.15\/$0.50 per M tokens; 3\u00d7 GLM-5.3 quota for every GLM Coding Plan tier. Self-hosting needs ~306 GiB FP8 weights on Hopper-or-newer; everyone else uses the API. Check out the\u00a0Model Weights and Blog.\u00a0Also,\u00a0feel free to follow us on\u00a0Twitter\u00a0and don\u2019t forget to join our\u00a0150k+ML SubReddit\u00a0and Subscribe to\u00a0our Newsletter. Wait! are you on telegram?\u00a0now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?\u00a0Connect with us The post Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context appeared first on MarkTechPost.<\/p>","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"pmpro_default_level":"","site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"_pvb_checkbox_block_on_post":false,"footnotes":""},"categories":[52,5,7,1],"tags":[],"class_list":["post-113985","post","type-post","status-publish","format-standard","hentry","category-ai-club","category-committee","category-news","category-uncategorized","pmpro-has-access"],"acf":[],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v25.3 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context - YouZum<\/title>\n<meta name=\"description\" content=\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/youzum.net\/zh\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/\" \/>\n<meta property=\"og:locale\" content=\"zh_CN\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context - YouZum\" \/>\n<meta property=\"og:description\" content=\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\" \/>\n<meta property=\"og:url\" content=\"https:\/\/youzum.net\/zh\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/\" \/>\n<meta property=\"og:site_name\" content=\"YouZum\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/DroneAssociationTH\/\" \/>\n<meta property=\"article:published_time\" content=\"2026-08-27T00:45:34+00:00\" \/>\n<meta name=\"author\" content=\"admin NU\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"\u4f5c\u8005\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin NU\" \/>\n\t<meta name=\"twitter:label2\" content=\"\u9884\u8ba1\u9605\u8bfb\u65f6\u95f4\" \/>\n\t<meta name=\"twitter:data2\" content=\"4 \u5206\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/youzum.net\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/#article\",\"isPartOf\":{\"@id\":\"https:\/\/youzum.net\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/\"},\"author\":{\"name\":\"admin NU\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c\"},\"headline\":\"Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context\",\"datePublished\":\"2026-08-27T00:45:34+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/youzum.net\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/\"},\"wordCount\":817,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\"},\"articleSection\":[\"AI\",\"Committee\",\"News\",\"Uncategorized\"],\"inLanguage\":\"zh-Hans\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/youzum.net\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/youzum.net\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/\",\"url\":\"https:\/\/youzum.net\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/\",\"name\":\"Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context - YouZum\",\"isPartOf\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#website\"},\"datePublished\":\"2026-08-27T00:45:34+00:00\",\"description\":\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\",\"breadcrumb\":{\"@id\":\"https:\/\/youzum.net\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/#breadcrumb\"},\"inLanguage\":\"zh-Hans\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/youzum.net\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/\"]}]},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/youzum.net\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/youzum.net\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/yousum.gpucore.co\/#website\",\"url\":\"https:\/\/yousum.gpucore.co\/\",\"name\":\"YouSum\",\"description\":\"\",\"publisher\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/yousum.gpucore.co\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"zh-Hans\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\",\"name\":\"Drone Association Thailand\",\"url\":\"https:\/\/yousum.gpucore.co\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"zh-Hans\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png\",\"contentUrl\":\"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png\",\"width\":300,\"height\":300,\"caption\":\"Drone Association Thailand\"},\"image\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/\"},\"sameAs\":[\"https:\/\/www.facebook.com\/DroneAssociationTH\/\"]},{\"@type\":\"Person\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c\",\"name\":\"admin NU\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"zh-Hans\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png\",\"contentUrl\":\"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png\",\"caption\":\"admin NU\"},\"url\":\"https:\/\/youzum.net\/zh\/members\/adminnu\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context - YouZum","description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/youzum.net\/zh\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/","og_locale":"zh_CN","og_type":"article","og_title":"Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context - YouZum","og_description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","og_url":"https:\/\/youzum.net\/zh\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/","og_site_name":"YouZum","article_publisher":"https:\/\/www.facebook.com\/DroneAssociationTH\/","article_published_time":"2026-08-27T00:45:34+00:00","author":"admin NU","twitter_card":"summary_large_image","twitter_misc":{"\u4f5c\u8005":"admin NU","\u9884\u8ba1\u9605\u8bfb\u65f6\u95f4":"4 \u5206"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/youzum.net\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/#article","isPartOf":{"@id":"https:\/\/youzum.net\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/"},"author":{"name":"admin NU","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c"},"headline":"Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context","datePublished":"2026-08-27T00:45:34+00:00","mainEntityOfPage":{"@id":"https:\/\/youzum.net\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/"},"wordCount":817,"commentCount":0,"publisher":{"@id":"https:\/\/yousum.gpucore.co\/#organization"},"articleSection":["AI","Committee","News","Uncategorized"],"inLanguage":"zh-Hans","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/youzum.net\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/youzum.net\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/","url":"https:\/\/youzum.net\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/","name":"Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context - YouZum","isPartOf":{"@id":"https:\/\/yousum.gpucore.co\/#website"},"datePublished":"2026-08-27T00:45:34+00:00","description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","breadcrumb":{"@id":"https:\/\/youzum.net\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/#breadcrumb"},"inLanguage":"zh-Hans","potentialAction":[{"@type":"ReadAction","target":["https:\/\/youzum.net\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/"]}]},{"@type":"BreadcrumbList","@id":"https:\/\/youzum.net\/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/youzum.net\/"},{"@type":"ListItem","position":2,"name":"Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context"}]},{"@type":"WebSite","@id":"https:\/\/yousum.gpucore.co\/#website","url":"https:\/\/yousum.gpucore.co\/","name":"YouSum","description":"","publisher":{"@id":"https:\/\/yousum.gpucore.co\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/yousum.gpucore.co\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"zh-Hans"},{"@type":"Organization","@id":"https:\/\/yousum.gpucore.co\/#organization","name":"Drone Association Thailand","url":"https:\/\/yousum.gpucore.co\/","logo":{"@type":"ImageObject","inLanguage":"zh-Hans","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/","url":"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png","contentUrl":"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png","width":300,"height":300,"caption":"Drone Association Thailand"},"image":{"@id":"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/DroneAssociationTH\/"]},{"@type":"Person","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c","name":"admin NU","image":{"@type":"ImageObject","inLanguage":"zh-Hans","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/image\/","url":"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png","contentUrl":"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png","caption":"admin NU"},"url":"https:\/\/youzum.net\/zh\/members\/adminnu\/"}]}},"rttpg_featured_image_url":null,"rttpg_author":{"display_name":"admin NU","author_link":"https:\/\/youzum.net\/zh\/members\/adminnu\/"},"rttpg_comment":0,"rttpg_category":"<a href=\"https:\/\/youzum.net\/zh\/category\/ai-club\/\" rel=\"category tag\">AI<\/a> <a href=\"https:\/\/youzum.net\/zh\/category\/committee\/\" rel=\"category tag\">Committee<\/a> <a href=\"https:\/\/youzum.net\/zh\/category\/news\/\" rel=\"category tag\">News<\/a> <a href=\"https:\/\/youzum.net\/zh\/category\/uncategorized\/\" rel=\"category tag\">Uncategorized<\/a>","rttpg_excerpt":"Z.ai has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series and the cheapest capable coding model the lab has shipped. It is a mixture-of-experts model with 320B total parameters and 18B active per token, a 1,048,576-token context window, and image and video input \u2014 released under an MIT license with weights on&hellip;","_links":{"self":[{"href":"https:\/\/youzum.net\/zh\/wp-json\/wp\/v2\/posts\/113985","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/youzum.net\/zh\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/youzum.net\/zh\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/youzum.net\/zh\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/youzum.net\/zh\/wp-json\/wp\/v2\/comments?post=113985"}],"version-history":[{"count":0,"href":"https:\/\/youzum.net\/zh\/wp-json\/wp\/v2\/posts\/113985\/revisions"}],"wp:attachment":[{"href":"https:\/\/youzum.net\/zh\/wp-json\/wp\/v2\/media?parent=113985"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/youzum.net\/zh\/wp-json\/wp\/v2\/categories?post=113985"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/youzum.net\/zh\/wp-json\/wp\/v2\/tags?post=113985"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}