{"id":120467,"date":"2026-09-27T02:26:49","date_gmt":"2026-09-27T02:26:49","guid":{"rendered":"https:\/\/youzum.net\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/"},"modified":"2026-09-27T02:26:49","modified_gmt":"2026-09-27T02:26:49","slug":"sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english","status":"publish","type":"post","link":"https:\/\/youzum.net\/es\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/","title":{"rendered":"Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English"},"content":{"rendered":"<p class=\"wp-block-paragraph\"><a href=\"https:\/\/www.sarvam.ai\/\">Sarvam AI<\/a> has released <a href=\"https:\/\/www.sarvam.ai\/blogs\/introducing-saaras-v4\">Saaras V4<\/a>, the newest generation of its speech recognition model. It covers all 22 scheduled Indian languages plus English, now including global English accents. Sarvam reports state-of-the-art accuracy across all 22 languages. <\/p>\n<p class=\"wp-block-paragraph\"><strong>Is it deployable?<\/strong> Yes, through Sarvam\u2019s API today, using <code>model=\"saaras:v4\"<\/code>. Weights are not public, and Sarvam\u2019s <a href=\"https:\/\/docs.sarvam.ai\/api\/self-hosted\/sagemaker\/deploy-saaras\">SageMaker self-hosting docs<\/a> currently cover Saaras v3 only.<\/p>\n<h2 class=\"wp-block-heading\"><strong>What is Inside Saaras V4<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">Saaras V4 is an encoder-decoder system. An audio encoder converts the waveform into embeddings that carry phonetic and acoustic detail. A temporal-downsampling adapter then shortens that sequence and projects it into the language model\u2019s embedding space. This keeps long recordings inside the decoder\u2019s context budget.<\/p>\n<p class=\"wp-block-paragraph\">The decoder is Sarvam-3B, a 3B-parameter hybrid state-space language model trained from scratch in-house. It reads the audio features alongside a text prompt. It then emits the transcript autoregressively, feeding each token back as input for the next.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Benchmark Results<\/strong><\/h2>\n<ul class=\"wp-block-list\">\n<li><strong>English<\/strong>: Sarvam evaluated 7 English datasets. Six come from Hugging Face\u2019s <a href=\"https:\/\/huggingface.co\/datasets\/hf-audio\/open-asr-leaderboard\">Open ASR Leaderboard<\/a>: AMI, GigaSpeech, LibriSpeech clean, LibriSpeech other, SPGISpeech and VoxPopuli. The seventh is AI4Bharat\u2019s Indian-accented <a href=\"https:\/\/huggingface.co\/datasets\/ai4bharat\/Svarah\">Svarah<\/a>. Scoring follows the leaderboard\u2019s <a href=\"https:\/\/github.com\/huggingface\/open_asr_leaderboard\">normalization code<\/a>. Saaras V4 posts the lowest average WER among the models Sarvam benchmarked.<\/li>\n<li><strong>Indic<\/strong>: On <a href=\"https:\/\/github.com\/AI4Bharat\/vistaar\">Vistaar<\/a>, Sarvam reports results across 10 Indian languages using both WER and <a href=\"https:\/\/www.sarvam.ai\/blogs\/evaluating-indian-language-asr\">LLM-WER<\/a>. LLM-WER adds a semantic check. It separates real meaning errors from harmless spelling or formatting variants common in Indic scripts.<\/li>\n<li><strong>Noisy audio<\/strong>: On Kathbath Noisy, measured with LLM-WER, Sarvam says Saaras V4\u2019s error rate is under half that of Deepgram Nova-3 and GPT-4o Transcribe. The set includes compressed, clipped and background-heavy recordings.<\/li>\n<li><strong>Language ID<\/strong>: On verified IndicVoices utterances, language identification error is 2.9% across the top 10 Indian languages. It is 5.22% across all 22.<\/li>\n<\/ul>\n<p class=\"wp-block-paragraph\"><em>It is important to note that <strong>all numbers above are vendor-reported. Independent reproduction has not been published yet.<\/strong><\/em><\/p>\n<h2 class=\"wp-block-heading\"><strong>5 Output Modes From 1 Model<\/strong><\/h2>\n<p class=\"wp-block-paragraph\"><strong>The same audio can return 5 representations, selected through the <code>mode<\/code> parameter:<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li><strong>transcribe<\/strong> (default): native script with numbers and dates normalized.<\/li>\n<li><strong>verbatim<\/strong>: every word as spoken, fillers and spoken numbers kept.<\/li>\n<li><strong>codemix<\/strong>: native script, with English words left in English.<\/li>\n<li><strong>translit<\/strong>: the full utterance in Latin script.<\/li>\n<li><strong>translate<\/strong>: an English translation with numbers normalized.<\/li>\n<\/ul>\n<p class=\"wp-block-paragraph\">Sarvam\u2019s argument is simple. Handling these inside the model removes post-processing steps that can compound errors.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Keyterm Prompting<\/strong><\/h2>\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/docs.sarvam.ai\/api\/api-guides-tutorials\/speech-to-text\/how-to\/keyterms\">Keyterm prompting<\/a> is new in V4 and works only with <code>saaras:v4<\/code>. You pass a JSON list under <code>keyterms<\/code>, with up to 50 terms of 64 characters each. Keyterms bias recognition; they do not guarantee output. Use <code>codemix<\/code> mode when a brand such as PhonePe must stay in Latin script.<\/p>\n<p class=\"wp-block-paragraph\">On <a href=\"https:\/\/huggingface.co\/datasets\/ai4bharat\/IndicContextEval\">IndicContextEval<\/a> (<a href=\"https:\/\/arxiv.org\/pdf\/2606.19157\">paper<\/a>, Interspeech 2026), Saaras V4 reports 16.03% WER in the L5 keyword-prompting setting. Sarvam says that is the lowest score on the benchmark.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Streaming, Long Audio and Pricing<\/strong><\/h2>\n<ul class=\"wp-block-list\">\n<li><strong>Streaming:<\/strong> <a href=\"https:\/\/docs.sarvam.ai\/api\/api-guides-tutorials\/speech-to-text\/realtime-streaming\">WebSocket<\/a> with partial results and time to first token below 150 ms.<\/li>\n<li><strong>REST:<\/strong> <a href=\"https:\/\/docs.sarvam.ai\/api\/api-guides-tutorials\/speech-to-text\/rest-api\">synchronous<\/a> transcription for clips up to 30 seconds.<\/li>\n<li><strong>Batch:<\/strong> <a href=\"https:\/\/docs.sarvam.ai\/api\/api-guides-tutorials\/speech-to-text\/batch-api\">asynchronous<\/a> jobs up to 2 hours per file, with optional speaker diarization.<\/li>\n<li><strong>SDKs:<\/strong> Python 3.9+ and Node.js 18+, plus <a href=\"https:\/\/docs.livekit.io\/agents\/models\/stt\/sarvam\/\">LiveKit Agents<\/a>, <a href=\"https:\/\/docs.pipecat.ai\/api-reference\/server\/services\/stt\/sarvam\">Pipecat<\/a> and <a href=\"https:\/\/docs.sarvam.ai\/api\/integration\/vercel-ai-sdk\">Vercel AI SDK<\/a> integrations.<\/li>\n<li><strong>Price:<\/strong> Sarvam lists speech-to-text at <a href=\"https:\/\/www.sarvam.ai\/api-pricing\">\u20b930 per hour<\/a> for real-time, streaming and batch, and \u20b945 per hour with diarization.<\/li>\n<\/ul>\n<p class=\"wp-block-paragraph\">Saaras v3 stays the <a href=\"https:\/\/docs.sarvam.ai\/api\/getting-started\/models\/saaras\">default model<\/a>. V4 uses the same request shape, so switching is a 1-line change.<\/p>\n<div>\n<\/div>\n<p class=\"wp-block-paragraph\">\n<h2 class=\"wp-block-heading\"><strong>Saaras V4 vs Closest Competitors<\/strong><\/h2>\n<\/p><p class=\"wp-block-paragraph\">These are the 3 systems Sarvam benchmarked against. Figures come from each vendor\u2019s public docs and pricing pages, checked on September 26, 2026.<\/p>\n<figure class=\"wp-block-table\">\n<table class=\"has-fixed-layout\">\n<thead>\n<tr>\n<th>Feature<\/th>\n<th>Sarvam Saaras V4<\/th>\n<th>Deepgram Nova-3<\/th>\n<th>ElevenLabs Scribe v2<\/th>\n<th>OpenAI GPT-4o Transcribe<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Indian scheduled languages (of 22)<\/td>\n<td><a href=\"https:\/\/docs.sarvam.ai\/api\/getting-started\/models\/saaras\">22<\/a><\/td>\n<td><a href=\"https:\/\/developers.deepgram.com\/docs\/models-languages-overview\">11<\/a><\/td>\n<td><a href=\"https:\/\/elevenlabs.io\/docs\/overview\/capabilities\/speech-to-text\">14<\/a><\/td>\n<td><a href=\"https:\/\/developers.openai.com\/api\/docs\/guides\/speech-to-text\">Not listed per language<\/a><\/td>\n<\/tr>\n<tr>\n<td>Total languages<\/td>\n<td>23 (22 Indian + English)<\/td>\n<td>45+<\/td>\n<td>90+<\/td>\n<td>Multilingual<\/td>\n<\/tr>\n<tr>\n<td>Keyterm biasing<\/td>\n<td>Up to 50 terms<\/td>\n<td><a href=\"https:\/\/deepgram.com\/pricing\">Yes, paid add-on<\/a><\/td>\n<td>Up to 1,000 (batch), 50 (realtime), paid add-on<\/td>\n<td>Free-text <code>prompt<\/code><\/td>\n<\/tr>\n<tr>\n<td>Built-in output modes<\/td>\n<td>5 (transcribe, verbatim, codemix, translit, translate)<\/td>\n<td>Transcript plus Smart Formatting<\/td>\n<td>Verbatim or no_verbatim<\/td>\n<td>Transcript<\/td>\n<\/tr>\n<tr>\n<td>Real-time streaming<\/td>\n<td>WebSocket, under 150 ms TTFT (vendor claim)<\/td>\n<td>Yes (WebSocket)<\/td>\n<td>Scribe v2 Realtime, about 150 ms<\/td>\n<td>File streaming; live via Realtime API<\/td>\n<\/tr>\n<tr>\n<td>Speaker diarization<\/td>\n<td>Batch API<\/td>\n<td>Yes<\/td>\n<td>Up to 32 speakers<\/td>\n<td>Separate <code>gpt-4o-transcribe-diarize<\/code> model<\/td>\n<\/tr>\n<tr>\n<td>List price<\/td>\n<td><a href=\"https:\/\/www.sarvam.ai\/api-pricing\">\u20b930\/hour<\/a><\/td>\n<td><a href=\"https:\/\/deepgram.com\/pricing\">$0.0052\/min<\/a> (multilingual, pre-recorded)<\/td>\n<td><a href=\"https:\/\/elevenlabs.io\/pricing\/api\">$0.22\/hour<\/a> (batch)<\/td>\n<td><a href=\"https:\/\/developers.openai.com\/api\/docs\/pricing\">~$0.006\/min<\/a><\/td>\n<\/tr>\n<tr>\n<td>Self-hosting<\/td>\n<td>Not for V4 yet (v3 on SageMaker)<\/td>\n<td><a href=\"https:\/\/developers.deepgram.com\/docs\/self-hosted-introduction\">Yes<\/a><\/td>\n<td>Cloud API<\/td>\n<td>Cloud API<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/figure>\n<h2 class=\"wp-block-heading\"><strong>Key Takeaways<\/strong><\/h2>\n<ul class=\"wp-block-list\">\n<li>Saaras V4 covers all 22 Indian languages plus global English in 1 model.<\/li>\n<li>A 3B hybrid state-space decoder, trained from scratch, sits behind an audio encoder.<\/li>\n<li>Keyterm prompting accepts up to 50 terms and scored 16.03% WER on IndicContextEval L5.<\/li>\n<li>5 output modes and sub-150 ms streaming TTFT come from the same model.<\/li>\n<li>API-only today at \u20b930 per hour; self-hosting docs still cover v3.<\/li>\n<\/ul>\n<p class=\"wp-block-paragraph\">\n<hr class=\"wp-block-separator has-alpha-channel-opacity\" \/>\n<\/p><p class=\"wp-block-paragraph\">\n<\/p><p class=\"wp-block-paragraph\">Check out the <a href=\"https:\/\/www.sarvam.ai\/blogs\/introducing-saaras-v4\" target=\"_blank\" rel=\"noreferrer noopener\"><strong>Technical Details<\/strong><\/a>. All credit goes to the researcher of this project. Also,\u00a0feel free to follow us on\u00a0<strong><a href=\"https:\/\/x.com\/intent\/follow?screen_name=marktechpost\" target=\"_blank\" rel=\"noopener\"><mark>Twitter<\/mark><\/a><\/strong>\u00a0and don\u2019t forget to join our\u00a0<strong><a href=\"https:\/\/www.reddit.com\/r\/machinelearningnews\/\" target=\"_blank\" rel=\"noopener\">150k+ML SubReddit<\/a><\/strong>\u00a0and Subscribe to\u00a0<strong><a href=\"https:\/\/magic.beehiiv.com\/v1\/f5e63dd4-5653-4f09-83e2-321a8b1ba526?email=%7B%7Bemail%7D%7D\" target=\"_blank\" rel=\"noopener\">our Newsletter<\/a><\/strong>. Wait! are you on telegram?\u00a0<strong><a href=\"https:\/\/t.me\/machinelearningresearchnews\" target=\"_blank\" rel=\"noopener\">now you can join us on telegram as well.<\/a><\/strong><\/p>\n<p class=\"wp-block-paragraph\">Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?\u00a0<strong><a href=\"https:\/\/forms.gle\/MJjjVDPS7whH8Ngs6\" target=\"_blank\" rel=\"noopener\"><mark>Connect with us<\/mark><\/a><\/strong><\/p>\n<p>The post <a href=\"https:\/\/www.marktechpost.com\/2026\/09\/26\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/\">Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English<\/a> appeared first on <a href=\"https:\/\/www.marktechpost.com\/\">MarkTechPost<\/a>.<\/p>","protected":false},"excerpt":{"rendered":"<p>Sarvam AI has released Saaras V4, the newest generation of its speech recognition model. It covers all 22 scheduled Indian languages plus English, now including global English accents. Sarvam reports state-of-the-art accuracy across all 22 languages. Is it deployable? Yes, through Sarvam\u2019s API today, using model=&#8221;saaras:v4&#8243;. Weights are not public, and Sarvam\u2019s SageMaker self-hosting docs currently cover Saaras v3 only. What is Inside Saaras V4 Saaras V4 is an encoder-decoder system. An audio encoder converts the waveform into embeddings that carry phonetic and acoustic detail. A temporal-downsampling adapter then shortens that sequence and projects it into the language model\u2019s embedding space. This keeps long recordings inside the decoder\u2019s context budget. The decoder is Sarvam-3B, a 3B-parameter hybrid state-space language model trained from scratch in-house. It reads the audio features alongside a text prompt. It then emits the transcript autoregressively, feeding each token back as input for the next. Benchmark Results English: Sarvam evaluated 7 English datasets. Six come from Hugging Face\u2019s Open ASR Leaderboard: AMI, GigaSpeech, LibriSpeech clean, LibriSpeech other, SPGISpeech and VoxPopuli. The seventh is AI4Bharat\u2019s Indian-accented Svarah. Scoring follows the leaderboard\u2019s normalization code. Saaras V4 posts the lowest average WER among the models Sarvam benchmarked. Indic: On Vistaar, Sarvam reports results across 10 Indian languages using both WER and LLM-WER. LLM-WER adds a semantic check. It separates real meaning errors from harmless spelling or formatting variants common in Indic scripts. Noisy audio: On Kathbath Noisy, measured with LLM-WER, Sarvam says Saaras V4\u2019s error rate is under half that of Deepgram Nova-3 and GPT-4o Transcribe. The set includes compressed, clipped and background-heavy recordings. Language ID: On verified IndicVoices utterances, language identification error is 2.9% across the top 10 Indian languages. It is 5.22% across all 22. It is important to note that all numbers above are vendor-reported. Independent reproduction has not been published yet. 5 Output Modes From 1 Model The same audio can return 5 representations, selected through the mode parameter: transcribe (default): native script with numbers and dates normalized. verbatim: every word as spoken, fillers and spoken numbers kept. codemix: native script, with English words left in English. translit: the full utterance in Latin script. translate: an English translation with numbers normalized. Sarvam\u2019s argument is simple. Handling these inside the model removes post-processing steps that can compound errors. Keyterm Prompting Keyterm prompting is new in V4 and works only with saaras:v4. You pass a JSON list under keyterms, with up to 50 terms of 64 characters each. Keyterms bias recognition; they do not guarantee output. Use codemix mode when a brand such as PhonePe must stay in Latin script. On IndicContextEval (paper, Interspeech 2026), Saaras V4 reports 16.03% WER in the L5 keyword-prompting setting. Sarvam says that is the lowest score on the benchmark. Streaming, Long Audio and Pricing Streaming: WebSocket with partial results and time to first token below 150 ms. REST: synchronous transcription for clips up to 30 seconds. Batch: asynchronous jobs up to 2 hours per file, with optional speaker diarization. SDKs: Python 3.9+ and Node.js 18+, plus LiveKit Agents, Pipecat and Vercel AI SDK integrations. Price: Sarvam lists speech-to-text at \u20b930 per hour for real-time, streaming and batch, and \u20b945 per hour with diarization. Saaras v3 stays the default model. V4 uses the same request shape, so switching is a 1-line change. Saaras V4 vs Closest Competitors These are the 3 systems Sarvam benchmarked against. Figures come from each vendor\u2019s public docs and pricing pages, checked on September 26, 2026. Feature Sarvam Saaras V4 Deepgram Nova-3 ElevenLabs Scribe v2 OpenAI GPT-4o Transcribe Indian scheduled languages (of 22) 22 11 14 Not listed per language Total languages 23 (22 Indian + English) 45+ 90+ Multilingual Keyterm biasing Up to 50 terms Yes, paid add-on Up to 1,000 (batch), 50 (realtime), paid add-on Free-text prompt Built-in output modes 5 (transcribe, verbatim, codemix, translit, translate) Transcript plus Smart Formatting Verbatim or no_verbatim Transcript Real-time streaming WebSocket, under 150 ms TTFT (vendor claim) Yes (WebSocket) Scribe v2 Realtime, about 150 ms File streaming; live via Realtime API Speaker diarization Batch API Yes Up to 32 speakers Separate gpt-4o-transcribe-diarize model List price \u20b930\/hour $0.0052\/min (multilingual, pre-recorded) $0.22\/hour (batch) ~$0.006\/min Self-hosting Not for V4 yet (v3 on SageMaker) Yes Cloud API Cloud API Key Takeaways Saaras V4 covers all 22 Indian languages plus global English in 1 model. A 3B hybrid state-space decoder, trained from scratch, sits behind an audio encoder. Keyterm prompting accepts up to 50 terms and scored 16.03% WER on IndicContextEval L5. 5 output modes and sub-150 ms streaming TTFT come from the same model. API-only today at \u20b930 per hour; self-hosting docs still cover v3. Check out the Technical Details. All credit goes to the researcher of this project. Also,\u00a0feel free to follow us on\u00a0Twitter\u00a0and don\u2019t forget to join our\u00a0150k+ML SubReddit\u00a0and Subscribe to\u00a0our Newsletter. Wait! are you on telegram?\u00a0now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?\u00a0Connect with us The post Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English appeared first on MarkTechPost.<\/p>","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"pmpro_default_level":"","site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"_pvb_checkbox_block_on_post":false,"footnotes":""},"categories":[52,5,7,1],"tags":[],"class_list":["post-120467","post","type-post","status-publish","format-standard","hentry","category-ai-club","category-committee","category-news","category-uncategorized","pmpro-has-access"],"acf":[],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v25.3 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English - YouZum<\/title>\n<meta name=\"description\" content=\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/youzum.net\/es\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/\" \/>\n<meta property=\"og:locale\" content=\"es_ES\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English - YouZum\" \/>\n<meta property=\"og:description\" content=\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\" \/>\n<meta property=\"og:url\" content=\"https:\/\/youzum.net\/es\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/\" \/>\n<meta property=\"og:site_name\" content=\"YouZum\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/DroneAssociationTH\/\" \/>\n<meta property=\"article:published_time\" content=\"2026-09-27T02:26:49+00:00\" \/>\n<meta name=\"author\" content=\"admin NU\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Escrito por\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin NU\" \/>\n\t<meta name=\"twitter:label2\" content=\"Tiempo de lectura\" \/>\n\t<meta name=\"twitter:data2\" content=\"4 minutos\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/youzum.net\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/#article\",\"isPartOf\":{\"@id\":\"https:\/\/youzum.net\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/\"},\"author\":{\"name\":\"admin NU\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c\"},\"headline\":\"Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English\",\"datePublished\":\"2026-09-27T02:26:49+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/youzum.net\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/\"},\"wordCount\":842,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\"},\"articleSection\":[\"AI\",\"Committee\",\"News\",\"Uncategorized\"],\"inLanguage\":\"es\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/youzum.net\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/youzum.net\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/\",\"url\":\"https:\/\/youzum.net\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/\",\"name\":\"Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English - YouZum\",\"isPartOf\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#website\"},\"datePublished\":\"2026-09-27T02:26:49+00:00\",\"description\":\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\",\"breadcrumb\":{\"@id\":\"https:\/\/youzum.net\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/#breadcrumb\"},\"inLanguage\":\"es\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/youzum.net\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/\"]}]},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/youzum.net\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/youzum.net\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/yousum.gpucore.co\/#website\",\"url\":\"https:\/\/yousum.gpucore.co\/\",\"name\":\"YouSum\",\"description\":\"\",\"publisher\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/yousum.gpucore.co\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"es\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\",\"name\":\"Drone Association Thailand\",\"url\":\"https:\/\/yousum.gpucore.co\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"es\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png\",\"contentUrl\":\"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png\",\"width\":300,\"height\":300,\"caption\":\"Drone Association Thailand\"},\"image\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/\"},\"sameAs\":[\"https:\/\/www.facebook.com\/DroneAssociationTH\/\"]},{\"@type\":\"Person\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c\",\"name\":\"admin NU\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"es\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png\",\"contentUrl\":\"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png\",\"caption\":\"admin NU\"},\"url\":\"https:\/\/youzum.net\/es\/members\/adminnu\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English - YouZum","description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/youzum.net\/es\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/","og_locale":"es_ES","og_type":"article","og_title":"Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English - YouZum","og_description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","og_url":"https:\/\/youzum.net\/es\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/","og_site_name":"YouZum","article_publisher":"https:\/\/www.facebook.com\/DroneAssociationTH\/","article_published_time":"2026-09-27T02:26:49+00:00","author":"admin NU","twitter_card":"summary_large_image","twitter_misc":{"Escrito por":"admin NU","Tiempo de lectura":"4 minutos"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/youzum.net\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/#article","isPartOf":{"@id":"https:\/\/youzum.net\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/"},"author":{"name":"admin NU","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c"},"headline":"Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English","datePublished":"2026-09-27T02:26:49+00:00","mainEntityOfPage":{"@id":"https:\/\/youzum.net\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/"},"wordCount":842,"commentCount":0,"publisher":{"@id":"https:\/\/yousum.gpucore.co\/#organization"},"articleSection":["AI","Committee","News","Uncategorized"],"inLanguage":"es","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/youzum.net\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/youzum.net\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/","url":"https:\/\/youzum.net\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/","name":"Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English - YouZum","isPartOf":{"@id":"https:\/\/yousum.gpucore.co\/#website"},"datePublished":"2026-09-27T02:26:49+00:00","description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","breadcrumb":{"@id":"https:\/\/youzum.net\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/#breadcrumb"},"inLanguage":"es","potentialAction":[{"@type":"ReadAction","target":["https:\/\/youzum.net\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/"]}]},{"@type":"BreadcrumbList","@id":"https:\/\/youzum.net\/sarvam-ai-releases-saaras-v4-a-speech-to-text-model-for-all-22-indian-languages-and-global-english\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/youzum.net\/"},{"@type":"ListItem","position":2,"name":"Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English"}]},{"@type":"WebSite","@id":"https:\/\/yousum.gpucore.co\/#website","url":"https:\/\/yousum.gpucore.co\/","name":"YouSum","description":"","publisher":{"@id":"https:\/\/yousum.gpucore.co\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/yousum.gpucore.co\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"es"},{"@type":"Organization","@id":"https:\/\/yousum.gpucore.co\/#organization","name":"Drone Association Thailand","url":"https:\/\/yousum.gpucore.co\/","logo":{"@type":"ImageObject","inLanguage":"es","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/","url":"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png","contentUrl":"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png","width":300,"height":300,"caption":"Drone Association Thailand"},"image":{"@id":"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/DroneAssociationTH\/"]},{"@type":"Person","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c","name":"admin NU","image":{"@type":"ImageObject","inLanguage":"es","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/image\/","url":"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png","contentUrl":"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png","caption":"admin NU"},"url":"https:\/\/youzum.net\/es\/members\/adminnu\/"}]}},"rttpg_featured_image_url":null,"rttpg_author":{"display_name":"admin NU","author_link":"https:\/\/youzum.net\/es\/members\/adminnu\/"},"rttpg_comment":0,"rttpg_category":"<a href=\"https:\/\/youzum.net\/es\/category\/ai-club\/\" rel=\"category tag\">AI<\/a> <a href=\"https:\/\/youzum.net\/es\/category\/committee\/\" rel=\"category tag\">Committee<\/a> <a href=\"https:\/\/youzum.net\/es\/category\/news\/\" rel=\"category tag\">News<\/a> <a href=\"https:\/\/youzum.net\/es\/category\/uncategorized\/\" rel=\"category tag\">Uncategorized<\/a>","rttpg_excerpt":"Sarvam AI has released Saaras V4, the newest generation of its speech recognition model. It covers all 22 scheduled Indian languages plus English, now including global English accents. Sarvam reports state-of-the-art accuracy across all 22 languages. Is it deployable? Yes, through Sarvam\u2019s API today, using model=\"saaras:v4\". Weights are not public, and Sarvam\u2019s SageMaker self-hosting docs&hellip;","_links":{"self":[{"href":"https:\/\/youzum.net\/es\/wp-json\/wp\/v2\/posts\/120467","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/youzum.net\/es\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/youzum.net\/es\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/youzum.net\/es\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/youzum.net\/es\/wp-json\/wp\/v2\/comments?post=120467"}],"version-history":[{"count":0,"href":"https:\/\/youzum.net\/es\/wp-json\/wp\/v2\/posts\/120467\/revisions"}],"wp:attachment":[{"href":"https:\/\/youzum.net\/es\/wp-json\/wp\/v2\/media?parent=120467"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/youzum.net\/es\/wp-json\/wp\/v2\/categories?post=120467"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/youzum.net\/es\/wp-json\/wp\/v2\/tags?post=120467"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}