{"id":116205,"date":"2026-09-07T01:25:27","date_gmt":"2026-09-07T01:25:27","guid":{"rendered":"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/"},"modified":"2026-09-07T01:25:27","modified_gmt":"2026-09-07T01:25:27","slug":"perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed","status":"publish","type":"post","link":"https:\/\/youzum.net\/de\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/","title":{"rendered":"Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed"},"content":{"rendered":"<p class=\"wp-block-paragraph\">Retrieval quality in an AI search product is bounded by two things: how good the embedding model is, and how cheaply you can run it across an index. This week, Perplexity Engineering team published <a href=\"https:\/\/www.perplexity.ai\/hub\/blog\/fast-embeddings-on-gpus\">Fast Embeddings on GPUs<\/a>, an under-the-hood account of the second \u2014 the serving infrastructure behind <a href=\"https:\/\/research.perplexity.ai\/articles\/pplx-embed-state-of-the-art-embedding-models-for-web-scale-retrieval\">pplx-embed<\/a> and the ranking models used across Perplexity Search, Computer and the API Platform.<\/p>\n<p class=\"wp-block-paragraph\">Perplexity team states that embedding inference on the GPU side has largely converged across engines on mature Hopper and Blackwell hardware. The wins sit in the runtime and harness around the model: CUDA graph management, an async result-tracking abstraction, and a Rust request path.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Two traffic patterns, one engine<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">Perplexity frames embedding serving as two workloads. <strong>Batch embedding<\/strong> happens when building or re-indexing the vector database, where throughput minimizes cost. <strong>Online embedding<\/strong> happens at query time, where a short query must be embedded fast. Scoring sits in between: after vector search, large document batches are ranked, balancing both.<\/p>\n<p class=\"wp-block-paragraph\">The key decision is that Perplexity did not build a separate embedding engine. Because embedding models are small Transformers, batch embedding resembles compute-bound prefill and online embedding, often a few tokens, resembles memory-bound decode. So the research team reuses the <a href=\"https:\/\/research.perplexity.ai\/articles\/cutedsl-at-perplexity\">prefill and decode kernels<\/a> from its LLM stack.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Ivy, Tulip and ROSE<\/strong><\/h2>\n<p class=\"wp-block-paragraph\"><strong>Three services handle a request:<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Ivy<\/strong> is a Rust HTTP gateway. It does the CPU-side work \u2014 JSON parsing, <a href=\"https:\/\/research.perplexity.ai\/articles\/improving-unigram-tokenizer-cpu-performance\">tokenization<\/a>, input templating, batch splitting \u2014 and translates requests into a custom gRPC protocol. It also splits large-batch requests into chunks and load-balances them across replicas, which corrects the load imbalance that arises when production payloads vary in size.<\/li>\n<li><strong>Tulip<\/strong> is the inference server interface: a gRPC server built with Rust, <code>tokio<\/code> and <a href=\"https:\/\/docs.rs\/tonic\/latest\/tonic\/\"><code>tonic<\/code><\/a>, handling scheduling and batching before dispatching to the engine.<\/li>\n<li><strong><a href=\"https:\/\/research.perplexity.ai\/articles\/gpt-oss-on-day-0\">ROSE<\/a><\/strong> (Runtime-Optimized Serving Engine) implements model inference. It is primarily Python, provides kernels, layers and model definitions, manages CUDA graphs, and exposes a <code>step()<\/code> function to Tulip.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\"><strong>Why the scheduler is deliberately simple<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">Tulip picks sequences first-come, first-served while requests accumulate. That simplicity is justified by a measurement: for small embedding models at the sequence lengths Perplexity serves, the linear cost of dense layers dominates the quadratic cost of attention. Latency is therefore roughly proportional to token count, not sequence count. Once a batch saturates the GPU, around <strong>512 tokens on a sub-billion-parameter model<\/strong>, packing in more sequences does not improve efficiency.<\/p>\n<h2 class=\"wp-block-heading\"><strong>CUDA graphs and LazyTensors<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">On small batches, CPU-side kernel launching can outweigh GPU execution. Perplexity builds <strong>whole-model CUDA graphs<\/strong> for all embedding models, capturing every launch into a single driver call. Because embedding models are small, the inflection point where GPU work exceeds launch cost arrives at batches of thousands of tokens and tens of sequences. Some attention implementations block full-model graphs by depending on dynamic host-side inputs; Perplexity <a href=\"https:\/\/github.com\/flashinfer-ai\/flashinfer\/issues\/626\">upstreamed changes to FlashInfer<\/a> to enable capture.<\/p>\n<p class=\"wp-block-paragraph\">Graphs must be captured per configuration, so token counts are padded to buckets that are multiples of 64 or 256. That still yields thousands of graphs and multiple minutes of capture per model. The fix is <strong>lazy capture<\/strong>: each configuration gets an eager warmup run, then triggers capture and replay on its second hit. This costs p99 latency at startup but spreads minutes of eager work across hours.<\/p>\n<p class=\"wp-block-paragraph\">The second piece is the <strong><code>LazyTensor<\/code><\/strong>, which tracks a page-locked host buffer plus a <code>cudaMemcpyAsync<\/code> and a CUDA event. Instead of <code>step()<\/code> blocking on the device, it returns a <code>LazyTensor<\/code>, letting a Rust async task wait on batch N while the CPU enqueues N+1.<\/p>\n<p> Send a request&lt;\/button&gt;&lt;\/div&gt;<br \/>\n&lt;\/div&gt;<br \/>\n&lt;\/div&gt;<\/p>\n<p>&lt;div class=&quot;&rdquo;pe-panel&rdquo;&quot; id=&quot;&rdquo;peP1&Prime;&quot;&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-card&rdquo;&quot;&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-txt&rdquo;&quot;&gt;On small batches, CPU-side kernel launches can outweigh GPU work. A whole-model &lt;b&gt;CUDA graph&lt;\/b&gt; captures every launch into one call to the driver, so the CPU is freed to enqueue the next batch. Toggle the two modes.&lt;\/div&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-ctl&rdquo;&quot; style=&quot;&rdquo;margin:0&quot; 0 12px&rdquo;&gt;<br \/>\n&lt;button class=&#8221;pe-btn ghost on&#8221; id=&#8221;peEager&#8221;&gt;Eager launches&lt;\/button&gt;<br \/>\n&lt;button class=&#8221;pe-btn ghost&#8221; id=&#8221;peGraph&#8221;&gt;CUDA graph&lt;\/button&gt;<br \/>\n&lt;\/div&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-lane&rdquo;&quot;&gt;&lt;div class=&quot;&rdquo;pe-lbl&rdquo;&quot;&gt;Host \/ CPU&lt;\/div&gt;&lt;div class=&quot;&rdquo;pe-track&rdquo;&quot; id=&quot;&rdquo;peCpuT&rdquo;&quot;&gt;&lt;\/div&gt;&lt;\/div&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-lane&rdquo;&quot;&gt;&lt;div class=&quot;&rdquo;pe-lbl&rdquo;&quot;&gt;Device \/ GPU&lt;\/div&gt;&lt;div class=&quot;&rdquo;pe-track&rdquo;&quot; id=&quot;&rdquo;peGpuT&rdquo;&quot;&gt;&lt;\/div&gt;&lt;\/div&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-stats&rdquo;&quot;&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-stat&rdquo;&quot;&gt;&lt;div class=&quot;&rdquo;v&rdquo;&quot; id=&quot;&rdquo;peLaunches&rdquo;&quot;&gt;&mdash;&lt;\/div&gt;&lt;div class=&quot;&rdquo;k&rdquo;&quot;&gt;Driver calls&lt;\/div&gt;&lt;\/div&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-stat&rdquo;&quot;&gt;&lt;div class=&quot;&rdquo;v&rdquo;&quot; id=&quot;&rdquo;peGap&rdquo;&quot;&gt;&mdash;&lt;\/div&gt;&lt;div class=&quot;&rdquo;k&rdquo;&quot;&gt;GPU idle gaps&lt;\/div&gt;&lt;\/div&gt;<br \/>\n&lt;\/div&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-txt&rdquo;&quot; style=&quot;&rdquo;margin:12px&quot; 0 0;font-size:11.5px;color:#6e8285&Prime;&gt;Schematic. Block widths illustrate the launch-overhead pattern described in the post, not measured timings.&lt;\/div&gt;<br \/>\n&lt;\/div&gt;<br \/>\n&lt;\/div&gt;<\/p>\n<p>&lt;div class=&quot;&rdquo;pe-panel&rdquo;&quot; id=&quot;&rdquo;peP2&Prime;&quot;&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-card&rdquo;&quot;&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-txt&rdquo;&quot;&gt;Reading results back normally forces a host sync. A &lt;b&gt;LazyTensor&lt;\/b&gt; tracks a page-locked host buffer plus an async device-to-host copy and a CUDA event, so Tulip can block on batch N while the CPU already prepares batch N+1.&lt;\/div&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-lane&rdquo;&quot;&gt;&lt;div class=&quot;&rdquo;pe-lbl&rdquo;&quot;&gt;CPU &mdash; prepare \/ sync&lt;\/div&gt;&lt;div class=&quot;&rdquo;pe-track&rdquo;&quot; id=&quot;&rdquo;peLzC&rdquo;&quot;&gt;&lt;\/div&gt;&lt;\/div&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-lane&rdquo;&quot;&gt;&lt;div class=&quot;&rdquo;pe-lbl&rdquo;&quot;&gt;GPU &mdash; forward pass&lt;\/div&gt;&lt;div class=&quot;&rdquo;pe-track&rdquo;&quot; id=&quot;&rdquo;peLzG&rdquo;&quot;&gt;&lt;\/div&gt;&lt;\/div&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-ctl&rdquo;&quot;&gt;<br \/>\n&lt;button class=&#8221;pe-btn&#8221; id=&#8221;peLzRun&#8221;&gt;<img decoding=\"async\" src=\"https:\/\/s.w.org\/images\/core\/emoji\/17.0.2\/72x72\/25b6.png\" alt=\"\u25b6\" class=\"wp-smiley\" \/> Run 3 batches&lt;\/button&gt;<br \/>\n&lt;button class=&#8221;pe-btn ghost on&#8221; id=&#8221;peLzOn&#8221;&gt;Overlapped&lt;\/button&gt;<br \/>\n&lt;button class=&#8221;pe-btn ghost&#8221; id=&#8221;peLzOff&#8221;&gt;Blocking&lt;\/button&gt;<br \/>\n&lt;\/div&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-note&rdquo;&quot; style=&quot;&rdquo;margin-top:12px&rdquo;&quot; id=&quot;&rdquo;peLzNote&rdquo;&quot;&gt;Overlapped: while the GPU chews batch N, the CPU is already tokenizing and packing batch N+1.&lt;\/div&gt;<br \/>\n&lt;\/div&gt;<br \/>\n&lt;\/div&gt;<\/p>\n<p>&lt;div class=&quot;&rdquo;pe-panel&rdquo;&quot; id=&quot;&rdquo;peP3&Prime;&quot;&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-card&rdquo;&quot;&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-txt&rdquo;&quot;&gt;For small embedding models at these sequence lengths, the linear cost of dense layers dominates the quadratic cost of attention, so latency tracks &lt;b&gt;token count, not sequence count&lt;\/b&gt;. Past roughly &lt;b&gt;512 tokens&lt;\/b&gt; on a sub-1B model, the GPU is saturated and packing in more sequences stops helping.&lt;\/div&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-lbl&rdquo;&quot; style=&quot;&rdquo;margin-top:6px&rdquo;&quot;&gt;Tokens in batch: &lt;span id=&quot;&rdquo;peTokV&rdquo;&quot; style=&quot;&rdquo;color:#3FB6C4&Prime;&quot;&gt;512&lt;\/span&gt;&lt;\/div&gt;<br \/>\n&lt;input type=&#8221;range&#8221; id=&#8221;peTok&#8221; min=&#8221;32&#8243; max=&#8221;4096&#8243; step=&#8221;32&#8243; value=&#8221;512&#8243;&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-lbl&rdquo;&quot;&gt;GPU utilisation&lt;\/div&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-bar&rdquo;&quot;&gt;&lt;div class=&quot;&rdquo;pe-fill&rdquo;&quot; id=&quot;&rdquo;peUtil&rdquo;&quot;&gt;&lt;\/div&gt;&lt;\/div&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-stats&rdquo;&quot;&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-stat&rdquo;&quot;&gt;&lt;div class=&quot;&rdquo;v&rdquo;&quot; id=&quot;&rdquo;peUtilV&rdquo;&quot;&gt;&mdash;&lt;\/div&gt;&lt;div class=&quot;&rdquo;k&rdquo;&quot;&gt;Saturation&lt;\/div&gt;&lt;\/div&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-stat&rdquo;&quot;&gt;&lt;div class=&quot;&rdquo;v&rdquo;&quot; id=&quot;&rdquo;peState&rdquo;&quot;&gt;&mdash;&lt;\/div&gt;&lt;div class=&quot;&rdquo;k&rdquo;&quot;&gt;Regime&lt;\/div&gt;&lt;\/div&gt;<br \/>\n&lt;\/div&gt;<br \/>\n&lt;div class=&quot;&rdquo;pe-txt&rdquo;&quot; style=&quot;&rdquo;margin:12px&quot; 0 0;font-size:11.5px;color:#6e8285&Prime;&gt;Illustrative curve. The ~512-token saturation point is the figure stated in the post; the shape between points is a stand-in, not a benchmark.&lt;\/div&gt;<br \/>\n&lt;\/div&gt;<br \/>\n&lt;\/div&gt;<\/p>\n<p>&lt;div class=&quot;&rdquo;pe-foot&rdquo;&quot;&gt;<br \/>\n&lt;span&gt;Source: Perplexity Engineering, &ldquo;Fast Embeddings on GPUs&rdquo; (Sep 4, 2026)&lt;\/span&gt;<br \/>\n&lt;span&gt;&lt;a href=&quot;\/de\/&rdquo;https:\/\/www.marktechpost.com&rdquo;\/&quot;&gt;Built by Marktechpost&lt;\/a&gt;&lt;\/span&gt;<br \/>\n&lt;\/div&gt;<\/p>\n<p>&lt;script&gt;<br \/>\n(function(){<br \/>\nvar R=document.getElementById(&#8216;pplxEmbedExplainer&#8217;);<br \/>\nvar NOTES=[<br \/>\n&lsquo;&lt;b&gt;Ivy&lt;\/b&gt; &mdash; parses JSON, tokenizes with the in-house unigram tokenizer, applies input templating and splits large batches, then translates to a custom gRPC protocol. It also load-balances chunks across replicas.&rsquo;,<br \/>\n&lsquo;&lt;b&gt;Tulip&lt;\/b&gt; &mdash; Rust gRPC server on tokio and tonic. Requests accumulate while it dispatches or waits; sequences are picked first-come, first-served and packed into a batch for the accelerator.&rsquo;,<br \/>\n&lsquo;&lt;b&gt;ROSE&lt;\/b&gt; &mdash; the Runtime-Optimized Serving Engine. Python-defined kernels and layers, CUDA-graph management, and a step() function that returns a handle to the GPU computation. No KV cache is allocated for embeddings.&rsquo;<br \/>\n];<br \/>\nfunction q(s){return R.querySelector(s)} function qa(s){return R.querySelectorAll(s)}<br \/>\n\/* tabs *\/<br \/>\nqa(&#8216;.pe-tab&#8217;).forEach(function(t){t.addEventListener(&#8216;click&#8217;,function(){<br \/>\nqa(&#8216;.pe-tab&#8217;).forEach(function(x){x.classList.remove(&#8216;on&#8217;)});t.classList.add(&#8216;on&#8217;);<br \/>\nqa(&#8216;.pe-panel&#8217;).forEach(function(p,i){p.classList.toggle(&#8216;on&#8217;,i==+t.dataset.p)});});});<br \/>\n\/* 1 flow *\/<br \/>\nvar note=q(&#8216;#peNote&#8217;),dot=q(&#8216;#peDot&#8217;);<br \/>\nfunction sel(i){qa(&#8216;.pe-node&#8217;).forEach(function(n,k){n.classList.toggle(&#8216;hot&#8217;,k==i)});note.innerHTML=NOTES[i];}<br \/>\nqa(&#8216;.pe-node&#8217;).forEach(function(n){n.addEventListener(&#8216;click&#8217;,function(){sel(+n.dataset.i)})});<br \/>\nsel(0);<br \/>\nvar busy=false;<br \/>\nq(&#8216;#peGo&#8217;).addEventListener(&#8216;click&#8217;,function(){<br \/>\nif(busy)return;busy=true;var steps=[[0,&#8217;3%&#8217;],[1,&#8217;40%&#8217;],[2,&#8217;76%&#8217;]],k=0;<br \/>\ndot.style.transition=&#8217;none&#8217;;dot.style.left=&#8217;3%&#8217;;dot.style.opacity=&#8217;1&#8242;;<br \/>\nsel(0);<br \/>\nvar iv=setInterval(function(){k++;if(k&gt;2){clearInterval(iv);dot.style.opacity=&#8217;0&#8242;;busy=false;return;}<br \/>\ndot.style.transition=&#8217;left .8s cubic-bezier(.4,0,.2,1)&#8217;;dot.style.left=steps[k][1];sel(k);},900);<br \/>\n});<br \/>\n\/* 2 cuda graphs *\/<br \/>\nvar cpuT=q(&#8216;#peCpuT&#8217;),gpuT=q(&#8216;#peGpuT&#8217;),mode=&#8217;eager&#8217;;<br \/>\nfunction blk(p,l,w,cls,txt){var d=document.createElement(&#8216;div&#8217;);d.className=&#8217;pe-blk &#8216;+cls;d.style.left=l+&#8217;%&#8217;;d.style.width=w+&#8217;%&#8217;;d.textContent=txt||&#8221;;p.appendChild(d);}<br \/>\nfunction drawG(){<br \/>\ncpuT.innerHTML=&#8221;;gpuT.innerHTML=&#8221;;<br \/>\nif(mode==&#8217;eager&#8217;){<br \/>\nfor(var i=0;i&lt;6;i++){blk(cpuT,1+i*16.4,7,&#8217;pe-cpu&#8217;,&#8217;launch&#8217;);blk(gpuT,8.4+i*16.4,7.6,&#8217;pe-gpu&#8217;,&#8217;kernel&#8217;);if(i&lt;5)blk(gpuT,16+i*16.4,7.2,&#8217;pe-idle&#8217;,&#8217;idle&#8217;);}<br \/>\nq(&#8216;#peLaunches&#8217;).textContent=&#8217;6&#8242;;q(&#8216;#peGap&#8217;).textContent=&#8217;5&#8242;;<br \/>\n}else{<br \/>\nblk(cpuT,1,12,&#8217;pe-cpu&#8217;,&#8217;graph launch&#8217;);blk(cpuT,15,26,&#8217;pe-cpu&#8217;,&#8217;prepare next batch&#8217;);<br \/>\nblk(gpuT,13.5,84,&#8217;pe-gpu&#8217;,&#8217;6 kernels \u2014 one replay, no gaps&#8217;);<br \/>\nq(&#8216;#peLaunches&#8217;).textContent=&#8217;1&#8242;;q(&#8216;#peGap&#8217;).textContent=&#8217;0&#8242;;<br \/>\n}}<br \/>\nq(&#8216;#peEager&#8217;).addEventListener(&#8216;click&#8217;,function(){mode=&#8217;eager&#8217;;q(&#8216;#peEager&#8217;).classList.add(&#8216;on&#8217;);q(&#8216;#peGraph&#8217;).classList.remove(&#8216;on&#8217;);drawG();});<br \/>\nq(&#8216;#peGraph&#8217;).addEventListener(&#8216;click&#8217;,function(){mode=&#8217;graph&#8217;;q(&#8216;#peGraph&#8217;).classList.add(&#8216;on&#8217;);q(&#8216;#peEager&#8217;).classList.remove(&#8216;on&#8217;);drawG();});<br \/>\ndrawG();<br \/>\n\/* 3 lazytensor *\/<br \/>\nvar lzC=q(&#8216;#peLzC&#8217;),lzG=q(&#8216;#peLzG&#8217;),lzMode=&#8217;on&#8217;,lzBusy=false,lzNote=q(&#8216;#peLzNote&#8217;);<br \/>\nfunction drawLz(){<br \/>\nlzC.innerHTML=&#8221;;lzG.innerHTML=&#8221;;<br \/>\nfor(var i=0;i&lt;3;i++){<br \/>\nif(lzMode==&#8217;on&#8217;){blk(lzC,2+i*32,14,&#8217;pe-cpu&#8217;,&#8217;prep &#8216;+(i+1));blk(lzG,17+i*32,26,&#8217;pe-gpu&#8217;,&#8217;batch &#8216;+(i+1));}<br \/>\nelse{blk(lzC,2+i*32,12,&#8217;pe-cpu&#8217;,&#8217;prep &#8216;+(i+1));blk(lzC,15+i*32,16,&#8217;pe-idle&#8217;,&#8217;wait&#8217;);blk(lzG,15+i*32,15,&#8217;pe-gpu&#8217;,&#8217;batch &#8216;+(i+1));}<br \/>\n}}<br \/>\nfunction lzRun(){<br \/>\nif(lzBusy)return;lzBusy=true;drawLz();<br \/>\nvar bs=lzG.querySelectorAll(&#8216;.pe-blk&#8217;),cs=lzC.querySelectorAll(&#8216;.pe-blk&#8217;);<br \/>\n[].forEach.call(bs,function(b){b.style.opacity=&#8217;.15&#8242;});[].forEach.call(cs,function(b){b.style.opacity=&#8217;.15&#8242;});<br \/>\nvar all=[].concat([].slice.call(cs),[].slice.call(bs)),j=0;<br \/>\nvar iv=setInterval(function(){if(j&gt;=all.length){clearInterval(iv);lzBusy=false;return;}all[j].style.opacity=&#8217;1&#8242;;j++;},220);<br \/>\n}<br \/>\nq(&#8216;#peLzRun&#8217;).addEventListener(&#8216;click&#8217;,lzRun);<br \/>\nq(&#8216;#peLzOn&#8217;).addEventListener(&#8216;click&#8217;,function(){lzMode=&#8217;on&#8217;;q(&#8216;#peLzOn&#8217;).classList.add(&#8216;on&#8217;);q(&#8216;#peLzOff&#8217;).classList.remove(&#8216;on&#8217;);lzNote.innerHTML=&#8217;Overlapped: while the GPU chews batch N, the CPU is already tokenizing and packing batch N+1.&#8217;;drawLz();});<br \/>\nq(&#8216;#peLzOff&#8217;).addEventListener(&#8216;click&#8217;,function(){lzMode=&#8217;off&#8217;;q(&#8216;#peLzOff&#8217;).classList.add(&#8216;on&#8217;);q(&#8216;#peLzOn&#8217;).classList.remove(&#8216;on&#8217;);lzNote.innerHTML=&#8217;Blocking: every step() waits for the device, so the CPU sits idle and the GPU starts late.&#8217;;drawLz();});<br \/>\ndrawLz();<br \/>\n\/* 4 batch shape *\/<br \/>\nvar tok=q(&#8216;#peTok&#8217;);<br \/>\nfunction drawT(){<br \/>\nvar v=+tok.value;q(&#8216;#peTokV&#8217;).textContent=v;<br \/>\nvar u=Math.min(100,Math.round(100*(1-Math.exp(-v\/230))));<br \/>\nq(&#8216;#peUtil&#8217;).style.width=u+&#8217;%&#8217;;q(&#8216;#peUtilV&#8217;).textContent=u+&#8217;%&#8217;;<br \/>\nq(&#8216;#peState&#8217;).textContent=v&lt;512?&#8217;Under-filled&#8217;:(v&lt;1200?&#8217;Saturated&#8217;:&#8217;Throughput-bound&#8217;);<br \/>\nq(&#8216;#peState&#8217;).style.color=v&lt;512?&#8217;#C7A24E&#8217;:&#8217;#3FB6C4&#8242;;<br \/>\n}<br \/>\ntok.addEventListener(&#8216;input&#8217;,drawT);drawT();<br \/>\n})();<br \/>\n&lt;\/script&gt;<br \/>\n&lt;\/div&gt;<\/p>\n<p>&lt;script&gt;<br \/>\n(function(){function h(){var e=document.getElementById(&#8216;pplxEmbedExplainer&#8217;);if(!e)return;parent.postMessage({pplxEmbedH:e.offsetHeight+40},&#8217;*&#8217;);}<br \/>\nwindow.addEventListener(&#8216;load&#8217;,h);setTimeout(h,300);setTimeout(h,1200);<br \/>\ndocument.addEventListener(&#8216;click&#8217;,function(){setTimeout(h,250)});<br \/>\ndocument.addEventListener(&#8216;input&#8217;,function(){setTimeout(h,120)});<br \/>\nif(window.ResizeObserver){var e=document.getElementById(&#8216;pplxEmbedExplainer&#8217;);if(e)new ResizeObserver(h).observe(e);}})();<br \/>\n&lt;\/script&gt;<br \/>\n&lt;\/body&gt;&lt;\/html&gt;&rdquo;&gt;<\/p>\n<p class=\"wp-block-paragraph\">\n<h2 class=\"wp-block-heading\"><strong>Kernels still matter<\/strong><\/h2>\n<\/p><p class=\"wp-block-paragraph\">ROSE supports multiple attention backends for ragged inputs: FlashInfer 2, FlashInfer 3 and FlashAttention 4. Perplexity team reports FlashAttention 4 is generally faster, but FlashInfer 3 outperforms it on <a href=\"https:\/\/research.perplexity.ai\/articles\/hosting-qwen-on-blackwell\">Qwen-based<\/a> models at very long sequence lengths, so backend selection is made case by case. Notably, when serving an embedding model ROSE does not instantiate a KV cache and dispatches to ragged attention variants to avoid padding.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Benchmarks<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">Perplexity benchmarks against <a href=\"https:\/\/github.com\/vllm-project\/vllm\">vLLM<\/a> v0.22.0 in BF16 on real weights and eval-derived inputs, with warmup runs verifying cosine similarity divergence within 0.1%. Four suites are charted: low-latency embeddings (batch 1; 128\/512\/4096 tokens), low-latency scoring (batch 5\/25\/50 at 512 tokens), high-throughput embeddings (batch 100, four concurrent processes) and high-concurrency embeddings (1 to 16 concurrent requests, including Ivy tokenization and network overhead).<\/p>\n<h2 class=\"wp-block-heading\"><strong>Key Takeaways<\/strong><\/h2>\n<ul class=\"wp-block-list\">\n<li>Perplexity\u2019s embedding stack reuses its LLM prefill\/decode kernels rather than running a separate engine.<\/li>\n<li>Latency tracks token count, not sequence count; ~512 tokens saturates a sub-1B model.<\/li>\n<li>Whole-model CUDA graphs plus lazy capture cut launch overhead without minutes-long startup.<\/li>\n<li><code>LazyTensor<\/code> overlaps CPU batch prep with in-flight GPU work instead of blocking on sync.<\/li>\n<li>Ivy, Tulip and ROSE are internal; pplx-embed is reachable via Perplexity\u2019s Embeddings API.<\/li>\n<\/ul>\n<p class=\"wp-block-paragraph\">\n<hr class=\"wp-block-separator has-alpha-channel-opacity\" \/>\n<\/p><p class=\"wp-block-paragraph\">\n<\/p><p class=\"wp-block-paragraph\">Check out the\u00a0<strong><a href=\"https:\/\/www.perplexity.ai\/hub\/blog\/fast-embeddings-on-gpus\" target=\"_blank\" rel=\"noreferrer noopener\">Technical details<\/a><\/strong>. Also,\u00a0feel free to follow us on\u00a0<strong><a href=\"https:\/\/x.com\/intent\/follow?screen_name=marktechpost\" target=\"_blank\" rel=\"noopener\"><mark>Twitter<\/mark><\/a><\/strong>\u00a0and don\u2019t forget to join our\u00a0<strong><a href=\"https:\/\/www.reddit.com\/r\/machinelearningnews\/\" target=\"_blank\" rel=\"noopener\">150k+ML SubReddit<\/a><\/strong>\u00a0and Subscribe to\u00a0<strong><a href=\"https:\/\/magic.beehiiv.com\/v1\/f5e63dd4-5653-4f09-83e2-321a8b1ba526?email=%7B%7Bemail%7D%7D\" target=\"_blank\" rel=\"noopener\">our Newsletter<\/a><\/strong>. Wait! are you on telegram?\u00a0<strong><a href=\"https:\/\/t.me\/machinelearningresearchnews\" target=\"_blank\" rel=\"noopener\">now you can join us on telegram as well.<\/a><\/strong><\/p>\n<p class=\"wp-block-paragraph\">Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?\u00a0<strong><a href=\"https:\/\/forms.gle\/wbash1wF6efRj8G58\" target=\"_blank\" rel=\"noopener\"><mark>Connect with us<\/mark><\/a><\/strong><\/p>\n<p>The post <a href=\"https:\/\/www.marktechpost.com\/2026\/09\/05\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/\">Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed<\/a> appeared first on <a href=\"https:\/\/www.marktechpost.com\/\">MarkTechPost<\/a>.<\/p>","protected":false},"excerpt":{"rendered":"<p>Retrieval quality in an AI search product is bounded by two things: how good the embedding model is, and how cheaply you can run it across an index. This week, Perplexity Engineering team published Fast Embeddings on GPUs, an under-the-hood account of the second &mdash; the serving infrastructure behind pplx-embed and the ranking models used across Perplexity Search, Computer and the API Platform. Perplexity team states that embedding inference on the GPU side has largely converged across engines on mature Hopper and Blackwell hardware. The wins sit in the runtime and harness around the model: CUDA graph management, an async result-tracking abstraction, and a Rust request path. Two traffic patterns, one engine Perplexity frames embedding serving as two workloads. Batch embedding happens when building or re-indexing the vector database, where throughput minimizes cost. Online embedding happens at query time, where a short query must be embedded fast. Scoring sits in between: after vector search, large document batches are ranked, balancing both. The key decision is that Perplexity did not build a separate embedding engine. Because embedding models are small Transformers, batch embedding resembles compute-bound prefill and online embedding, often a few tokens, resembles memory-bound decode. So the research team reuses the prefill and decode kernels from its LLM stack. Ivy, Tulip and ROSE Three services handle a request: Ivy is a Rust HTTP gateway. It does the CPU-side work &mdash; JSON parsing, tokenization, input templating, batch splitting &mdash; and translates requests into a custom gRPC protocol. It also splits large-batch requests into chunks and load-balances them across replicas, which corrects the load imbalance that arises when production payloads vary in size. Tulip is the inference server interface: a gRPC server built with Rust, tokio and tonic, handling scheduling and batching before dispatching to the engine. ROSE (Runtime-Optimized Serving Engine) implements model inference. It is primarily Python, provides kernels, layers and model definitions, manages CUDA graphs, and exposes a step() function to Tulip. Why the scheduler is deliberately simple Tulip picks sequences first-come, first-served while requests accumulate. That simplicity is justified by a measurement: for small embedding models at the sequence lengths Perplexity serves, the linear cost of dense layers dominates the quadratic cost of attention. Latency is therefore roughly proportional to token count, not sequence count. Once a batch saturates the GPU, around 512 tokens on a sub-billion-parameter model, packing in more sequences does not improve efficiency. CUDA graphs and LazyTensors On small batches, CPU-side kernel launching can outweigh GPU execution. Perplexity builds whole-model CUDA graphs for all embedding models, capturing every launch into a single driver call. Because embedding models are small, the inflection point where GPU work exceeds launch cost arrives at batches of thousands of tokens and tens of sequences. Some attention implementations block full-model graphs by depending on dynamic host-side inputs; Perplexity upstreamed changes to FlashInfer to enable capture. Graphs must be captured per configuration, so token counts are padded to buckets that are multiples of 64 or 256. That still yields thousands of graphs and multiple minutes of capture per model. The fix is lazy capture: each configuration gets an eager warmup run, then triggers capture and replay on its second hit. This costs p99 latency at startup but spreads minutes of eager work across hours. The second piece is the LazyTensor, which tracks a page-locked host buffer plus a cudaMemcpyAsync and a CUDA event. Instead of step() blocking on the device, it returns a LazyTensor, letting a Rust async task wait on batch N while the CPU enqueues N+1. Send a request&lt;\/button&gt;&lt;\/div&gt; &lt;\/div&gt; &lt;\/div&gt; &lt;div class=&quot;&rdquo;pe-panel&rdquo;&quot; id=&quot;&rdquo;peP1&Prime;&quot;&gt; &lt;div class=&quot;&rdquo;pe-card&rdquo;&quot;&gt; &lt;div class=&quot;&rdquo;pe-txt&rdquo;&quot;&gt;On small batches, CPU-side kernel launches can outweigh GPU work. A whole-model &lt;b&gt;CUDA graph&lt;\/b&gt; captures every launch into one call to the driver, so the CPU is freed to enqueue the next batch. Toggle the two modes.&lt;\/div&gt; &lt;div class=&quot;&rdquo;pe-ctl&rdquo;&quot; style=&quot;&rdquo;margin:0&quot; 0 12px&rdquo;&gt; &lt;button class=&quot;&rdquo;pe-btn&quot; ghost on&rdquo; id=&quot;&rdquo;peEager&rdquo;&quot;&gt;Eager launches&lt;\/button&gt; &lt;button class=&quot;&rdquo;pe-btn&quot; ghost&rdquo; id=&quot;&rdquo;peGraph&rdquo;&quot;&gt;CUDA graph&lt;\/button&gt; &lt;\/div&gt; &lt;div class=&quot;&rdquo;pe-lane&rdquo;&quot;&gt;&lt;div class=&quot;&rdquo;pe-lbl&rdquo;&quot;&gt;Host \/ CPU&lt;\/div&gt;&lt;div class=&quot;&rdquo;pe-track&rdquo;&quot; id=&quot;&rdquo;peCpuT&rdquo;&quot;&gt;&lt;\/div&gt;&lt;\/div&gt; &lt;div class=&quot;&rdquo;pe-lane&rdquo;&quot;&gt;&lt;div class=&quot;&rdquo;pe-lbl&rdquo;&quot;&gt;Device \/ GPU&lt;\/div&gt;&lt;div class=&quot;&rdquo;pe-track&rdquo;&quot; id=&quot;&rdquo;peGpuT&rdquo;&quot;&gt;&lt;\/div&gt;&lt;\/div&gt; &lt;div class=&quot;&rdquo;pe-stats&rdquo;&quot;&gt; &lt;div class=&quot;&rdquo;pe-stat&rdquo;&quot;&gt;&lt;div class=&quot;&rdquo;v&rdquo;&quot; id=&quot;&rdquo;peLaunches&rdquo;&quot;&gt;&mdash;&lt;\/div&gt;&lt;div class=&quot;&rdquo;k&rdquo;&quot;&gt;Driver calls&lt;\/div&gt;&lt;\/div&gt; &lt;div class=&quot;&rdquo;pe-stat&rdquo;&quot;&gt;&lt;div class=&quot;&rdquo;v&rdquo;&quot; id=&quot;&rdquo;peGap&rdquo;&quot;&gt;&mdash;&lt;\/div&gt;&lt;div class=&quot;&rdquo;k&rdquo;&quot;&gt;GPU idle gaps&lt;\/div&gt;&lt;\/div&gt; &lt;\/div&gt; &lt;div class=&quot;&rdquo;pe-txt&rdquo;&quot; style=&quot;&rdquo;margin:12px&quot; 0 0;font-size:11.5px;color:#6e8285&Prime;&gt;Schematic. Block widths illustrate the launch-overhead pattern described in the post, not measured timings.&lt;\/div&gt; &lt;\/div&gt; &lt;\/div&gt; &lt;div class=&quot;&rdquo;pe-panel&rdquo;&quot; id=&quot;&rdquo;peP2&Prime;&quot;&gt; &lt;div class=&quot;&rdquo;pe-card&rdquo;&quot;&gt; &lt;div class=&quot;&rdquo;pe-txt&rdquo;&quot;&gt;Reading results back normally forces a host sync. A &lt;b&gt;LazyTensor&lt;\/b&gt; tracks a page-locked host buffer plus an async device-to-host copy and a CUDA event, so Tulip can block on batch N while the CPU already prepares batch N+1.&lt;\/div&gt; &lt;div class=&quot;&rdquo;pe-lane&rdquo;&quot;&gt;&lt;div class=&quot;&rdquo;pe-lbl&rdquo;&quot;&gt;CPU &mdash; prepare \/ sync&lt;\/div&gt;&lt;div class=&quot;&rdquo;pe-track&rdquo;&quot; id=&quot;&rdquo;peLzC&rdquo;&quot;&gt;&lt;\/div&gt;&lt;\/div&gt; &lt;div class=&quot;&rdquo;pe-lane&rdquo;&quot;&gt;&lt;div class=&quot;&rdquo;pe-lbl&rdquo;&quot;&gt;GPU &mdash; forward pass&lt;\/div&gt;&lt;div class=&quot;&rdquo;pe-track&rdquo;&quot; id=&quot;&rdquo;peLzG&rdquo;&quot;&gt;&lt;\/div&gt;&lt;\/div&gt; &lt;div class=&quot;&rdquo;pe-ctl&rdquo;&quot;&gt; &lt;button class=&quot;&rdquo;pe-btn&rdquo;&quot; id=&quot;&rdquo;peLzRun&rdquo;&quot;&gt; Run 3 batches&lt;\/button&gt; &lt;button class=&quot;&rdquo;pe-btn&quot; ghost on&rdquo; id=&quot;&rdquo;peLzOn&rdquo;&quot;&gt;Overlapped&lt;\/button&gt; &lt;button class=&quot;&rdquo;pe-btn&quot; ghost&rdquo; id=&quot;&rdquo;peLzOff&rdquo;&quot;&gt;Blocking&lt;\/button&gt; &lt;\/div&gt; &lt;div class=&quot;&rdquo;pe-note&rdquo;&quot; style=&quot;&rdquo;margin-top:12px&rdquo;&quot; id=&quot;&rdquo;peLzNote&rdquo;&quot;&gt;Overlapped: while the GPU chews batch N, the CPU is already tokenizing and packing batch N+1.&lt;\/div&gt; &lt;\/div&gt; &lt;\/div&gt; &lt;div class=&quot;&rdquo;pe-panel&rdquo;&quot; id=&quot;&rdquo;peP3&Prime;&quot;&gt; &lt;div class=&quot;&rdquo;pe-card&rdquo;&quot;&gt; &lt;div class=&quot;&rdquo;pe-txt&rdquo;&quot;&gt;For small embedding models at these sequence lengths, the linear cost of dense layers dominates the quadratic cost of attention, so latency tracks &lt;b&gt;token count, not sequence count&lt;\/b&gt;. Past roughly &lt;b&gt;512 tokens&lt;\/b&gt; on a sub-1B model, the GPU is saturated and packing in more sequences stops helping.&lt;\/div&gt; &lt;div class=&quot;&rdquo;pe-lbl&rdquo;&quot; style=&quot;&rdquo;margin-top:6px&rdquo;&quot;&gt;Tokens in batch: &lt;span id=&quot;&rdquo;peTokV&rdquo;&quot; style=&quot;&rdquo;color:#3FB6C4&Prime;&quot;&gt;512&lt;\/span&gt;&lt;\/div&gt; &lt;input type=&quot;&rdquo;range&rdquo;&quot; id=&quot;&rdquo;peTok&rdquo;&quot; min=&quot;&rdquo;32&Prime;&quot; max=&quot;&rdquo;4096&Prime;&quot; step=&quot;&rdquo;32&Prime;&quot; value=&quot;&rdquo;512&Prime;&quot;&gt; &lt;div class=&quot;&rdquo;pe-lbl&rdquo;&quot;&gt;GPU utilisation&lt;\/div&gt; &lt;div class=&quot;&rdquo;pe-bar&rdquo;&quot;&gt;&lt;div class=&quot;&rdquo;pe-fill&rdquo;&quot; id=&quot;&rdquo;peUtil&rdquo;&quot;&gt;&lt;\/div&gt;&lt;\/div&gt; &lt;div class=&quot;&rdquo;pe-stats&rdquo;&quot;&gt; &lt;div class=&quot;&rdquo;pe-stat&rdquo;&quot;&gt;&lt;div class=&quot;&rdquo;v&rdquo;&quot; id=&quot;&rdquo;peUtilV&rdquo;&quot;&gt;&mdash;&lt;\/div&gt;&lt;div class=&quot;&rdquo;k&rdquo;&quot;&gt;Saturation&lt;\/div&gt;&lt;\/div&gt; &lt;div class=&quot;&rdquo;pe-stat&rdquo;&quot;&gt;&lt;div class=&quot;&rdquo;v&rdquo;&quot; id=&quot;&rdquo;peState&rdquo;&quot;&gt;&mdash;&lt;\/div&gt;&lt;div class=&quot;&rdquo;k&rdquo;&quot;&gt;Regime&lt;\/div&gt;&lt;\/div&gt; &lt;\/div&gt; &lt;div class=&quot;&rdquo;pe-txt&rdquo;&quot; style=&quot;&rdquo;margin:12px&quot; 0 0;font-size:11.5px;color:#6e8285&Prime;&gt;Illustrative curve. The ~512-token saturation point is the figure stated in the post; the shape between points is a stand-in, not a benchmark.&lt;\/div&gt; &lt;\/div&gt; &lt;\/div&gt; &lt;div class=&quot;&rdquo;pe-foot&rdquo;&quot;&gt; &lt;span&gt;Source: Perplexity Engineering, &ldquo;Fast Embeddings on GPUs&rdquo; (Sep 4, 2026)&lt;\/span&gt; &lt;span&gt;&lt;a href=&quot;\/de\/&rdquo;https:\/\/www.marktechpost.com&rdquo;\/&quot;&gt;Built by Marktechpost&lt;\/a&gt;&lt;\/span&gt; &lt;\/div&gt; &lt;script&gt; (function(){ var R=document.getElementById(&lsquo;pplxEmbedExplainer&rsquo;); var NOTES=[ &lsquo;&lt;b&gt;Ivy&lt;\/b&gt; &mdash; parses JSON, tokenizes with the in-house unigram tokenizer, applies input templating and splits large batches, then translates to a custom gRPC protocol. It also load-balances chunks across replicas.&rsquo;, &lsquo;&lt;b&gt;Tulip&lt;\/b&gt; &mdash; Rust gRPC server on tokio and tonic. Requests accumulate while it dispatches or waits; sequences are picked first-come, first-served and packed into a batch for the accelerator.&rsquo;,<\/p>","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"pmpro_default_level":"","site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"_pvb_checkbox_block_on_post":false,"footnotes":""},"categories":[52,5,7,1],"tags":[],"class_list":["post-116205","post","type-post","status-publish","format-standard","hentry","category-ai-club","category-committee","category-news","category-uncategorized","pmpro-has-access"],"acf":[],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v25.3 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed - YouZum<\/title>\n<meta name=\"description\" content=\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/youzum.net\/de\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/\" \/>\n<meta property=\"og:locale\" content=\"de_DE\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed - YouZum\" \/>\n<meta property=\"og:description\" content=\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\" \/>\n<meta property=\"og:url\" content=\"https:\/\/youzum.net\/de\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/\" \/>\n<meta property=\"og:site_name\" content=\"YouZum\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/DroneAssociationTH\/\" \/>\n<meta property=\"article:published_time\" content=\"2026-09-07T01:25:27+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/s.w.org\/images\/core\/emoji\/17.0.2\/72x72\/25b6.png\" \/>\n<meta name=\"author\" content=\"admin NU\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Verfasst von\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin NU\" \/>\n\t<meta name=\"twitter:label2\" content=\"Gesch\u00e4tzte Lesezeit\" \/>\n\t<meta name=\"twitter:data2\" content=\"12\u00a0Minuten\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/#article\",\"isPartOf\":{\"@id\":\"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/\"},\"author\":{\"name\":\"admin NU\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c\"},\"headline\":\"Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed\",\"datePublished\":\"2026-09-07T01:25:27+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/\"},\"wordCount\":2360,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\"},\"image\":{\"@id\":\"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/#primaryimage\"},\"thumbnailUrl\":\"https:\/\/s.w.org\/images\/core\/emoji\/17.0.2\/72x72\/25b6.png\",\"articleSection\":[\"AI\",\"Committee\",\"News\",\"Uncategorized\"],\"inLanguage\":\"de\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/\",\"url\":\"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/\",\"name\":\"Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed - YouZum\",\"isPartOf\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/#primaryimage\"},\"image\":{\"@id\":\"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/#primaryimage\"},\"thumbnailUrl\":\"https:\/\/s.w.org\/images\/core\/emoji\/17.0.2\/72x72\/25b6.png\",\"datePublished\":\"2026-09-07T01:25:27+00:00\",\"description\":\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\",\"breadcrumb\":{\"@id\":\"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/#breadcrumb\"},\"inLanguage\":\"de\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"de\",\"@id\":\"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/#primaryimage\",\"url\":\"https:\/\/s.w.org\/images\/core\/emoji\/17.0.2\/72x72\/25b6.png\",\"contentUrl\":\"https:\/\/s.w.org\/images\/core\/emoji\/17.0.2\/72x72\/25b6.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/youzum.net\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/yousum.gpucore.co\/#website\",\"url\":\"https:\/\/yousum.gpucore.co\/\",\"name\":\"YouSum\",\"description\":\"\",\"publisher\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/yousum.gpucore.co\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"de\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\",\"name\":\"Drone Association Thailand\",\"url\":\"https:\/\/yousum.gpucore.co\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"de\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png\",\"contentUrl\":\"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png\",\"width\":300,\"height\":300,\"caption\":\"Drone Association Thailand\"},\"image\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/\"},\"sameAs\":[\"https:\/\/www.facebook.com\/DroneAssociationTH\/\"]},{\"@type\":\"Person\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c\",\"name\":\"admin NU\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"de\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png\",\"contentUrl\":\"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png\",\"caption\":\"admin NU\"},\"url\":\"https:\/\/youzum.net\/de\/members\/adminnu\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed - YouZum","description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/youzum.net\/de\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/","og_locale":"de_DE","og_type":"article","og_title":"Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed - YouZum","og_description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","og_url":"https:\/\/youzum.net\/de\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/","og_site_name":"YouZum","article_publisher":"https:\/\/www.facebook.com\/DroneAssociationTH\/","article_published_time":"2026-09-07T01:25:27+00:00","og_image":[{"url":"https:\/\/s.w.org\/images\/core\/emoji\/17.0.2\/72x72\/25b6.png","type":"","width":"","height":""}],"author":"admin NU","twitter_card":"summary_large_image","twitter_misc":{"Verfasst von":"admin NU","Gesch\u00e4tzte Lesezeit":"12\u00a0Minuten"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/#article","isPartOf":{"@id":"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/"},"author":{"name":"admin NU","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c"},"headline":"Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed","datePublished":"2026-09-07T01:25:27+00:00","mainEntityOfPage":{"@id":"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/"},"wordCount":2360,"commentCount":0,"publisher":{"@id":"https:\/\/yousum.gpucore.co\/#organization"},"image":{"@id":"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/#primaryimage"},"thumbnailUrl":"https:\/\/s.w.org\/images\/core\/emoji\/17.0.2\/72x72\/25b6.png","articleSection":["AI","Committee","News","Uncategorized"],"inLanguage":"de","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/","url":"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/","name":"Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed - YouZum","isPartOf":{"@id":"https:\/\/yousum.gpucore.co\/#website"},"primaryImageOfPage":{"@id":"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/#primaryimage"},"image":{"@id":"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/#primaryimage"},"thumbnailUrl":"https:\/\/s.w.org\/images\/core\/emoji\/17.0.2\/72x72\/25b6.png","datePublished":"2026-09-07T01:25:27+00:00","description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","breadcrumb":{"@id":"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/#breadcrumb"},"inLanguage":"de","potentialAction":[{"@type":"ReadAction","target":["https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/"]}]},{"@type":"ImageObject","inLanguage":"de","@id":"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/#primaryimage","url":"https:\/\/s.w.org\/images\/core\/emoji\/17.0.2\/72x72\/25b6.png","contentUrl":"https:\/\/s.w.org\/images\/core\/emoji\/17.0.2\/72x72\/25b6.png"},{"@type":"BreadcrumbList","@id":"https:\/\/youzum.net\/perplexity-details-its-gpu-embedding-stack-how-ivy-tulip-and-rose-serve-pplx-embed\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/youzum.net\/"},{"@type":"ListItem","position":2,"name":"Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed"}]},{"@type":"WebSite","@id":"https:\/\/yousum.gpucore.co\/#website","url":"https:\/\/yousum.gpucore.co\/","name":"YouSum","description":"","publisher":{"@id":"https:\/\/yousum.gpucore.co\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/yousum.gpucore.co\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"de"},{"@type":"Organization","@id":"https:\/\/yousum.gpucore.co\/#organization","name":"Drone Association Thailand","url":"https:\/\/yousum.gpucore.co\/","logo":{"@type":"ImageObject","inLanguage":"de","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/","url":"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png","contentUrl":"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png","width":300,"height":300,"caption":"Drone Association Thailand"},"image":{"@id":"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/DroneAssociationTH\/"]},{"@type":"Person","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c","name":"admin NU","image":{"@type":"ImageObject","inLanguage":"de","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/image\/","url":"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png","contentUrl":"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png","caption":"admin NU"},"url":"https:\/\/youzum.net\/de\/members\/adminnu\/"}]}},"rttpg_featured_image_url":null,"rttpg_author":{"display_name":"admin NU","author_link":"https:\/\/youzum.net\/de\/members\/adminnu\/"},"rttpg_comment":0,"rttpg_category":"<a href=\"https:\/\/youzum.net\/de\/category\/ai-club\/\" rel=\"category tag\">AI<\/a> <a href=\"https:\/\/youzum.net\/de\/category\/committee\/\" rel=\"category tag\">Committee<\/a> <a href=\"https:\/\/youzum.net\/de\/category\/news\/\" rel=\"category tag\">News<\/a> <a href=\"https:\/\/youzum.net\/de\/category\/uncategorized\/\" rel=\"category tag\">Uncategorized<\/a>","rttpg_excerpt":"Retrieval quality in an AI search product is bounded by two things: how good the embedding model is, and how cheaply you can run it across an index. This week, Perplexity Engineering team published Fast Embeddings on GPUs, an under-the-hood account of the second \u2014 the serving infrastructure behind pplx-embed and the ranking models used&hellip;","_links":{"self":[{"href":"https:\/\/youzum.net\/de\/wp-json\/wp\/v2\/posts\/116205","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/youzum.net\/de\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/youzum.net\/de\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/youzum.net\/de\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/youzum.net\/de\/wp-json\/wp\/v2\/comments?post=116205"}],"version-history":[{"count":0,"href":"https:\/\/youzum.net\/de\/wp-json\/wp\/v2\/posts\/116205\/revisions"}],"wp:attachment":[{"href":"https:\/\/youzum.net\/de\/wp-json\/wp\/v2\/media?parent=116205"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/youzum.net\/de\/wp-json\/wp\/v2\/categories?post=116205"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/youzum.net\/de\/wp-json\/wp\/v2\/tags?post=116205"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}