{"id":120231,"date":"2026-09-26T02:22:58","date_gmt":"2026-09-26T02:22:58","guid":{"rendered":"https:\/\/youzum.net\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/"},"modified":"2026-09-26T02:22:58","modified_gmt":"2026-09-26T02:22:58","slug":"perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation","status":"publish","type":"post","link":"https:\/\/youzum.net\/it\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/","title":{"rendered":"Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation"},"content":{"rendered":"<p class=\"wp-block-paragraph\"><a href=\"https:\/\/www.perplexity.ai\/hub\/blog\/learning-from-real-world-experience\">Perplexity Research published a new post-training study<\/a>. It trains a model inside <a href=\"https:\/\/www.perplexity.ai\/computer\">Perplexity Computer<\/a> on real user sessions, including failed ones. The method pairs rejection sampling fine-tuning with hint-guided self-distillation. In a live A\/B test, tool-call failures fell from 2.24% to 1.77% between 2 trained checkpoints. Perplexity team reports this as a statistically significant 21.2% relative reduction.<\/p>\n<p class=\"wp-block-paragraph\"><strong>Is it deployable?<\/strong> Not directly. Perplexity has not released the post-trained weights or training code. The model runs only as a model option inside Perplexity Computer. The base model, <a href=\"https:\/\/huggingface.co\/zai-org\/GLM-5.2\">GLM 5.2<\/a>, is openly available on Hugging Face.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Why Outcome-Only Filtering Falls Short<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">Standard <a href=\"https:\/\/arxiv.org\/abs\/2308.01825\">rejection sampling fine-tuning<\/a> (RFT) judges each session and imitates only the successful ones. A successful outcome does not mean every step was correct. An agent can recover from a bad tool call and still deliver the right answer. Imitating that full trajectory can reinforce the error. Discarding failed sessions also throws away clear evidence of avoidable mistakes.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Imitate, Correct, or Keep as Context<\/strong><\/h2>\n<p class=\"wp-block-paragraph\"><strong>Perplexity team separates 2 decisions:<\/strong> which sessions hold behavior worth imitating, and which turns hold mistakes worth correcting. <\/p>\n<p class=\"wp-block-paragraph\"><strong>Each assistant turn gets 1 of 3 treatments:<\/strong><\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Imitate:<\/strong> non-error turns in successful sessions receive cross-entropy (CE) loss.<\/li>\n<li><strong>Correct:<\/strong> error turns with a validated hint receive Kullback-Leibler (KL) divergence loss, in any session.<\/li>\n<li><strong>Keep as context:<\/strong> remaining turns stay in the input but receive no loss.<\/li>\n<\/ul>\n<p class=\"wp-block-paragraph\">Successful sessions can supply both imitation and correction targets. Unsuccessful sessions supply only correction targets.<\/p>\n<h2 class=\"wp-block-heading\"><strong>How a Hint Becomes a Training Signal<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">A hint is a short corrective instruction grounded in information the model already had. In one example, a search call set <code>recency_filter<\/code> to \u2018year.\u2019 The schema allowed only \u2018day,\u2019 \u2018week,\u2019 or \u2018month.\u2019 The hint names the failed call, includes the validation error, and suggests an allowed value or omitting the optional field.<\/p>\n<p class=\"wp-block-paragraph\">The corrective part uses <a href=\"https:\/\/arxiv.org\/abs\/2601.18734\">On-Policy Self-Distillation<\/a> (OPSD). The trainer runs the same GLM 5.2 checkpoint twice on the recorded turn. The teacher pass sees the hint; the student pass does not. Both use teacher forcing, so no replacement answer is generated. The teacher\u2019s next-token probabilities are detached and act as a soft target through forward KL.<\/p>\n<p class=\"wp-block-paragraph\">The combined loss is (CE + \u03bb \u00d7 KL), divided by the number of imitated tokens. Setting \u03bb to 0 recovers standard SFT. The CE term matters. Correction-only training can let teacher and student agree by ignoring context.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Tracing Complaints to the Real Mistake<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">The pipeline draws from training-eligible Computer sessions served by GLM 5.2. Sessions with personally identifiable information and users who opted out are excluded. An LLM judge keeps tasks rated 4 or 5 on a 5-point difficulty scale. Two LLM judges must both approve the final delivery for a session to count as successful.<\/p>\n<p class=\"wp-block-paragraph\">For user feedback, threes LLM judges locate the responsible turn, and at least 2 must agree. This is important because the last assistant turn before a complaint is the root cause only about half the time. Each hint is also checked against information available before the mistake. That check reduces hindsight bias.<\/p>\n<p class=\"wp-block-paragraph\">One example: a user asked for their \u2018w3\u2019 on Paychex. The model assumed a W-2 typo and searched for the wrong form. The hint targets that earlier interpretation, not just the final answer.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Interactive Explainer<\/strong><\/h2>\n<div>\n<\/div>\n<p class=\"wp-block-paragraph\">\n<h2 class=\"wp-block-heading\"><strong>What the Evaluations Show<\/strong><\/h2>\n<ul class=\"wp-block-list\">\n<li><strong>Hints work before training<\/strong>: On 985 held-out tool-error turns, the unchanged base model avoided the original failure in 93.7% of cases with hints, up from 75.1%. The share taking the corrected action rose from 60.6% to 82.3%. On user-feedback turns, fixed or on-track rates rose from 40.0% to 75.0% for explicit evidence. For inferred intent, they rose from 32.5% to 80.0%.<\/li>\n<li><strong>Offline tool errors fell<\/strong>: Recorded tool-error rates were 2.79% for stock GLM 5.2 and 1.35% for RFT only. The RFT plus OPSD checkpoint reached 0.87%. Perplexity notes these checkpoints used different training data, so this is not a matched ablation. Task-level benchmark results on suites like <a href=\"https:\/\/openai.com\/index\/browsecomp\/\">BrowseComp<\/a> and <a href=\"https:\/\/spreadsheetbench.github.io\/\">SpreadsheetBench<\/a> were mixed.<\/li>\n<li><strong>Live results are narrower<\/strong>: Each A\/B test used about 100,000 users per condition. An early checkpoint versus stock GLM 5.2 showed 2.82% versus 2.94% failures, which was not significant. The later checkpoint comparison produced the significant 21.2% drop, without hints at inference. Strong dissatisfaction moved from 2.58% to 2.54%, which was also not significant. Perplexity did not compare the later checkpoint directly against stock GLM 5.2 online.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\"><strong>Key Takeaways<\/strong><\/h2>\n<ul class=\"wp-block-list\">\n<li>Perplexity learns from failed sessions, not just successful ones.<\/li>\n<li>Validated hints turn avoidable mistakes into KL correction targets.<\/li>\n<li>1 model acts as teacher (with hint) and student (without).<\/li>\n<li>Live tool-call failures fell from 2.24% to 1.77%.<\/li>\n<li>User dissatisfaction showed no significant change.<\/li>\n<\/ul>\n<\/p><p class=\"wp-block-paragraph\">\n<hr class=\"wp-block-separator has-alpha-channel-opacity\" \/>\n<\/p><p class=\"wp-block-paragraph\">\n<\/p><p class=\"wp-block-paragraph\">Check out the <strong><a href=\"https:\/\/www.perplexity.ai\/hub\/blog\/learning-from-real-world-experience\" target=\"_blank\" rel=\"noreferrer noopener\">Technical Details<\/a><\/strong>. All credit goes to the researcher of this project. Also,\u00a0feel free to follow us on\u00a0<strong><a href=\"https:\/\/x.com\/intent\/follow?screen_name=marktechpost\" target=\"_blank\" rel=\"noopener\"><mark>Twitter<\/mark><\/a><\/strong>\u00a0and don\u2019t forget to join our\u00a0<strong><a href=\"https:\/\/www.reddit.com\/r\/machinelearningnews\/\" target=\"_blank\" rel=\"noopener\">150k+ML SubReddit<\/a><\/strong>\u00a0and Subscribe to\u00a0<strong><a href=\"https:\/\/magic.beehiiv.com\/v1\/f5e63dd4-5653-4f09-83e2-321a8b1ba526?email=%7B%7Bemail%7D%7D\" target=\"_blank\" rel=\"noopener\">our Newsletter<\/a><\/strong>. Wait! are you on telegram?\u00a0<strong><a href=\"https:\/\/t.me\/machinelearningresearchnews\" target=\"_blank\" rel=\"noopener\">now you can join us on telegram as well.<\/a><\/strong><\/p>\n<p class=\"wp-block-paragraph\">Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?\u00a0<strong><a href=\"https:\/\/forms.gle\/MJjjVDPS7whH8Ngs6\" target=\"_blank\" rel=\"noopener\"><mark>Connect with us<\/mark><\/a><\/strong><\/p>\n<p>The post <a href=\"https:\/\/www.marktechpost.com\/2026\/09\/25\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/\">Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation<\/a> appeared first on <a href=\"https:\/\/www.marktechpost.com\/\">MarkTechPost<\/a>.<\/p>","protected":false},"excerpt":{"rendered":"<p>Perplexity Research published a new post-training study. It trains a model inside Perplexity Computer on real user sessions, including failed ones. The method pairs rejection sampling fine-tuning with hint-guided self-distillation. In a live A\/B test, tool-call failures fell from 2.24% to 1.77% between 2 trained checkpoints. Perplexity team reports this as a statistically significant 21.2% relative reduction. Is it deployable? Not directly. Perplexity has not released the post-trained weights or training code. The model runs only as a model option inside Perplexity Computer. The base model, GLM 5.2, is openly available on Hugging Face. Why Outcome-Only Filtering Falls Short Standard rejection sampling fine-tuning (RFT) judges each session and imitates only the successful ones. A successful outcome does not mean every step was correct. An agent can recover from a bad tool call and still deliver the right answer. Imitating that full trajectory can reinforce the error. Discarding failed sessions also throws away clear evidence of avoidable mistakes. Imitate, Correct, or Keep as Context Perplexity team separates 2 decisions: which sessions hold behavior worth imitating, and which turns hold mistakes worth correcting. Each assistant turn gets 1 of 3 treatments: Imitate: non-error turns in successful sessions receive cross-entropy (CE) loss. Correct: error turns with a validated hint receive Kullback-Leibler (KL) divergence loss, in any session. Keep as context: remaining turns stay in the input but receive no loss. Successful sessions can supply both imitation and correction targets. Unsuccessful sessions supply only correction targets. How a Hint Becomes a Training Signal A hint is a short corrective instruction grounded in information the model already had. In one example, a search call set recency_filter to \u2018year.\u2019 The schema allowed only \u2018day,\u2019 \u2018week,\u2019 or \u2018month.\u2019 The hint names the failed call, includes the validation error, and suggests an allowed value or omitting the optional field. The corrective part uses On-Policy Self-Distillation (OPSD). The trainer runs the same GLM 5.2 checkpoint twice on the recorded turn. The teacher pass sees the hint; the student pass does not. Both use teacher forcing, so no replacement answer is generated. The teacher\u2019s next-token probabilities are detached and act as a soft target through forward KL. The combined loss is (CE + \u03bb \u00d7 KL), divided by the number of imitated tokens. Setting \u03bb to 0 recovers standard SFT. The CE term matters. Correction-only training can let teacher and student agree by ignoring context. Tracing Complaints to the Real Mistake The pipeline draws from training-eligible Computer sessions served by GLM 5.2. Sessions with personally identifiable information and users who opted out are excluded. An LLM judge keeps tasks rated 4 or 5 on a 5-point difficulty scale. Two LLM judges must both approve the final delivery for a session to count as successful. For user feedback, threes LLM judges locate the responsible turn, and at least 2 must agree. This is important because the last assistant turn before a complaint is the root cause only about half the time. Each hint is also checked against information available before the mistake. That check reduces hindsight bias. One example: a user asked for their \u2018w3\u2019 on Paychex. The model assumed a W-2 typo and searched for the wrong form. The hint targets that earlier interpretation, not just the final answer. Interactive Explainer What the Evaluations Show Hints work before training: On 985 held-out tool-error turns, the unchanged base model avoided the original failure in 93.7% of cases with hints, up from 75.1%. The share taking the corrected action rose from 60.6% to 82.3%. On user-feedback turns, fixed or on-track rates rose from 40.0% to 75.0% for explicit evidence. For inferred intent, they rose from 32.5% to 80.0%. Offline tool errors fell: Recorded tool-error rates were 2.79% for stock GLM 5.2 and 1.35% for RFT only. The RFT plus OPSD checkpoint reached 0.87%. Perplexity notes these checkpoints used different training data, so this is not a matched ablation. Task-level benchmark results on suites like BrowseComp and SpreadsheetBench were mixed. Live results are narrower: Each A\/B test used about 100,000 users per condition. An early checkpoint versus stock GLM 5.2 showed 2.82% versus 2.94% failures, which was not significant. The later checkpoint comparison produced the significant 21.2% drop, without hints at inference. Strong dissatisfaction moved from 2.58% to 2.54%, which was also not significant. Perplexity did not compare the later checkpoint directly against stock GLM 5.2 online. Key Takeaways Perplexity learns from failed sessions, not just successful ones. Validated hints turn avoidable mistakes into KL correction targets. 1 model acts as teacher (with hint) and student (without). Live tool-call failures fell from 2.24% to 1.77%. User dissatisfaction showed no significant change. Check out the Technical Details. All credit goes to the researcher of this project. Also,\u00a0feel free to follow us on\u00a0Twitter\u00a0and don\u2019t forget to join our\u00a0150k+ML SubReddit\u00a0and Subscribe to\u00a0our Newsletter. Wait! are you on telegram?\u00a0now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?\u00a0Connect with us The post Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation appeared first on MarkTechPost.<\/p>","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"pmpro_default_level":"","site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"_pvb_checkbox_block_on_post":false,"footnotes":""},"categories":[52,5,7,1],"tags":[],"class_list":["post-120231","post","type-post","status-publish","format-standard","hentry","category-ai-club","category-committee","category-news","category-uncategorized","pmpro-has-access"],"acf":[],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v25.3 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation - YouZum<\/title>\n<meta name=\"description\" content=\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/youzum.net\/it\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/\" \/>\n<meta property=\"og:locale\" content=\"it_IT\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation - YouZum\" \/>\n<meta property=\"og:description\" content=\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\" \/>\n<meta property=\"og:url\" content=\"https:\/\/youzum.net\/it\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/\" \/>\n<meta property=\"og:site_name\" content=\"YouZum\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/DroneAssociationTH\/\" \/>\n<meta property=\"article:published_time\" content=\"2026-09-26T02:22:58+00:00\" \/>\n<meta name=\"author\" content=\"admin NU\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Scritto da\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin NU\" \/>\n\t<meta name=\"twitter:label2\" content=\"Tempo di lettura stimato\" \/>\n\t<meta name=\"twitter:data2\" content=\"4 minuti\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/youzum.net\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/#article\",\"isPartOf\":{\"@id\":\"https:\/\/youzum.net\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/\"},\"author\":{\"name\":\"admin NU\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c\"},\"headline\":\"Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation\",\"datePublished\":\"2026-09-26T02:22:58+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/youzum.net\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/\"},\"wordCount\":830,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\"},\"articleSection\":[\"AI\",\"Committee\",\"News\",\"Uncategorized\"],\"inLanguage\":\"it-IT\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/youzum.net\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/youzum.net\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/\",\"url\":\"https:\/\/youzum.net\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/\",\"name\":\"Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation - YouZum\",\"isPartOf\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#website\"},\"datePublished\":\"2026-09-26T02:22:58+00:00\",\"description\":\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\",\"breadcrumb\":{\"@id\":\"https:\/\/youzum.net\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/#breadcrumb\"},\"inLanguage\":\"it-IT\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/youzum.net\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/\"]}]},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/youzum.net\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/youzum.net\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/yousum.gpucore.co\/#website\",\"url\":\"https:\/\/yousum.gpucore.co\/\",\"name\":\"YouSum\",\"description\":\"\",\"publisher\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/yousum.gpucore.co\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"it-IT\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\",\"name\":\"Drone Association Thailand\",\"url\":\"https:\/\/yousum.gpucore.co\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"it-IT\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png\",\"contentUrl\":\"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png\",\"width\":300,\"height\":300,\"caption\":\"Drone Association Thailand\"},\"image\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/\"},\"sameAs\":[\"https:\/\/www.facebook.com\/DroneAssociationTH\/\"]},{\"@type\":\"Person\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c\",\"name\":\"admin NU\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"it-IT\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png\",\"contentUrl\":\"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png\",\"caption\":\"admin NU\"},\"url\":\"https:\/\/youzum.net\/it\/members\/adminnu\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation - YouZum","description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/youzum.net\/it\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/","og_locale":"it_IT","og_type":"article","og_title":"Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation - YouZum","og_description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","og_url":"https:\/\/youzum.net\/it\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/","og_site_name":"YouZum","article_publisher":"https:\/\/www.facebook.com\/DroneAssociationTH\/","article_published_time":"2026-09-26T02:22:58+00:00","author":"admin NU","twitter_card":"summary_large_image","twitter_misc":{"Scritto da":"admin NU","Tempo di lettura stimato":"4 minuti"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/youzum.net\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/#article","isPartOf":{"@id":"https:\/\/youzum.net\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/"},"author":{"name":"admin NU","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c"},"headline":"Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation","datePublished":"2026-09-26T02:22:58+00:00","mainEntityOfPage":{"@id":"https:\/\/youzum.net\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/"},"wordCount":830,"commentCount":0,"publisher":{"@id":"https:\/\/yousum.gpucore.co\/#organization"},"articleSection":["AI","Committee","News","Uncategorized"],"inLanguage":"it-IT","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/youzum.net\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/youzum.net\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/","url":"https:\/\/youzum.net\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/","name":"Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation - YouZum","isPartOf":{"@id":"https:\/\/yousum.gpucore.co\/#website"},"datePublished":"2026-09-26T02:22:58+00:00","description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","breadcrumb":{"@id":"https:\/\/youzum.net\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/#breadcrumb"},"inLanguage":"it-IT","potentialAction":[{"@type":"ReadAction","target":["https:\/\/youzum.net\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/"]}]},{"@type":"BreadcrumbList","@id":"https:\/\/youzum.net\/perplexity-trains-its-computer-agent-on-real-mistakes-with-hint-guided-self-distillation\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/youzum.net\/"},{"@type":"ListItem","position":2,"name":"Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation"}]},{"@type":"WebSite","@id":"https:\/\/yousum.gpucore.co\/#website","url":"https:\/\/yousum.gpucore.co\/","name":"YouSum","description":"","publisher":{"@id":"https:\/\/yousum.gpucore.co\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/yousum.gpucore.co\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"it-IT"},{"@type":"Organization","@id":"https:\/\/yousum.gpucore.co\/#organization","name":"Drone Association Thailand","url":"https:\/\/yousum.gpucore.co\/","logo":{"@type":"ImageObject","inLanguage":"it-IT","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/","url":"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png","contentUrl":"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png","width":300,"height":300,"caption":"Drone Association Thailand"},"image":{"@id":"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/DroneAssociationTH\/"]},{"@type":"Person","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c","name":"admin NU","image":{"@type":"ImageObject","inLanguage":"it-IT","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/image\/","url":"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png","contentUrl":"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png","caption":"admin NU"},"url":"https:\/\/youzum.net\/it\/members\/adminnu\/"}]}},"rttpg_featured_image_url":null,"rttpg_author":{"display_name":"admin NU","author_link":"https:\/\/youzum.net\/it\/members\/adminnu\/"},"rttpg_comment":0,"rttpg_category":"<a href=\"https:\/\/youzum.net\/it\/category\/ai-club\/\" rel=\"category tag\">AI<\/a> <a href=\"https:\/\/youzum.net\/it\/category\/committee\/\" rel=\"category tag\">Committee<\/a> <a href=\"https:\/\/youzum.net\/it\/category\/news\/\" rel=\"category tag\">News<\/a> <a href=\"https:\/\/youzum.net\/it\/category\/uncategorized\/\" rel=\"category tag\">Uncategorized<\/a>","rttpg_excerpt":"Perplexity Research published a new post-training study. It trains a model inside Perplexity Computer on real user sessions, including failed ones. The method pairs rejection sampling fine-tuning with hint-guided self-distillation. In a live A\/B test, tool-call failures fell from 2.24% to 1.77% between 2 trained checkpoints. Perplexity team reports this as a statistically significant 21.2%&hellip;","_links":{"self":[{"href":"https:\/\/youzum.net\/it\/wp-json\/wp\/v2\/posts\/120231","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/youzum.net\/it\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/youzum.net\/it\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/youzum.net\/it\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/youzum.net\/it\/wp-json\/wp\/v2\/comments?post=120231"}],"version-history":[{"count":0,"href":"https:\/\/youzum.net\/it\/wp-json\/wp\/v2\/posts\/120231\/revisions"}],"wp:attachment":[{"href":"https:\/\/youzum.net\/it\/wp-json\/wp\/v2\/media?parent=120231"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/youzum.net\/it\/wp-json\/wp\/v2\/categories?post=120231"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/youzum.net\/it\/wp-json\/wp\/v2\/tags?post=120231"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}