{"id":117261,"date":"2026-09-12T01:44:24","date_gmt":"2026-09-12T01:44:24","guid":{"rendered":"https:\/\/youzum.net\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/"},"modified":"2026-09-12T01:44:24","modified_gmt":"2026-09-12T01:44:24","slug":"can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize","status":"publish","type":"post","link":"https:\/\/youzum.net\/fr\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/","title":{"rendered":"Can LLMs Engineer Their Own Agent Harness? ByteDance Seed\u2019s HarnessDev Says Only 34 of 64 Changes Generalize"},"content":{"rendered":"<p class=\"wp-block-paragraph\">An agent harness is the code around a model: execution loop, tools, context, state, recovery, and verification. Per the <a href=\"https:\/\/www.tbench.ai\/\">Terminal-Bench 2.1<\/a> leaderboard, GPT-5 solves 35.2% of tasks inside Terminus 2 but 49.6% inside Codex CLI with identical weights. Most benchmarks keep that harness fixed. <a href=\"https:\/\/arxiv.org\/abs\/2609.01437\">HarnessDev<\/a> proposed by team of researchers from ByteDance Seed, Singapore University of Technology and Design, Georgia Institute of Technology, M-A-P, and TokenWave.AI, flips the target: the artifact under evaluation is the runnable harness the model writes, not the answer it produces.<\/p>\n<h2 class=\"wp-block-heading\"><strong>2 stages: Creation and Evolution<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">In <strong>Creation<\/strong>, every creator receives the same weak seed: passive file, search, and process primitives plus result and trajectory writers, with no loop, planner, verifier, retry, or stopping rule. Unmodified, it scores 0 everywhere. The creator gets a task-family spec, a short design tutorial, and 1 to 3 development cases, builds a full harness, and the harness is frozen before hidden tasks.<\/p>\n<p class=\"wp-block-paragraph\">In <strong>Evolution<\/strong>, the creator starts from its own frozen Creation code harness and revises it using execution feedback from a fixed set of 100 <a href=\"https:\/\/arxiv.org\/abs\/2509.16941\">SWE-bench Pro<\/a> tasks and all 89 Terminal-Bench 2.1 tasks. Each official candidate must complete both evaluations as a pair, with a budget of 10 pairs and at most 2 five-task probes between pairs. Every official version is later scored on 630 held-out SWE-Pro instances the creator never sees.<\/p>\n<p class=\"wp-block-paragraph\">Harnesses are graded on capability (task success) and efficiency (executor tokens, with creator tokens excluded).<\/p>\n<h2 class=\"wp-block-heading\"><strong>Setup<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">6 creator LLMs were tested: Opus 4.8, GPT-5.5, Gemini 3.1 Pro, DeepSeek V4 Pro, Qwen 3.7 Max, and Seed 2.0 Pro, working inside Claude Code 2.1.177 (GPT-5.5 used Codex 0.144.3). Creation spans 4 domains and 5 benchmarks totaling 2,207 instances: SWE-bench Pro public split (731), Terminal-Bench 2.1 (89), <a href=\"https:\/\/arxiv.org\/abs\/2410.07095\">MLE-bench<\/a> (75), <a href=\"https:\/\/github.com\/EQ-bench\/eqbench3\">EQ-Bench3<\/a> (46), and <a href=\"https:\/\/arxiv.org\/abs\/2504.12516\">BrowseComp<\/a> (1,266). Each creator builds 3 harnesses per benchmark, reported as avg@3. Self-Eval runs each harness with its creator; Unified-Eval runs all with Gemini 3.1 Pro.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Creation results<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">Under Self-Eval, Opus 4.8 posts the highest average score at 67.8 against a human-engineered reference of 86.2. The gap depends on domain:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Code:<\/strong> Opus 4.8 reaches 69.3 on SWE-Pro versus the 80.0 reference. Gemini 3.1 Pro leads Terminal-Bench at 68.8 versus 88.8.<\/li>\n<li><strong>Search:<\/strong> the widest gap. The best BrowseComp score is 52.6 (GPT-5.5) against a 92.2 reference.<\/li>\n<li><strong>Writing:<\/strong> Opus 4.8 scores 84.6 on EQ-Bench3, above the 83.7 reference.<\/li>\n<li><strong>ML experimentation:<\/strong> Opus 4.8 (32.9) and Gemini (32.4) beat the 24.0 MLE-bench reference.<\/li>\n<\/ul>\n<p class=\"wp-block-paragraph\">The SWE-Pro, Terminal-Bench, and BrowseComp references are external results from OpenAI\u2019s GPT-5.6 report, not re-runs.<\/p>\n<p class=\"wp-block-paragraph\">Code volume did not predict quality: the 18 code harnesses added 17,111 net lines, yet Gemini added the fewest (1,006) and led Terminal-Bench. Self-test count barely correlated with score (Spearman 0.13 to 0.26); revision calls reached 0.57.<\/p>\n<p class=\"wp-block-paragraph\">Much generated machinery is inert. Of 108 code component instances, 72 trigger in real runs and 18 never fire, all of them state and memory. 11 of 18 harnesses define a State class, yet no checkpoint event appears across 26,679 trajectories. 124 of 587 writing features are dead code.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Cost and executor transfer<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">MLE-bench token use varied roughly 19-fold. GPT-5.5 hit a 19.1 medal rate with 29.3M tokens while DeepSeek V4 hit 19.6 with 208.4M. Swapping the executor to Gemini reshuffled rankings: Qwen gained 17.6 points on BrowseComp and 12.9 on MLE-bench, while Opus 4.8\u2019s SWE-Pro score fell from 69.3 to 33.0, partly because one harness hard-coded a 120-step limit around its original executor. The Opus search harness\u2019s duplicate-query rate jumped from 10.1% to 88.2% after the switch.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Evolution results<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">9 lineages (5 self-runtime, 4 fixed-Gemini) produced 73 official versions and 64 adjacent switches. All 5 self-runtime creators improved on held-out tasks, from +1.43 to +4.44 points (mean +3.11). Under fixed Gemini, only Opus improved; GPT-5.5 regressed 10.32 points.<\/p>\n<p class=\"wp-block-paragraph\">Progress was not monotonic. Of 64 switches, 8 regressed on both benchmarks, 16 on one, 27 gained only within the noise band, and 2 showed clear positive evidence. A single commit can vary by about \u00b14.75 pair-score points. Feedback and held-out scores moved in the same direction only 34 of 64 times (53.1%), and only 2 of 9 declared final versions were held-out optimal. Of 169 new functions or classes, 25 have no caller.<\/p>\n<p class=\"wp-block-paragraph\"><strong>The clearest win:<\/strong> Opus 4.8 noticed 99 of 100 runs reported success while only 48 passed, traced it to premature completion, and added a completion gate. Failure diagnosis was otherwise the weakest step: the dedicated trajectory interface was called only twice.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Interactive explainer<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">\n<h2 class=\"wp-block-heading\"><strong>Key Takeaways<\/strong><\/h2>\n<ul class=\"wp-block-list\">\n<li>HarnessDev scores the harness a model builds, not the answer it returns.<\/li>\n<li>Self-built harnesses match or beat references on writing and ML experimentation but trail badly on code and search.<\/li>\n<li>Harness quality is executor-specific; Opus 4.8 drops from 69.3 to 33.0 on SWE-Pro under Gemini.<\/li>\n<li>Evolution gains are small, noisy, and only 34 of 64 changes point the same way on held-out tasks.<\/li>\n<li>Much generated state and memory code never executes.<\/li>\n<\/ul>\n<\/p><p class=\"wp-block-paragraph\">\n<hr class=\"wp-block-separator has-alpha-channel-opacity\" \/>\n<\/p><p class=\"wp-block-paragraph\">\n<\/p><p class=\"wp-block-paragraph\">Check out the\u00a0<a href=\"https:\/\/arxiv.org\/abs\/2609.01437\"><strong>Paper<\/strong><\/a> and <a href=\"https:\/\/self-developing-agents.github.io\/\"><strong>Project Page<\/strong><\/a>. Also,\u00a0feel free to follow us on\u00a0<strong><a href=\"https:\/\/x.com\/intent\/follow?screen_name=marktechpost\" target=\"_blank\" rel=\"noopener\"><mark>Twitter<\/mark><\/a><\/strong>\u00a0and don\u2019t forget to join our\u00a0<strong><a href=\"https:\/\/www.reddit.com\/r\/machinelearningnews\/\" target=\"_blank\" rel=\"noopener\">150k+ML SubReddit<\/a><\/strong>\u00a0and Subscribe to\u00a0<strong><a href=\"https:\/\/magic.beehiiv.com\/v1\/f5e63dd4-5653-4f09-83e2-321a8b1ba526?email=%7B%7Bemail%7D%7D\" target=\"_blank\" rel=\"noopener\">our Newsletter<\/a><\/strong>. Wait! are you on telegram?\u00a0<strong><a href=\"https:\/\/t.me\/machinelearningresearchnews\" target=\"_blank\" rel=\"noopener\">now you can join us on telegram as well.<\/a><\/strong><\/p>\n<p class=\"wp-block-paragraph\">Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?\u00a0<strong><a href=\"https:\/\/forms.gle\/wbash1wF6efRj8G58\" target=\"_blank\" rel=\"noopener\"><mark>Connect with us<\/mark><\/a><\/strong><\/p>\n<p>The post <a href=\"https:\/\/www.marktechpost.com\/2026\/09\/11\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/\">Can LLMs Engineer Their Own Agent Harness? ByteDance Seed\u2019s HarnessDev Says Only 34 of 64 Changes Generalize<\/a> appeared first on <a href=\"https:\/\/www.marktechpost.com\/\">MarkTechPost<\/a>.<\/p>","protected":false},"excerpt":{"rendered":"<p>An agent harness is the code around a model: execution loop, tools, context, state, recovery, and verification. Per the Terminal-Bench 2.1 leaderboard, GPT-5 solves 35.2% of tasks inside Terminus 2 but 49.6% inside Codex CLI with identical weights. Most benchmarks keep that harness fixed. HarnessDev proposed by team of researchers from ByteDance Seed, Singapore University of Technology and Design, Georgia Institute of Technology, M-A-P, and TokenWave.AI, flips the target: the artifact under evaluation is the runnable harness the model writes, not the answer it produces. 2 stages: Creation and Evolution In Creation, every creator receives the same weak seed: passive file, search, and process primitives plus result and trajectory writers, with no loop, planner, verifier, retry, or stopping rule. Unmodified, it scores 0 everywhere. The creator gets a task-family spec, a short design tutorial, and 1 to 3 development cases, builds a full harness, and the harness is frozen before hidden tasks. In Evolution, the creator starts from its own frozen Creation code harness and revises it using execution feedback from a fixed set of 100 SWE-bench Pro tasks and all 89 Terminal-Bench 2.1 tasks. Each official candidate must complete both evaluations as a pair, with a budget of 10 pairs and at most 2 five-task probes between pairs. Every official version is later scored on 630 held-out SWE-Pro instances the creator never sees. Harnesses are graded on capability (task success) and efficiency (executor tokens, with creator tokens excluded). Setup 6 creator LLMs were tested: Opus 4.8, GPT-5.5, Gemini 3.1 Pro, DeepSeek V4 Pro, Qwen 3.7 Max, and Seed 2.0 Pro, working inside Claude Code 2.1.177 (GPT-5.5 used Codex 0.144.3). Creation spans 4 domains and 5 benchmarks totaling 2,207 instances: SWE-bench Pro public split (731), Terminal-Bench 2.1 (89), MLE-bench (75), EQ-Bench3 (46), and BrowseComp (1,266). Each creator builds 3 harnesses per benchmark, reported as avg@3. Self-Eval runs each harness with its creator; Unified-Eval runs all with Gemini 3.1 Pro. Creation results Under Self-Eval, Opus 4.8 posts the highest average score at 67.8 against a human-engineered reference of 86.2. The gap depends on domain: Code: Opus 4.8 reaches 69.3 on SWE-Pro versus the 80.0 reference. Gemini 3.1 Pro leads Terminal-Bench at 68.8 versus 88.8. Search: the widest gap. The best BrowseComp score is 52.6 (GPT-5.5) against a 92.2 reference. Writing: Opus 4.8 scores 84.6 on EQ-Bench3, above the 83.7 reference. ML experimentation: Opus 4.8 (32.9) and Gemini (32.4) beat the 24.0 MLE-bench reference. The SWE-Pro, Terminal-Bench, and BrowseComp references are external results from OpenAI\u2019s GPT-5.6 report, not re-runs. Code volume did not predict quality: the 18 code harnesses added 17,111 net lines, yet Gemini added the fewest (1,006) and led Terminal-Bench. Self-test count barely correlated with score (Spearman 0.13 to 0.26); revision calls reached 0.57. Much generated machinery is inert. Of 108 code component instances, 72 trigger in real runs and 18 never fire, all of them state and memory. 11 of 18 harnesses define a State class, yet no checkpoint event appears across 26,679 trajectories. 124 of 587 writing features are dead code. Cost and executor transfer MLE-bench token use varied roughly 19-fold. GPT-5.5 hit a 19.1 medal rate with 29.3M tokens while DeepSeek V4 hit 19.6 with 208.4M. Swapping the executor to Gemini reshuffled rankings: Qwen gained 17.6 points on BrowseComp and 12.9 on MLE-bench, while Opus 4.8\u2019s SWE-Pro score fell from 69.3 to 33.0, partly because one harness hard-coded a 120-step limit around its original executor. The Opus search harness\u2019s duplicate-query rate jumped from 10.1% to 88.2% after the switch. Evolution results 9 lineages (5 self-runtime, 4 fixed-Gemini) produced 73 official versions and 64 adjacent switches. All 5 self-runtime creators improved on held-out tasks, from +1.43 to +4.44 points (mean +3.11). Under fixed Gemini, only Opus improved; GPT-5.5 regressed 10.32 points. Progress was not monotonic. Of 64 switches, 8 regressed on both benchmarks, 16 on one, 27 gained only within the noise band, and 2 showed clear positive evidence. A single commit can vary by about \u00b14.75 pair-score points. Feedback and held-out scores moved in the same direction only 34 of 64 times (53.1%), and only 2 of 9 declared final versions were held-out optimal. Of 169 new functions or classes, 25 have no caller. The clearest win: Opus 4.8 noticed 99 of 100 runs reported success while only 48 passed, traced it to premature completion, and added a completion gate. Failure diagnosis was otherwise the weakest step: the dedicated trajectory interface was called only twice. Interactive explainer Key Takeaways HarnessDev scores the harness a model builds, not the answer it returns. Self-built harnesses match or beat references on writing and ML experimentation but trail badly on code and search. Harness quality is executor-specific; Opus 4.8 drops from 69.3 to 33.0 on SWE-Pro under Gemini. Evolution gains are small, noisy, and only 34 of 64 changes point the same way on held-out tasks. Much generated state and memory code never executes. Check out the\u00a0Paper and Project Page. Also,\u00a0feel free to follow us on\u00a0Twitter\u00a0and don\u2019t forget to join our\u00a0150k+ML SubReddit\u00a0and Subscribe to\u00a0our Newsletter. Wait! are you on telegram?\u00a0now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?\u00a0Connect with us The post Can LLMs Engineer Their Own Agent Harness? ByteDance Seed\u2019s HarnessDev Says Only 34 of 64 Changes Generalize appeared first on MarkTechPost.<\/p>","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"pmpro_default_level":"","site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"_pvb_checkbox_block_on_post":false,"footnotes":""},"categories":[52,5,7,1],"tags":[],"class_list":["post-117261","post","type-post","status-publish","format-standard","hentry","category-ai-club","category-committee","category-news","category-uncategorized","pmpro-has-access"],"acf":[],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v25.3 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>Can LLMs Engineer Their Own Agent Harness? ByteDance Seed\u2019s HarnessDev Says Only 34 of 64 Changes Generalize - YouZum<\/title>\n<meta name=\"description\" content=\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/youzum.net\/fr\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/\" \/>\n<meta property=\"og:locale\" content=\"fr_FR\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Can LLMs Engineer Their Own Agent Harness? ByteDance Seed\u2019s HarnessDev Says Only 34 of 64 Changes Generalize - YouZum\" \/>\n<meta property=\"og:description\" content=\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\" \/>\n<meta property=\"og:url\" content=\"https:\/\/youzum.net\/fr\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/\" \/>\n<meta property=\"og:site_name\" content=\"YouZum\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/DroneAssociationTH\/\" \/>\n<meta property=\"article:published_time\" content=\"2026-09-12T01:44:24+00:00\" \/>\n<meta name=\"author\" content=\"admin NU\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"\u00c9crit par\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin NU\" \/>\n\t<meta name=\"twitter:label2\" content=\"Dur\u00e9e de lecture estim\u00e9e\" \/>\n\t<meta name=\"twitter:data2\" content=\"4 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/youzum.net\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/#article\",\"isPartOf\":{\"@id\":\"https:\/\/youzum.net\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/\"},\"author\":{\"name\":\"admin NU\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c\"},\"headline\":\"Can LLMs Engineer Their Own Agent Harness? ByteDance Seed\u2019s HarnessDev Says Only 34 of 64 Changes Generalize\",\"datePublished\":\"2026-09-12T01:44:24+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/youzum.net\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/\"},\"wordCount\":817,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\"},\"articleSection\":[\"AI\",\"Committee\",\"News\",\"Uncategorized\"],\"inLanguage\":\"fr-FR\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/youzum.net\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/youzum.net\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/\",\"url\":\"https:\/\/youzum.net\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/\",\"name\":\"Can LLMs Engineer Their Own Agent Harness? ByteDance Seed\u2019s HarnessDev Says Only 34 of 64 Changes Generalize - YouZum\",\"isPartOf\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#website\"},\"datePublished\":\"2026-09-12T01:44:24+00:00\",\"description\":\"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19\",\"breadcrumb\":{\"@id\":\"https:\/\/youzum.net\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/#breadcrumb\"},\"inLanguage\":\"fr-FR\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/youzum.net\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/\"]}]},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/youzum.net\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/youzum.net\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Can LLMs Engineer Their Own Agent Harness? ByteDance Seed\u2019s HarnessDev Says Only 34 of 64 Changes Generalize\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/yousum.gpucore.co\/#website\",\"url\":\"https:\/\/yousum.gpucore.co\/\",\"name\":\"YouSum\",\"description\":\"\",\"publisher\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/yousum.gpucore.co\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"fr-FR\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/yousum.gpucore.co\/#organization\",\"name\":\"Drone Association Thailand\",\"url\":\"https:\/\/yousum.gpucore.co\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"fr-FR\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png\",\"contentUrl\":\"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png\",\"width\":300,\"height\":300,\"caption\":\"Drone Association Thailand\"},\"image\":{\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/\"},\"sameAs\":[\"https:\/\/www.facebook.com\/DroneAssociationTH\/\"]},{\"@type\":\"Person\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c\",\"name\":\"admin NU\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"fr-FR\",\"@id\":\"https:\/\/yousum.gpucore.co\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png\",\"contentUrl\":\"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png\",\"caption\":\"admin NU\"},\"url\":\"https:\/\/youzum.net\/fr\/members\/adminnu\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Can LLMs Engineer Their Own Agent Harness? ByteDance Seed\u2019s HarnessDev Says Only 34 of 64 Changes Generalize - YouZum","description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/youzum.net\/fr\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/","og_locale":"fr_FR","og_type":"article","og_title":"Can LLMs Engineer Their Own Agent Harness? ByteDance Seed\u2019s HarnessDev Says Only 34 of 64 Changes Generalize - YouZum","og_description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","og_url":"https:\/\/youzum.net\/fr\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/","og_site_name":"YouZum","article_publisher":"https:\/\/www.facebook.com\/DroneAssociationTH\/","article_published_time":"2026-09-12T01:44:24+00:00","author":"admin NU","twitter_card":"summary_large_image","twitter_misc":{"\u00c9crit par":"admin NU","Dur\u00e9e de lecture estim\u00e9e":"4 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/youzum.net\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/#article","isPartOf":{"@id":"https:\/\/youzum.net\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/"},"author":{"name":"admin NU","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c"},"headline":"Can LLMs Engineer Their Own Agent Harness? ByteDance Seed\u2019s HarnessDev Says Only 34 of 64 Changes Generalize","datePublished":"2026-09-12T01:44:24+00:00","mainEntityOfPage":{"@id":"https:\/\/youzum.net\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/"},"wordCount":817,"commentCount":0,"publisher":{"@id":"https:\/\/yousum.gpucore.co\/#organization"},"articleSection":["AI","Committee","News","Uncategorized"],"inLanguage":"fr-FR","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/youzum.net\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/youzum.net\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/","url":"https:\/\/youzum.net\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/","name":"Can LLMs Engineer Their Own Agent Harness? ByteDance Seed\u2019s HarnessDev Says Only 34 of 64 Changes Generalize - YouZum","isPartOf":{"@id":"https:\/\/yousum.gpucore.co\/#website"},"datePublished":"2026-09-12T01:44:24+00:00","description":"\u0e01\u0e34\u0e08\u0e01\u0e23\u0e23\u0e21\u0e40\u0e01\u0e35\u0e48\u0e22\u0e27\u0e01\u0e31\u0e1a\u0e42\u0e14\u0e23\u0e19","breadcrumb":{"@id":"https:\/\/youzum.net\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/#breadcrumb"},"inLanguage":"fr-FR","potentialAction":[{"@type":"ReadAction","target":["https:\/\/youzum.net\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/"]}]},{"@type":"BreadcrumbList","@id":"https:\/\/youzum.net\/can-llms-engineer-their-own-agent-harness-bytedance-seeds-harnessdev-says-only-34-of-64-changes-generalize\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/youzum.net\/"},{"@type":"ListItem","position":2,"name":"Can LLMs Engineer Their Own Agent Harness? ByteDance Seed\u2019s HarnessDev Says Only 34 of 64 Changes Generalize"}]},{"@type":"WebSite","@id":"https:\/\/yousum.gpucore.co\/#website","url":"https:\/\/yousum.gpucore.co\/","name":"YouSum","description":"","publisher":{"@id":"https:\/\/yousum.gpucore.co\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/yousum.gpucore.co\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"fr-FR"},{"@type":"Organization","@id":"https:\/\/yousum.gpucore.co\/#organization","name":"Drone Association Thailand","url":"https:\/\/yousum.gpucore.co\/","logo":{"@type":"ImageObject","inLanguage":"fr-FR","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/","url":"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png","contentUrl":"https:\/\/youzum.net\/wp-content\/uploads\/2024\/11\/tranparent-logo.png","width":300,"height":300,"caption":"Drone Association Thailand"},"image":{"@id":"https:\/\/yousum.gpucore.co\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/DroneAssociationTH\/"]},{"@type":"Person","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/97fa48242daf3908e4d9a5f26f4a059c","name":"admin NU","image":{"@type":"ImageObject","inLanguage":"fr-FR","@id":"https:\/\/yousum.gpucore.co\/#\/schema\/person\/image\/","url":"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png","contentUrl":"https:\/\/youzum.net\/wp-content\/uploads\/avatars\/2\/1746849356-bpfull.png","caption":"admin NU"},"url":"https:\/\/youzum.net\/fr\/members\/adminnu\/"}]}},"rttpg_featured_image_url":null,"rttpg_author":{"display_name":"admin NU","author_link":"https:\/\/youzum.net\/fr\/members\/adminnu\/"},"rttpg_comment":0,"rttpg_category":"<a href=\"https:\/\/youzum.net\/fr\/category\/ai-club\/\" rel=\"category tag\">AI<\/a> <a href=\"https:\/\/youzum.net\/fr\/category\/committee\/\" rel=\"category tag\">Committee<\/a> <a href=\"https:\/\/youzum.net\/fr\/category\/news\/\" rel=\"category tag\">News<\/a> <a href=\"https:\/\/youzum.net\/fr\/category\/uncategorized\/\" rel=\"category tag\">Uncategorized<\/a>","rttpg_excerpt":"An agent harness is the code around a model: execution loop, tools, context, state, recovery, and verification. Per the Terminal-Bench 2.1 leaderboard, GPT-5 solves 35.2% of tasks inside Terminus 2 but 49.6% inside Codex CLI with identical weights. Most benchmarks keep that harness fixed. HarnessDev proposed by team of researchers from ByteDance Seed, Singapore University\u2026","_links":{"self":[{"href":"https:\/\/youzum.net\/fr\/wp-json\/wp\/v2\/posts\/117261","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/youzum.net\/fr\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/youzum.net\/fr\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/youzum.net\/fr\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/youzum.net\/fr\/wp-json\/wp\/v2\/comments?post=117261"}],"version-history":[{"count":0,"href":"https:\/\/youzum.net\/fr\/wp-json\/wp\/v2\/posts\/117261\/revisions"}],"wp:attachment":[{"href":"https:\/\/youzum.net\/fr\/wp-json\/wp\/v2\/media?parent=117261"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/youzum.net\/fr\/wp-json\/wp\/v2\/categories?post=117261"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/youzum.net\/fr\/wp-json\/wp\/v2\/tags?post=117261"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}