YouZum

Uncategorized

AI, Committee, Noticias, Uncategorized

The Download: biotech’s future and cheaper, cleaner steel

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Meet the under-35s shaping the future of biotech Every year, MIT Technology Review puts together our 35 Innovators Under 35, a list of some of the brightest and best young minds working across science and technology. This year’s honorees include nine people transforming biotech, whose work spans everything from lifesaving innovations to groundbreaking longevity tech. Their innovations include a “reprogramming” therapy that reverses vision loss, tiny brain electrodes inspired by Japanese art, and a personalized gene-editing treatment for a baby with a rare genetic disorder. There are even efforts to design new viruses with generative AI, which (hopefully) will produce new drugs or soak up pollution. Get to know the biotech innovators behind these breakthroughs. —Jessica Hamzelou This story is from The Checkup, our weekly biotech newsletter. Sign up to receive it in your inbox every Thursday. Biotechnology is one of four categories in our 35 Innovators Under 35 list for 2026, featuring young people worldwide doing groundbreaking work in science and technology. Meet the rest of them here, or explore the full list across the AI, computing and robotics, biotechnology, and climate and energy categories. This founder is making cheaper, cleaner steel The steel industry isn’t exactly known for innovation. Very little has changed about purifying iron ore since the process was invented and commercialized in the 1850s. But Laureen Meroueh, founder of Hertha Metals, has an idea that could change that. Meroueh may have found a way to clean up steelmaking without driving up the price. Her new furnace turns iron ore into refined liquid steel in a single step and swaps coal for natural gas. Together, those changes slash emissions by at least half, she says, and cut costs by 25% compared with steelmaking as usual. Here’s how she plans to make steel cleaner without making it more expensive. —Bridget Reed Morawski Laureen Meroueh is one of the climate change and energy honorees on our 35 Innovators Under 35 list. The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 Anthropic says it has blocked potential plots to build biological weaponsThe company identified five such cases. (NYT $)+ And six cases of using AI to build software for conventional weapons (BBC)+ Governments are also using Claude for surveillance.(Axios)+ While Russia-linked hackers used it to automate attacks on Ukraine. (Quartz)+ The threats were revealed in a new Anthropic report. (Guardian)+ Bill Gates says AI needs new guardrails. (MIT Technology Review) 2 California has banned addictive social media features for under-16sThe law prohibits infinite scroll and autoplay. (Guardian)+ It also introduces new rules for AI and companion chatbots. (Reuters $)+ It’s the first law of its kind in the US. (NYT $)+ Social media encourages the worst AI boosterism. (MIT Technology Review) 3 Two AI researchers have left Anthropic and Google over safety risksThey left a day after Jacob Coxon’s viral departure from Anthropic. (NBC News)+ Elon Musk called their concerns a “setup” and a “psyop.” (Guardian)+ AI fears are pushing Congress toward tougher regulation. (WSJ $) 4 Sam Altman is pitching OpenAI’s cyber defenses to power companiesThe meetings followed reports of AI attacks on critical systems. (Politico $)+ Altman also told staff that OpenAI is open to slowing down AI. Bloomberg $) 5 After years of fighting AI, music labels are starting to embrace itUniversal is partnering with ElevenLabs on an AI remix platform.(Gizmodo)+ AI is complicating definitions of creativity. (MIT Technology Review) 6 Chinese drugmakers are challenging US dominance in weight-loss drugsThey’re developing hundreds of GLP-1 treatments for global markets. (WSJ $) 7 Electric air taxis have begun official test flights in TexasThey’re the first flights under the White House’s new pilot program. (Verge) 8 Chinese drones are helping to rescue survivors of Nepal’s floodsThey’re delivering food and airlifting bodies from flood-hit areas. (Ars Technica) 9 NASA and IBM have built an AI model to map the moonIt could help locate ice and identify safer landing sites. (Register) 10 One man is on a quest to digitally preserve America’s public restroomsHis Restroom Archive is a museum-style repository of 3D scans. (404 Media) Quote of the day “I didn’t ask Facebook to build a profile of my family—I posted a video of me singing in the car with my kids.”  —Kalie Roberts, a travel content creator, says in an Instagram reel that Meta AI used years of Facebook posts to piece together her children’s identities and pinpoint where her family lives. One more thing Chinese tech workers are starting to train their AI doubles—and pushing back In April, a GitHub project called Colleague Skill struck a nerve by claiming to “distill” a worker’s skills and personality—and replicate them with an AI agent. Though the project was a spoof, it prompted a wave of soul-searching among otherwise enthusiastic early adopters. A number of tech workers told MIT Technology Review that their bosses are already encouraging them to document their workflows for automation via tools like OpenClaw. Many now fear that they are being flattened into code and losing their professional identity. In response, some are fighting back with tools designed to sabotage the automation process. Read the full story on their battle with clone workers. —Caiwei Chen We can still have nice things A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.) + Worried about Flock cameras? These guys designed a car to fool them.+ Webb’s Near-Infrared Camera has captured a galactic merger’s dazzling final phase.+ An exquisitely preserved 66-million-year-old bird feather was found in a fossilised dinosaur dropping.+ A plucky preservationist travelled 1,700 miles and made 52 calls from a rare phone box to keep it in service.

The Download: biotech’s future and cheaper, cleaner steel Leer entrada »

AI, Committee, Noticias, Uncategorized

Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills

Anthropic has published a new plugin evals workflow for Claude Code. The claude plugin eval command runs a plugin against realistic prompts, grades what Claude produced, and compares the result with a run where the plugin is not loaded. It answers 3 questions plugin developers could not previously measure: does the skill trigger, does it survive an edit or a new model, and does it beat a bare model. Deployable: Yes. It runs on Claude Code v2.1.269 or later against any directory with a plugin.json or .claude-plugin/plugin.json manifest, or a skills-directory plugin. Every eval run and judge grader is a real model call billed to your plan or API account. What a case looks like An eval suite lives in an evals/ directory inside the plugin. Each case is a subdirectory holding a prompt.md and a graders/ folder. The prompt body goes to Claude exactly as written, and @path mentions are not expanded. Frontmatter on prompt.md can set max_turns (default 10), timeout_seconds (default 300), model, tags, and allowed_tools. Graders are markdown files whose frontmatter sets a type, an optional weight, and an optional arm. There are 6 types. Four cost nothing because they are computed from the transcript and the files on disk: regex, tool_used, tool_order, and file_exists. Two call a judge model and add to the bill: llm, which scores the reply against prose criteria you write, and baseline, which compares it against a reference answer. claude plugin eval init reads the plugin, asks what a good result looks like, proposes cases and graders, tries them, and writes the files. In CI, –bare <name> writes a blank template instead. The number that matters is Δ By default every case runs twice: a with-arm where the plugin is loaded and a without-arm where it is not. Their difference, Δ, is what the plugin contributed. If a case scores 1.0 in both arms, the plugin is not why it passed. The docs example output shows a single case at WITH 1.00, W/OUT 0.33, Δ +0.67 across 6 runs, costing an estimated $0.41 and taking 74 seconds. A grader marked with-only, typically tool_used: Skill, is reported as an indicator and excluded from the score, since the without-arm has no skill to fire. Anthropic calls out the most common first finding: a Δ near zero with the tool_used: Skill grader failing, which means Claude is not choosing the skill on natural phrasing. That is the defect claude plugin validate cannot see, because it checks manifest syntax and schema rather than behavior. Results land under evals/results/<timestamp>/report.html with per-grader verdicts and judge votes. Where the account supports it, the report is also published to claude.ai unless –no-publish is set. Cost and CI A suite makes roughly cases × runs × arms agent runs, plus 3 short judge calls per llm or baseline grader per run, and results vary between runs. The documented CI invocation is: Copy CodeCopiedUse a different Browser claude plugin eval . –trust-plugin –json results.json –threshold 0.8 –model claude-sonnet-5 –judge-model claude-haiku-4-5 –no-publish –max-cost-usd 20 The runner needs a Claude Code install and credentials such as ANTHROPIC_API_KEY. Without –trust-plugin, an untrusted checkout is refused with exit 1 when there is no terminal. Report problems never change the exit code, and –json suppresses progress output. Interactive explainer Key Takeaways claude plugin eval scores realistic prompts with 6 grader types; 4 are free, llm and baseline bill a judge model. Every case runs with and without the plugin by default; Δ is the only number that proves the plugin did the work. A Δ near zero with a failing tool_used: Skill grader means the skill never triggers on natural phrasing. –threshold, –max-cost-usd, and –trust-plugin turn it into a CI gate; usage-limit errors can fake a regression. Requires Claude Code v2.1.269+; claude plugin eval init writes the first suite for you. Check out the Technical details. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills appeared first on MarkTechPost.

Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills Leer entrada »

AI, Committee, Noticias, Uncategorized

Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize

An agent harness is the code around a model: execution loop, tools, context, state, recovery, and verification. Per the Terminal-Bench 2.1 leaderboard, GPT-5 solves 35.2% of tasks inside Terminus 2 but 49.6% inside Codex CLI with identical weights. Most benchmarks keep that harness fixed. HarnessDev proposed by team of researchers from ByteDance Seed, Singapore University of Technology and Design, Georgia Institute of Technology, M-A-P, and TokenWave.AI, flips the target: the artifact under evaluation is the runnable harness the model writes, not the answer it produces. 2 stages: Creation and Evolution In Creation, every creator receives the same weak seed: passive file, search, and process primitives plus result and trajectory writers, with no loop, planner, verifier, retry, or stopping rule. Unmodified, it scores 0 everywhere. The creator gets a task-family spec, a short design tutorial, and 1 to 3 development cases, builds a full harness, and the harness is frozen before hidden tasks. In Evolution, the creator starts from its own frozen Creation code harness and revises it using execution feedback from a fixed set of 100 SWE-bench Pro tasks and all 89 Terminal-Bench 2.1 tasks. Each official candidate must complete both evaluations as a pair, with a budget of 10 pairs and at most 2 five-task probes between pairs. Every official version is later scored on 630 held-out SWE-Pro instances the creator never sees. Harnesses are graded on capability (task success) and efficiency (executor tokens, with creator tokens excluded). Setup 6 creator LLMs were tested: Opus 4.8, GPT-5.5, Gemini 3.1 Pro, DeepSeek V4 Pro, Qwen 3.7 Max, and Seed 2.0 Pro, working inside Claude Code 2.1.177 (GPT-5.5 used Codex 0.144.3). Creation spans 4 domains and 5 benchmarks totaling 2,207 instances: SWE-bench Pro public split (731), Terminal-Bench 2.1 (89), MLE-bench (75), EQ-Bench3 (46), and BrowseComp (1,266). Each creator builds 3 harnesses per benchmark, reported as avg@3. Self-Eval runs each harness with its creator; Unified-Eval runs all with Gemini 3.1 Pro. Creation results Under Self-Eval, Opus 4.8 posts the highest average score at 67.8 against a human-engineered reference of 86.2. The gap depends on domain: Code: Opus 4.8 reaches 69.3 on SWE-Pro versus the 80.0 reference. Gemini 3.1 Pro leads Terminal-Bench at 68.8 versus 88.8. Search: the widest gap. The best BrowseComp score is 52.6 (GPT-5.5) against a 92.2 reference. Writing: Opus 4.8 scores 84.6 on EQ-Bench3, above the 83.7 reference. ML experimentation: Opus 4.8 (32.9) and Gemini (32.4) beat the 24.0 MLE-bench reference. The SWE-Pro, Terminal-Bench, and BrowseComp references are external results from OpenAI’s GPT-5.6 report, not re-runs. Code volume did not predict quality: the 18 code harnesses added 17,111 net lines, yet Gemini added the fewest (1,006) and led Terminal-Bench. Self-test count barely correlated with score (Spearman 0.13 to 0.26); revision calls reached 0.57. Much generated machinery is inert. Of 108 code component instances, 72 trigger in real runs and 18 never fire, all of them state and memory. 11 of 18 harnesses define a State class, yet no checkpoint event appears across 26,679 trajectories. 124 of 587 writing features are dead code. Cost and executor transfer MLE-bench token use varied roughly 19-fold. GPT-5.5 hit a 19.1 medal rate with 29.3M tokens while DeepSeek V4 hit 19.6 with 208.4M. Swapping the executor to Gemini reshuffled rankings: Qwen gained 17.6 points on BrowseComp and 12.9 on MLE-bench, while Opus 4.8’s SWE-Pro score fell from 69.3 to 33.0, partly because one harness hard-coded a 120-step limit around its original executor. The Opus search harness’s duplicate-query rate jumped from 10.1% to 88.2% after the switch. Evolution results 9 lineages (5 self-runtime, 4 fixed-Gemini) produced 73 official versions and 64 adjacent switches. All 5 self-runtime creators improved on held-out tasks, from +1.43 to +4.44 points (mean +3.11). Under fixed Gemini, only Opus improved; GPT-5.5 regressed 10.32 points. Progress was not monotonic. Of 64 switches, 8 regressed on both benchmarks, 16 on one, 27 gained only within the noise band, and 2 showed clear positive evidence. A single commit can vary by about ±4.75 pair-score points. Feedback and held-out scores moved in the same direction only 34 of 64 times (53.1%), and only 2 of 9 declared final versions were held-out optimal. Of 169 new functions or classes, 25 have no caller. The clearest win: Opus 4.8 noticed 99 of 100 runs reported success while only 48 passed, traced it to premature completion, and added a completion gate. Failure diagnosis was otherwise the weakest step: the dedicated trajectory interface was called only twice. Interactive explainer Key Takeaways HarnessDev scores the harness a model builds, not the answer it returns. Self-built harnesses match or beat references on writing and ML experimentation but trail badly on code and search. Harness quality is executor-specific; Opus 4.8 drops from 69.3 to 33.0 on SWE-Pro under Gemini. Evolution gains are small, noisy, and only 34 of 64 changes point the same way on held-out tasks. Much generated state and memory code never executes. Check out the Paper and Project Page. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize appeared first on MarkTechPost.

Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize Leer entrada »

AI, Committee, Noticias, Uncategorized

The Download: a “God-driven” cryptocurrency and a solar engineering roadmap

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. God told them to sell crypto. Their investors lost everything. When Eli Regalado first heard God speak to him, he wondered whether he was hallucinating. According to Eli and his wife, Kaitlyn, He told them to get married, buy a house, and start having kids. Then in 2021, divine guidance steered them in an unexpected new direction: crypto. That October, the Regalados later testified in court, they received holdings in a little-known digital coin. “Take this to my people for a wealth transfer,” Eli heard God say. Over time, they came to believe that He wanted them to launch their own coin. The Regalados created INDXcoin, which they promoted through family, friends, and contacts in evangelical Christian circles. In all, more than 500 people handed over more than $3 million. But within a year, the project collapsed. Investors lost it all, leaving many to wonder where the funds went and whether they had fallen victim to an elaborate fraud. Read the full story on the collapse of a pastor’s “God-driven” cryptocurrency. —Katia Savchuk This article is part of the Big Story series, the home of MIT Technology Review’s most important and ambitious reporting. You can read the rest of the series here.  The story was produced in partnership with Type Investigations and with support from the Fund for Investigative Journalism. This road map could help us decide whether to deploy solar geoengineering Scientists have spent half a century exploring whether we could counteract climate change by releasing reflective particles into the stratosphere, mimicking the cooling effects of volcanic eruptions. But even after hundreds of studies, we still don’t know how well it would work or what else it might do—and there’s no systematic plan for clearing up that uncertainty. Reflective, a research organization, has now attempted to fill that gap. The San Francisco nonprofit has published a detailed road map of the experiments, studies, and infrastructure that it says would be needed to make informed decisions about the use of solar geoengineering, MIT Technology Review can reveal. Find out what it would take to make informed decisions about solar geoengineering. —James Temple This founder is teaching chips how to recycle (their energy) Throughout the history of the computer chip, engineers have treated waste heat as an inevitable cost of a calculation. Hannah Earley, however, thinks it’s a design choice. Earley, 31, is cofounder and CTO of Vaire Computing, which builds chips that recycle energy usually thrown away as heat, a strategy known as reversible computing. The approach could make data centers (and our laptops and phones) much more energy efficient. Last year, Vaire announced a key breakthrough: a chip with a resonator that recovered more energy than it lost, even after the energy needed to power the component was taken into account. Here’s how she plans to bring an old idea about energy-efficient computers into the future. —Eshan Raul Hannah Earley is one of the computing and robotics honorees on our 35 Innovators Under 35 list for 2026. Meet the rest of them here, or explore the full list across the biotechnology, AI, computing and robotics, and climate and energy categories. Can the US battery market untangle from China? —Casey Crownhart The US energy storage market is growing at a record pace, which could shore up the grid and cut emissions. Crucially, this is all happening with the help of cheap Chinese batteries, which the Trump administration is trying to phase out. Reducing reliance on any single source of crucial energy technology makes sense. But the tension raises a broader question for me: how much should countries take advantage of cheap, available tech, and how much should they cut themselves off from foreign sources to develop their own, even if it costs more? Dive into the difficult choices facing America’s booming battery market. This story is from The Spark, our weekly climate tech newsletter. Sign up to receive it in your inbox every Wednesday. The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 OpenAI’s agents used at least 10 websites for unauthorized communicationsResearchers found they bypassed restrictions on posting online.(Reuters $)+ The company faces a Senate probe into the Hugging Face breach. (Axios)+ Its hacking issues may indicate cultural problems. (MIT Technology Review) 2 Another Anthropic model hacked a real system during testingA misconfigured environment gave it internet access. (CBS News)+ The January incident went undetected until last month. (Reuters $)+ AI agents are not your “coworkers.” (MIT Technology Review) 3 Apple has entered the foldable phone market with the $1,999 iPhone DuoIt opens into a 7.6-inch display and launches October 23. (NPR)+ Apple is betting its design and privacy will give it an edge. (Reuters $)+ And that foldables can solve the smartphone’s sameness problem. (NPR $)+ Samsung responded with a campaign touting its foldable lead. (CNBC)+ In China, Apple enters a crowded market dominated by Huawei. (SCMP) 4 US prosecutors have called Huawei a criminal enterprise at trialThey accuse the company of stealing American technology. (Reuters $)+ And helping Iran snoop on its citizens. (AP News)+ The trial could impact Trump’s upcoming meeting with Xi. (WSJ $) 5 California is warming to nuclear power after decades of oppositionThe state may extend Diablo Canyon and lift its ban on new reactors. (NYT $)+ China is betting on big nuclear reactors. (MIT Technology Review) 6 Chinese professionals are becoming gig workers training AILawyers and engineers are training models for extra income. (Rest of World)+ Gig workers are training humanoids at home. (MIT Technology Review) 7 The new Apple Watch can listen to conversations happening nearbyApple says users must opt in, but others cannot. (Wired $) 8 Pink noise during sleep could help the brain clear away wasteTimed bursts boosted brain fluid flow in a small study. (New Scientist $) 9 A lost supercontinent may have triggered the explosion of

The Download: a “God-driven” cryptocurrency and a solar engineering roadmap Leer entrada »

AI, Committee, Noticias, Uncategorized

NVIDIA Details BioNeMo Inference Runtime (BioIR): 2.90x Higher Boltz-2 Folding Throughput and 58.5K Residues per GPU-Hour on 8xH100

Biomolecular structure prediction has shifted from single-target runs to proteome-scale worklists. The bottleneck is no longer whether a model can fold a protein. It is how fast an entire queue of independent targets moves through parsing, featurization, GPU inference, and output writing. NVIDIA’s new technical deep dive walks through BioNeMo Inference Runtime (BioIR), a Python library that accelerates supported structure-prediction models on NVIDIA GPUs while keeping the standard PyTorch workflow. BioIR has already run at production scale. It powered the recent expansion of the AlphaFold Database, generating protein-complex structures across 4,777 proteomes, about 31 million candidate complexes, with 1.81 million released as high-confidence predictions. Is it deployable? Yes. BioIR is available now as an open GitHub repository with a wheel containing precompiled CUBINs. Runtime use needs Python 3.12+, a compatible NVIDIA GPU and driver, a staged model checkpoint, and per-chain A3M MSAs. It does not require nvcc, CUDA source, CMake, or the CUDA toolkit. What is BioIR BioIR targets the operations that general-purpose inference stacks do not fully optimize. These include Pairformer and Evoformer stacks, triangle operations, pairwise attention, diffusion transformers, and atom-level modules. Models stay ordinary torch.nn.Module objects. There is no engine build, export step, or separate artifact between a checkpoint and a forward pass. There are 2 ways to use it. The end-to-end processor moves an InputRequest through parsing, tokenization, feature generation, GPU inference, and PDB or mmCIF writing. Direct PyTorch integration lets developers construct a supported model or reuse selected optimized modules inside custom code. The tutorial demonstrates the processor path with Boltz-2 (model_source=”boltz-2″). Each protein chain requires an A3M MSA. Paired or unpaired MSAs are accepted for inputs with multiple non-identical protein chains. Templates can be supplied manually because BioIR does not run HHsearch or HMMsearch. The processor supports ligand structure prediction but not ligand-affinity prediction. Three Layers of Acceleration BioIR optimizes at 3 distinct layers, each targeting a different bottleneck: Kernel selection: Supported operations pick compatible BioIR custom, cuEquivariance, or PyTorch fallback implementations based on model configuration, GPU, data type, and tensor shape. Module optimization: A separate optimize() mechanism enables CUDA Graph capture for compatible modules, cutting launch overhead. Pipeline scaling: A Ray executor places 1 complete model replica on each visible GPU in a node and distributes independent inputs among them. CPU stages (parsing, featurization, writing) overlap with GPU folding. Note: Ray does not split a single forward pass across GPUs. Replica mode scales worklists, not individual targets. Per the support matrix, context-parallel folding is planned but not yet available. The capacity rule is simple: engine_stage.compute x num_gpus must not exceed visible GPUs. At the model-forward level, NVIDIA’s early benchmarking reports geometric-mean speedups over an OSS torch.compile baseline of 1.55x (OpenFold3), 1.78x (Boltz2), and 2.56x (OpenFold2 monomer) on H100. H200 numbers are similar at 1.54x, 1.75x, and 2.61x. These were measured across 17 inputs spanning 29 to 1,734 residues. The Benchmark: 1,000 Human Dimers on 8xH100 To quantify end-to-end delivery, NVIDIA team ran a matched benchmark on 1,000 human dimer targets with combined sequence lengths below 2,800 residues. The comparison pitted BioIR-accelerated Boltz-2 against a torch-compiled open-source Boltz-2 implementation on 8xH100 80GB GPUs. Both used identical targets, staged MSAs, inference recipe (3 recycles, 200 sampling steps, 5 diffusion samples), and GPU configuration. The results: BioIR completed all 1,000 targets and delivered 58.5K successfully folded residues per allocated GPU-hour. The public implementation delivered 20.2K residues per GPU-hour and ran out of memory on 29 targets. Net result: a 2.90x improvement in residue-normalized throughput. These numbers are folding-stage measurements specific to this dataset and hardware. They exclude MSA generation, preprocessing CPU allocations, storage, data transfer, and retries. The blog explicitly warns against generalizing them to all BioIR-supported models or datasets. Energy at One Million Targets Extrapolating the benchmark linearly to 1 million comparable targets, BioIR is estimated to need 11 MWh versus 35 MWh for the public implementation using 8-GPU TDP equivalents. Using full-node maximum-power equivalents, the estimate is 21 MWh versus 64 MWh. These are rated-power, folding-only estimates for IT equipment, not metered measurements, and exclude data center overhead such as PUE. Still, a 23 to 43 MWh saving per million targets is a material number for proteome-scale campaigns. Key Takeaways BioIR accelerates Boltz-2, OpenFold2, and OpenFold3 inference on NVIDIA GPUs while staying in plain PyTorch. Matched 8xH100 benchmark: 58.5K vs 20.2K folded residues per GPU-hour, a 2.90x throughput gain. Ray replica mode scales independent worklists; it never splits 1 forward pass across GPUs. Estimated energy for 1M targets drops from 35 MWh to 11 MWh at 8-GPU TDP equivalents. Already proven at scale: 31M candidate complexes generated for the AlphaFold Database expansion. Check out the technical blog, GitHub repo, docs, and the BioNeMo Agent Toolkit for agentic orchestration. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post NVIDIA Details BioNeMo Inference Runtime (BioIR): 2.90x Higher Boltz-2 Folding Throughput and 58.5K Residues per GPU-Hour on 8xH100 appeared first on MarkTechPost.

NVIDIA Details BioNeMo Inference Runtime (BioIR): 2.90x Higher Boltz-2 Folding Throughput and 58.5K Residues per GPU-Hour on 8xH100 Leer entrada »

AI, Committee, Noticias, Uncategorized

OpenAI Launches the Agents API in Public Beta, Putting the Codex Harness Behind One API Call

OpenAI has released the Agents API in public beta. It gives developers the same harness and infrastructure that run Codex. OpenAI hosts and maintains the harness. Developers run the agent’s compute in an OpenAI-managed sandbox, their own infrastructure, or a partner sandbox. Is it deployable? Yes. It is live for all developers in public beta. Data stays US-only, and Zero Data Retention is unsupported. What OpenAI Shipped The Agents API is a managed service built on the open-source Codex harness. OpenAI team states scaling Codex and ChatGPT for Work showed what long-running agents need. They need a harness that manages context, uses tools efficiently, and coordinates subagents. They also need infrastructure that keeps them running reliably for days. The official docs organize the API around 4 concepts: Agent: the model, instructions, tools, and MCP servers available to it. Environment: an optional sandbox where the agent accesses files, loads skills, and runs commands. Session: a durable agent instance that works on tasks and responds to input. Events and items: the inputs sent to the agent and the output it produces. A session runs in 4 steps. You create it and give it a task. Then you follow progress through streaming or webhooks. Finally, you continue with a new task or steer the current turn. One API Call OpenAI’s announcement shows an incident-investigation agent created in a single call: Copy CodeCopiedUse a different Browser import OpenAI from “openai”; const client = new OpenAI(); const session = await client.beta.agents.sessions.create({ agent: { model: “gpt-6-astra”, tools: [ { type: “mcp”, server_label: “observability”, transport: { type: “http”, server_url: “https://observability.example.com/mcp”, }, }, ], multi_agent: { enabled: true, max_concurrent_subagents: 3 }, }, vault_ids: [“vault_YOUR_VAULT_ID”], environment: { type: “openai_hosted”, capability_directories: [“/workspace/capabilities/skills”], }, input: “Investigate service-api’s elevated 5xx rate over the last 30 minutes. ” + “Delegate deployment, error, and dependency analysis to subagents. ” + “Save findings, evidence, and recommended mitigation in /workspace/outputs.”, }); The quickstart covers API key permissions and SDK setup. Where the Agent Runs Environment choice is the main architectural decision. The Agents API supports 3 sandbox options, and it can also run without a sandbox. OpenAI-hosted sandbox: uses the sandboxing infrastructure behind Codex and ChatGPT. You can configure it with files, packages, skills, and plugins. Self-hosted: you run codex exec-server inside your environment. It registers with a restricted key and connects over WebSocket. All connections are outbound. Partner sandboxes: Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle, Runloop, and Vercel have first-class integrations. What the Harness Handles OpenAI maintains the harness alongside its models, with versioned access at each model launch. Long sessions: The API automatically compacts earlier context as a session nears its limit. Developers do not write their own compaction logic. Efficient tool use: Tool search loads tool definitions only when needed. This reduces token usage and cost while preserving the model’s cache. Programmatic tool calling lets agents run calls in parallel and chain operations. Agents filter or combine results in code, so only relevant data returns into context. Supported tools include MCP, custom functions, and built-in tools like web search. Subagents: With multi-agent support, the main agent splits complex tasks into independent pieces. Each subagent keeps its own context. The main agent coordinates them and combines the results. Agents API vs Agents SDK vs Responses API OpenAI’s runtime comparison positions the 3 options this way: Agents API Agents SDK Responses API Where the agent runs OpenAI runs a managed Codex harness Inside your application Your application, with optional hosted orchestration Integration effort Low Medium High State between tasks Saved session configuration, turns, and items Your storage and SDK sessions Manual history, response chaining, or Conversations Execution environment OpenAI-hosted, self-hosted, or no sandbox Your runtime and sandbox providers Your own environment Early Customer Results OpenAI published these customer-reported numbers. They are vendor-supplied, not independent benchmarks. Ciridae: evaluation score rose from 0.71 to 0.85, with a 4x latency reduction on subagent flows. SafetyKit: 60% lower cost per case after migrating its case review workflow. Hypha: 86% fewer failed agent responses after separating the harness from the sandbox. Nash.ai: runs thousands of long-running agents across global logistics networks. Key Takeaways OpenAI’s Agents API exposes the managed Codex harness as a public beta API. Agents run in OpenAI-hosted, self-hosted, or 9 partner sandboxes. Compaction, tool search, programmatic tool calling, and subagents come built in. There is no extra fee; you pay for tokens, tools, and container time. US-only data residency and no ZDR limit regulated workloads for now. Check out the Technical details. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post OpenAI Launches the Agents API in Public Beta, Putting the Codex Harness Behind One API Call appeared first on MarkTechPost.

OpenAI Launches the Agents API in Public Beta, Putting the Codex Harness Behind One API Call Leer entrada »

AI, Committee, Noticias, Uncategorized

Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster

Production LLM applications rarely receive a question nobody has asked before. Support assistants and RAG pipelines field the same intents thousands of times a day, each phrased differently, and most stacks treat every phrasing as a fresh, fully billed request. Redis LangCache is a fully managed semantic caching service that sits between the application and the model, matches incoming prompts against previously answered ones by meaning rather than exact text, and returns the stored response when a close enough match exists. Redis reports API cost savings of up to 90% and cache-hit responses up to 15x faster than re-querying the model. Is it deployable? Yes. LangCache is available today as a public preview on Redis Cloud, accessed through a REST API with Python and JavaScript SDKs, and Redis notes that features and behavior may change during the preview. The Problem: Paraphrases Are Still Full LLM Calls Consider three requests to a customer-support assistant: “Can I get a refund after buying the monthly plan?” “Is the monthly subscription refundable?” “Can I cancel the plan and get my money back?” The wording differs, but the question and answer are identical. Without a semantic cache, each version triggers a complete generation: input tokens processed, output tokens decoded, user waiting. Prefix caching only removes part of that cost. When requests share a system prompt or context, the engine reuses the KV states computed for that prefix, but the request still reaches the LLM, new tokens still get processed, and the full answer still gets decoded. A prefix-cache hit is a cheaper generation call, not an avoided one. How LangCache Works LangCache moves the cache outside the model and stores the generated response itself. The architecture is a two-call loop: Before invoking the model, the app sends the prompt to POST /v1/caches/{cacheId}/entries/search. LangCache generates an embedding for the prompt and runs a vector search over stored entries. If a semantically similar entry clears the configured similarity threshold, the cached response is returned and no LLM call occurs. On a miss, the app calls its chosen LLM as usual, then stores the prompt and new response through POST /v1/caches/{cacheId}/entries for future matches. Embedding generation is handled by the service, with default models or bring-your-own. Cache behavior is controlled through similarity thresholds, TTLs, and eviction policies, plus adaptive controls that tune precision and recall. Built on Redis’s vector database and exposed as a REST API, it works with any LLM provider and language. Hit rates and savings are monitored from the Redis Cloud console. What a Cache Hit Actually Saves A cache hit removes the input tokens, the output tokens, and the decoding latency of an additional model call. In a demo run comparing both paths on a paraphrased question, direct inference took 2.232 seconds and consumed 514 input tokens plus 250 output tokens. LangCache returned the earlier response in 0.37 seconds with zero LLM input or output tokens, roughly 6x faster in that run. The Redis documentation is careful about how savings accrue. On a cached response you do not pay for output tokens, while input token costs are typically offset by embedding and storage costs. The suggested estimate is: Est. monthly savings = (Monthly output token costs) x (Cache hit rate) With $200 of monthly LLM spend, 60% of it on output tokens, and a 50% hit rate, that works out to $60 saved per month. Redis also publishes a savings calculator for annual estimates. Redis’s public preview announcement cited up to 15x faster responses on cache hits and up to 70% lower token usage, while the current product page states savings of up to 90%. Customer Mangoes.ai reports a 70% hit rate on its patient-care voice app, cutting LLM spend by 70% with 4x faster responses. The actual result depends on how much safe repetition exists in the traffic. Where Semantic Caching Needs Care Deciding which questions can safely share an answer is a production concern, not a configuration detail. A threshold set too low returns a refund policy to a customer asking about upgrades. Set too high, nearly every paraphrase goes back to the model and the cache stops paying for itself. Production setups need well-tuned thresholds, expiration policies so stale answers age out, data isolation between tenants, and monitoring for incorrect matches. LangCache covers these with access scopes, custom filtering, TTL and eviction controls, and monitoring through Redis Cloud. Data stays on the customer’s Redis servers, and Redis states it does not access that data or use it to train models. Key Takeaways Prefix caching cuts prompt-processing cost; semantic caching eliminates the LLM call entirely on a hit. LangCache is a two-call REST integration: search before the model, store after it. Savings come mainly from avoided output tokens; the docs give the formula output cost x hit rate. Redis claims up to 90% cost savings and up to 15x faster cache hits; a demo run showed 6x. Thresholds, TTLs, isolation, and false-match monitoring decide whether a semantic cache is safe. Check out redis.io/langcache and follow the API and SDK examples. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster appeared first on MarkTechPost.

Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster Leer entrada »

We use cookies to improve your experience and performance on our website. You can learn more at Política de privacidad and manage your privacy settings by clicking Settings.

Privacy Preferences

You can choose your cookie settings by turning on/off each type of cookie as you wish, except for essential cookies.

Allow All
Manage Consent Preferences
  • Always Active

Save
es_ES