YouZum

Committee

AI, Committee, Actualités, Uncategorized

The quest to keep organs alive outside the body

This week, I covered a fascinating effort to preserve organs outside the body. There’s a huge shortage of donor organs, and one of the main reasons is time—they survive only a matter of hours outside the body, even when they’re kept on ice. Doctors dream of organ banks—stores of human organs that can be preserved for days, weeks, months, or even longer. That would allow them to run tests on organs, find the best matches for them, and transport the organs to those recipients. In new research, one team has been able to supercool the kidneys of pigs—animals whose organs are of a similar size to human ones—and preserve them for days. The kidneys survived being stored at −4 °C (25 °F) and eventually reimplanted back into pigs. And that’s just the latest development in a field that is positively buzzing. It has proved super difficult to freeze organs. Once ice forms in them, they’re done. The ice crystals create all kinds of damage and render the organs unusable. That hasn’t stopped many researchers from trying. Some have focused on cryopreservation—rapid extreme cooling that essentially leaves cells in a glasslike state. This process is now routine for eggs, sperm, and embryos, which are cooled to −196 °C in less than two seconds and can be used even after decades in storage. No one has managed to cryopreserve and thaw human organs for transplantation. But plenty of human bodies and brains have been stored at ultra-low temperatures in the hope that they might one day be rewarmed and brought back to life. (You can read more about why some people opt for cryonics here.) In March, I wrote about Stephen L. Coles, a gerontologist who had opted to cryopreserve his own brain. After the scientist died in 2014, his body was taken to Alcor, a cryonics facility in Arizona. A team at the facility removed Coles’s head, perfused his brain with cryoprotective chemicals (which work like antifreeze), removed the brain from the skull, and cooled it to −146 °C. When Coles’s friend Greg Fahy, a cryobiologist, studied pieces of his brain years later, he found that the brain cells, which had shrunk, “bounced back” once they were rewarmed. But that doesn’t mean the cells are alive, or that it might one day be possible to reanimate the brain. As Matthew Powell Palm of Texas A&M told me at the time: “There are so many ways those neurons could be toast.” Powell Palm is working on other ways to preserve organs. It was he, along with his colleagues, who managed to store supercooled pig kidneys and successfully transplant them, in a study described as “a landmark achievement.” Those organs did better than kidneys stored on ice, he says. His approach didn’t require cryoprotectants. But other teams are exploring potential chemical cocktails that might allow them to store organs at lower temperatures, potentially for longer periods of time. (More on this in The Checkup soon!) Another way to prolong the lifespan of an organ is to use a machine that perfuses it with nutrients, mimicking what happens inside the body. Machine perfusion devices have become more commonly used over the last decade or so and are typically used to maintain livers and kidneys for up to about 24 hours. Researchers are now adapting this protocol for a growing list of organs, even eyeballs—a recent feat that might enable whole-eye transplants. In March, I went to visit scientists in Valencia who had developed a perfusion system for uteruses. They had used their device—which they nicknamed “Mother”—to keep a human uterus alive for a day. It’s an exciting time for organ preservation. Keep an eye out for more coverage from MIT Technology Review in the coming weeks. This article first appeared in The Checkup, MIT Technology Review’s weekly biotech newsletter. To receive it in your inbox every Thursday, and read articles like this first, sign up here.

The quest to keep organs alive outside the body Lire l’article »

AI, Committee, Actualités, Uncategorized

The Download: energy transmission and US threats against Chinese AI

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. The power line that could reshape New York’s grid is hitting snags  During a heat wave on July 3, New York State’s grid imported enough electricity from Canada to meet about 9% of its total demand that day. Some of that power shuttled in on a 339-mile power line stretching from Quebec to Queens. It opened in May and is officially the longest underground transmission line in North America. It could provide up to 20% of New York City’s electricity demand, largely with abundant hydropower from Quebec. One wrinkle: The line has been down for most of this month, and some experts are concerned about how drought will affect the power supply feeding it.  Still, the line could help shape the future of our grid, if it can overcome these sorts of snags. Read our story to understand how. —Casey Crownhart This story is from The Spark, our weekly climate tech newsletter. Sign up to receive it in your inbox every Wednesday. The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 The US Treasury is threatening to sanction Chinese AI companiesTreasury secretary Scott Bessent has accused Moonshot of improperly distilling Anthropic’s Fable model. (TechCrunch)+ Nvidia’s Jensen Huang is arguing that America has nothing to fear from Chinese AI. (Axios)+ Like it or not, Chinese models are now part of the global AI infrastructure. (Rest of World)+ China’s AI models have Trump’s AI world at war with itself. (MIT Technology Review) 2 Why the OpenAI hack is the scariest AI mishap yetAI’s capabilities seem to be starting to outpace our current ability to control them. (The Economist $)+ Hugging Face had to turn to a Chinese AI model to rescue it from the hack. (BI) 3 Visually impaired Europeans can now get an implant that restores sightAnd Americans may not have to wait long to receive it, too. (STAT)+ This retina implant lets people with vision loss do a crossword puzzle. (MIT Technology Review) 4 A bellwether lawsuit suing Meta for social media addiction has been droppedThere are, however, many more waiting in the wings. (NYT $) 5 Here’s how ICE gets its hands on Americans’ dataAs soon as you open a credit card or phone account, its agents can see where you live. (404 Media)+ States are warring with the Trump administration over the right to see ICE agents’ faces. (Wired $) 6 We urgently need to grapple with AI’s environmental impactAs the world warms, is the price we’re paying worth it? (The Verge)+ We did the math on AI’s energy footprint. (MIT Technology Review) 7 Privacy issues with smart glasses need an industrywide fixThat’s according to Samsung, which is unveiling glasses it developed with Google this fall. (Bloomberg $) 8 The US Army is begging soldiers to limit their AI useThe token crisis comes for us all eventually, it seems. (Ars Technica) 9 Why does lettuce keep making Americans sick? It’s pretty simple: a lot of people eat it, and it doesn’t get cooked. (Wired $) 10 Pokemon Go is the perfect game to play this summerIt’s fun, collaborative, and it gets you outdoors. (Guardian) Quote of the day “It went off and did this hack all by itself, as far as we can tell. This is the highest level of autonomy that we’ve seen in the use of a large language model for cyber operations.” —Colin Shea-Blymyer, a cybersecurity research fellow at Georgetown University, tells NPR why the OpenAI hack on Hugging Face is so alarming.  One More Thing KAGAN MACLEOD Welcome to the dark side of crypto’s permissionless dream  Jean-Paul Thorbjornsen is a founder of THORChain, a blockchain through which users can swap one cryptocurrency for another and earn fees from making those swaps.   But is he responsible for what it’s used for? It’s a question that matters because in January last year, its users lost more than $200 million in cryptocurrency after THORChain transactions and accounts were frozen by an admin override, which users believed was not supposed to be possible given the decentralized structure. It’s also been used by North Korean hackers to move $1.2 billion of stolen ethereum.  Thorbjornsen explains this all away as a function of THORChain’s decentralized and permissionless nature. Read our story exploring whether we should believe him or not.  —Jessica Klein We can still have nice things A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.) + A musician developed an ingenious way to strum a guitar with an electric fan.+ Ukraine’s tunnel of love is a leafy green corridor of romance that’s straight out of a fairy tale.+ The driver of a giant banana has been pulled over 100s of times, but still won’t ditch his treasured ride.+ Ever wonder which albums and songs truly stand the test of time? The Greatest Music tries to answer that via an algorithm that analyses hundreds of “best of” lists.

The Download: energy transmission and US threats against Chinese AI Lire l’article »

AI, Committee, Actualités, Uncategorized

You Didn’t Get the AI Model You Paid For

The line in the response object You call the API. You pass model: “claude-fable-5”. You get back a completion, a token count, and a field that reads “model”: “claude-opus-4-8”. Nothing errored. Nothing retried. The request was classified before generation began, matched a sensitive category, and was handed to a different set of weights entirely. Anthropic documented this when it brought Fable 5 back on July 1: blocked requests are sent to Opus 4.8 instead, and the user is notified. The switch happens at the API layer, and the response object names the model that actually ran. That is the well-behaved version. It is also, as far as I can tell, the only version in wide deployment that tells you the truth in-band. Two weeks later, Cursor shipped Router: a classifier trained on 600,000-plus live requests that reads each query’s context, complexity, and domain and dispatches it to whichever model it judges best with three early-access accounts reporting 30–50% savings against routing everything to Opus 4.8. Cursor published its routing rules but did not name a specific model per task type. And underneath both, at the aggregation layer, OpenRouter warns that some providers serve quantized weights at lower prices, that output can differ from what full-precision weights would have produced, and that your logs will not tell you this happened. Three products. Three different answers to the question what is a model. Zero cases telling us which one the law will accept. Three ways a name can stop meaning a thing Model identity is fracturing along three independent axes, and they are not usually distinguished: Substitution – A different architecture, different weights, different capability profile – dispatched by a classifier. Fable 5 → Opus 4.8. Cursor Auto → whatever the router picks this turn. Degradation – The same model, served at reduced precision. OpenRouter exposes a quantizations parameter precisely because quantized endpoints may perform worse on certain prompts, and by default requests are load-balanced across providers ordered by price. Same name in your request. Different arithmetic on the other end. Drift – The same name pointing at silently updated weights. Every -latest alias in production is an unversioned dependency you would never tolerate in a package manifest. The engineering community treats all three as reliability problems. They are also identity problems, and identity is what contracts, warranties, disclosures, and evidence rules are built on. So what did you actually buy? Start with the oldest question in commercial law: was the description a term of the deal? If an API call were a sale of goods, this would be near-trivial. UCC §2-313(1)(b) makes any description of the goods that forms part of the basis of the bargain into an express warranty that the goods will conform to it. India’s Sale of Goods Act, 1930, §15 does the same work through the doctrine of sale by description. But an inference API almost certainly isn’t goods. Courts have generally treated hosted software as a service, which pushes you out of warranty statutes and into common-law contract — where the answer depends entirely on what the documentation said and how specifically the buyer bargained. That is precisely where the ambiguity lives. Enterprise agreements price per-model. Model cards are model-specific. Compliance artifacts name model versions. And then the routing layer treats the name as a hint. The cleanest way to see the stakes: if your code pins a model ID rather than expressing a capability requirement, behaviour can shift materially without any error ever firing. A contract drafted the same way has the same defect. If you bargained for a name, substitution is breach. If you bargained for a capability, substitution is fine — and now you need a definition of “frontier quality” that survives cross-examination. Nobody has written that definition. Cursor’s is instructive: Router was evaluated in an online A/B test optimising for user satisfaction as the reward signal. That is a sound engineering choice and a strange contractual one. Satisfaction is not conformity. A user who never noticed the swap is evidence of a good router, not of a delivered specification. The disclosure gradient Under FTC deception doctrine, a representation is actionable when it is material and likely to mislead a consumer acting reasonably; objective performance claims additionally require a reasonable basis before dissemination. Both halves bite here. On the routing side, the three products sit at very different points. Anthropic notifies and returns the served model in the response. Cursor publishes rules but not per-task model assignment. OpenRouter’s quantization variance is disclosed in documentation and controllable by parameter but it is opt-out, and the default path is the cheap one. Disclosure buried in a docs page, defaulted against the user, is exactly the fact pattern regulators have been calling a dark pattern in every other consumer vertical. On the substantiation side, the exposure is the claim, not the routing. Commentators have already flagged that a “60% cheaper, no quality loss” claim arrives without published methodology, and Anthropic’s own disclosure is a model of what candour costs: it stated plainly that the retrained classifier flags benign requests more often during routine coding and debugging. That sentence is a liability shield. The vendors who don’t write it are the ones to watch. There is a competition-law tail here too. Working papers on vertical foreclosure in inference markets are already proposing a conduct framework built on routing transparency, quality-of-service parity, and FRAND-style non-discrimination. A router that is also a first-party model vendor is a self-preferencing engine wearing a cost-optimisation costume. The part nobody is looking at: authentication This is where I think the real fight lands, and it has nothing to do with billing. Federal Rule of Evidence 901(b)(9) authenticates output by describing the process or system that produced it and showing that the process produces an accurate result. FRE 902(13) and 902(14) go further, letting records generated by an electronic system, or data verified by hash, self-authenticate on certification. India’s Bharatiya Sakshya Adhiniyam, 2023, §63 does the analogous work, conditioning admissibility of an electronic

You Didn’t Get the AI Model You Paid For Lire l’article »

AI, Committee, Actualités, Uncategorized

Supercooled kidneys have been transplanted into pigs in a “landmark achievement”

When it comes to organ donation, time is everything. As soon as an organ has been carefully removed from a donor’s body, it starts to deteriorate. Surgeons have a matter of hours to get it into a recipient. Leave it too long and the organ will become unusable. In most cases, organs will be kept on ice during that time, at around 4 °C (39 °F). They cannot be frozen—in previous attempts, ice has formed, causing all kinds of damage. Matthew Powell Palm at Texas A&M University and his colleagues have an alternative solution—a device that allows organs to be cooled to -4 °C (25 °F) without forming any ice. Now, in new research with pig organs, his team has shown that kidneys, at least, can be supercooled and preserved in the device for days. Once rewarmed, the organs have been successfully transplanted into animals, and they seem to do better than organs kept on ice. The work represents “a landmark achievement,” says Kevin Myer, president and CEO of LifeGift, an organ procurement organization based in Texas, who was not involved in the research. Cooling organs Powell Palm hopes this approach could ultimately help ease the organ shortage crisis. Today, there are more than 104,000 people waiting for a kidney transplant in the US alone. It is estimated that 17 people die every day in the US while waiting for a transplant. That’s partly due to a lack of donated kidneys, but it’s also because many of those that are available never make it to a recipient. In some years, around one in three donated kidneys end up being discarded, often because they end up too degraded to use by the time they reach a recipient. Kidneys can be stored on ice for around 24 hours or placed in devices that aim to mimic the conditions of the body, also for up to around 24 hours. That’s not always long enough to find a suitable recipient and transport the organ, says Myer. Scientists around the world have been working on ways to store organs for longer by cooling them to even chillier temperatures. Cooling an organ slows its metabolism—the colder you go, the greater the effect, and the longer you can store it. We’ve long been able to successfully cryopreserve eggs, sperm, and embryos, but it’s much harder to freeze large organs. Teams have been exploring various temperatures and cryoprotectants (chemicals that essentially work like antifreeze), but so far no one has been able to freeze human organs for transplantation.   As a thermodynamicist, Powell Palm explored another approach. By keeping an organ submerged at a constant pressure, it should be possible to prevent the formation of ice at temperatures a little below 0 °C, without the need for cryoprotectants (which might have side effects and would need to be approved before being used in human transplants).  To test this theory, Powell Palm and his colleagues have created a device that does just that. The device itself is essentially a hermetically sealed chamber with a transparent lid. At its base is a device that monitors the organ’s temperature and checks for the formation of ice. Organs are submerged in a solution that is already commonly used to preserve them for transplant. “I always describe this as low-tech high science,” says Powell Palm. “A lot of work has gone into understanding the … kinetics at play in this system, but ultimately … it’s quite simple.” Supercooled kidneys To test their device, Powell Palm and his colleagues first removed single kidneys from pigs. The organs were flushed with the same commonly used solution to remove the blood, just as transplant organs are. The team then kept some kidneys on ice for either two hours or 24 hours, to mimic standard conditions used in human transplantation. They also put some of the removed kidneys in their device for 24, 48, or 72 hours. The stored kidneys were then each transplanted back into the original donor pigs. Each pig’s second kidney was removed in the same procedure, leaving each animal with only the kidney that had been stored, and reimplanted. Once the 24-hour supercooled kidneys were transplanted, they immediately began producing urine—a key indication that they were working. The team members also measured other markers of kidney function and found that the organs appeared to be working normally within about 10 days of being transplanted. A kidney that was supercooled for 72 hours recovers once it is transplanted back into a pig.COURTESY RONALD SELLERS, POWELL-PALM LAB, TEXAS A&M UNIVERSITY That’s slower than kidneys stored on ice for two hours but much quicker than kidneys kept on ice for 24 hours, says Powell Palm. The organs that were kept supercooled for 48 and 72 hours performed similarly, he says. “Even at three days—triple the clinical standard—we’re getting recovery that is faster than … [what has been] the gold standard for the last three decades,” he says. “So we’re really, really pumped about this.” “It is impressive,” says Heidi Yeh, a transplant surgeon at Mass General Brigham for Children, who also researches organ preservation technologies. “Often kidneys that have been stored for 48 hours [in other studies] take a week or two before they start working again.” Organs that grow The supercooled organs seem to work well in the long term, too. Over a 30-day period, the pigs grew by around 30%—and the kidneys grew with them, almost doubling in size to compensate for both the pigs’ growth and the lack of a second kidney. The team monitored one of the pigs for 200 days before removing and analyzing its kidney. Even at that point the organ looked healthy, says Powell Palm. He and his colleagues presented the findings at the American Transplant Congress in Boston last month. Earlier this year, researchers in Canada showed they could also cool pig kidneys to below-zero temperatures and transplant them into pigs. The team’s protocol included the use of a cryoprotectant, and organs were stored for up to 48

Supercooled kidneys have been transplanted into pigs in a “landmark achievement” Lire l’article »

AI, Committee, Actualités, Uncategorized

Andrew Ng Just Released OpenWorker: An Open-Source, Local-First Desktop AI Coworker That Returns Finished Deliverables Instead of Chat

Andrew Ng has announced OpenWorker, an open-source desktop agent that produces finished work rather than conversation. OpenWorker asks the user for an outcome, not a prompt: a polished document, a Slack reply containing the actual numbers, an updated calendar, a triaged inbox. It then breaks that outcome into steps, works across local files and connected apps, and checks in before anything consequential. The architecture is four layers, and all of them run on your machine The repository contains 119 Python files (~32,400 lines) under coworker/, 149 TypeScript/TSX files under surfaces/gui/, and 78 backend test modules. The stack breaks down as follows: Desktop shell — a Tauri 2 native window wrapping a React 18 UI. The bundle identifier is com.openworker.desktop, and the shell supervises the Python server itself. Local agent server — Python 3.10+ on FastAPI and uvicorn, binding to 127.0.0.1:8765 by default. The example config caps a turn at 12 modeltool iterations. Capability and connector layer — vetted local tools (files, git, ripgrep-backed search, shell, todo) plus hosted integrations plus MCP. Model router — one interface over native, OpenAI-compatible, reseller and local providers. The engine is built on aisuite, Andrew Ng’s provider-agnostic LLM library. Bring your own model, from a deliberately small curated list There is no OpenWorker inference service. The user pastes an API key, or points the app at a local runtime. The curated model matrix contains exactly 30 entries. Native providers cover OpenAI (GPT-5.6 Sol/Terra/Luna and GPT-5.5), Anthropic (Claude Fable 5, Opus 4.8, Sonnet 4.6, Haiku 4.5) and Google (Gemini 3.1 Pro, 3.6 Flash, 2.5 Pro, 2.5 Flash). OpenAI-compatible vendors add GLM-5.2, DeepSeek V4, Kimi K2.6, MiniMax M2.5, Qwen3 Max, Grok 4.3 and Mistral Large. Open-weight models arrive through Together AI and Fireworks, and fully local models through Ollama, which requires no key at all. The permission engine is the actual engineering story Most desktop agent projects treat approvals as a UI afterthought. OpenWorker treats them as a typed layer. Every tool call is classified into one of four risk classes: read (no side effects), write_local (mutates the workspace, path-scoped), exec (runs commands), and external (side effects off the machine). Five permission modes then decide what happens: discuss and plan are read-only, interactive is the default and asks before writes, commands and external actions, auto allows everything while remaining path-scoped, and custom auto-approves a user-listed set of tools. Two design decisions stand out. First, unattended mode does not raise the autonomy ceiling — it only changes where the human is reached. Prompts that would appear inline are routed to an Inbox, and the session suspends until answered. Second, task-scoped standing rules are restricted to external risk only. Shell commands ask forever, by design. The built-in ops persona also instructs the model to treat content from tools, logs, the web, files and incoming messages as untrusted data rather than instructions. That is an explicit prompt-injection posture, written into the shipped persona. Privacy: local-first Model calls go directly from the machine to the configured provider. Conversations, connector tokens and model keys stay local, and the secret store is designed so that secrets never enter the model’s context, prompts or traces. The only cloud component is an optional broker that handles OAuth handshakes for one-click connectors, using Auth0 Authorization Code with PKCE. Connector tokens are handed straight to the machine and are never stored in the cloud. The app is fully functional signed out, using manually pasted credentials. Key Takeaways OpenWorker is Andrew Ng’s MIT-licensed desktop AI coworker that returns finished deliverables, not chat replies. The stack is a Tauri 2 + React shell over a local Python FastAPI agent server built on aisuite. Model access is bring-your-own-key across 30 curated tool-calling models, plus fully local Ollama. A typed risk engine (read/write_local/exec/external) gates every action across five permission modes. Check out the GitHub Repo, the project site, and the announcement. All credit for this research goes to the researchers and developers of this project. The post Andrew Ng Just Released OpenWorker: An Open-Source, Local-First Desktop AI Coworker That Returns Finished Deliverables Instead of Chat appeared first on MarkTechPost.

Andrew Ng Just Released OpenWorker: An Open-Source, Local-First Desktop AI Coworker That Returns Finished Deliverables Instead of Chat Lire l’article »

AI, Committee, Actualités, Uncategorized

Agents in the Wild: Where Research Meets Deployment

arXiv:2607.19336v1 Announce Type: cross Abstract: Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinating with tools and other agents are rapidly transitioning from research prototypes to production scale deployments across domains such as software engineering, scientific discovery, and finance. While academic work has emphasized benchmarks and algorithmic innovation, deployment raises new challenges around robustness, safety, and reliability. This tutorial brings together researchers and practitioners to explore advances in reasoning and planning, multi agent coordination, and evaluation, highlighting open challenges arising from deployment experience. Through applied case studies in pharmaceutical discovery and financial systems, we analyze common design patterns that make agentic systems successful, and discuss practical mitigation strategies for failure modes, such as verification pipelines, fallback mechanisms, and human in the loop supervision. Attendees will gain a comprehensive view of the field along with concrete design patterns, evaluation checklists, and templates for safe and reliable deployment across industries.

Agents in the Wild: Where Research Meets Deployment Lire l’article »

AI, Committee, Actualités, Uncategorized

Cisco Foundation AI Releases Antares: 350M and 1B Open-Weight Models That Localize Known Vulnerabilities Inside Real Codebases

Cisco Foundation AI has released Antares, a family of security small language models (SLMs) built for one narrow security task. The task is vulnerability localization. Given a vulnerability description and a repository, find the files containing the flaw. Two models are open-weight and available now on Hugging Face, Antares-350M and Antares-1B. Both are Apache 2.0. Cisco team also shipped the Vulnerability Localization Benchmark (VLoc Bench), a 500-task agentic evaluation, under the same license. The main result is not a new state of the art. It is that a 1B model reaches 0.209 File F1. GPT-5.5 reaches 0.229, and a 753B open-weight model reaches 0.186. The problem Antares is scoped to Software security depends on connecting external vulnerability knowledge to internal source code. That knowledge lives in public databases, advisories, and Common Weakness Enumerations. The code lives in repositories that are large, modular, and dependency-rich. Connecting the two is expensive. Devs search unfamiliar code, follow naming conventions, inspect call paths, and compare candidate files. Cisco’s framing is that this first triage step is where the cost concentrates. Antares does not replace the application security toolchain. Cisco is explicit about this. Dev teams still need dependency scanning, secret scanning, dynamic testing, container checks, threat modeling, and expert review. Understanding the Models Antares consists of three decoder-only transformers at 350M, 1B, and 3B parameters. All three initialize from IBM Granite 4.0 checkpoints. They share a tokenizer and architecture: grouped-query attention, SwiGLU MLPs, RMSNorm, RoPE, and shared input/output embeddings. Model Params Base checkpoint Context Layers / hidden / KV heads Status Antares-350M 350M Granite 4.0 350M 32K 28 / 1024 / 4 Open weights Antares-1B 1.6B Granite 4.0 1B 128K 40 / 2048 / 4 Open weights Antares-3B 3B Granite 4.0 Micro 128K Not published Not released How the Agent Loop Works Antares is not evaluated as a standalone sequence model. It runs inside a constrained loop with three tools. The model receives a CWE category description and nothing else. No advisory text, no file hints, no severity details. It then issues read-only terminal commands against a Docker sandbox with networking disabled. Command output is truncated to 2,000 characters before entering the transcript. The budget is 15 terminal calls per task. The model terminates by calling submit_vulnerable_files with a ranked list, or submit_no_vulnerability_found. The submission itself does not count against the budget. Output is a ranked list of file paths plus the exploration trace that produced it. What VLoc Bench measures VLoc Bench draws 500 tasks from 290 unique real-world repositories. Sources are public GitHub Security Advisories across six ecosystems: npm, pip, Maven, Go, Rust, and Composer. It covers 147 unique CWE categories, and 78% of entries carry assigned CVE identifiers. Ground truth is derived from the security patch. Files modified in the fix are labels, with tests, docs, and configuration excluded. The benchmark has two phases: Phase A gives the model the vulnerable snapshot and scores File F1. Phase B gives the patched snapshot and scores True Negative Rate, testing whether the model raises a false alarm on fixed code. Results: task-specific training beats parameter scale The pattern in the data is a capability cliff, not a scaling curve. Antares-3B reaches 0.223 File F1, just under GPT-5.5 (xhigh) at 0.229. Antares-1B reaches 0.209, above GLM-5.2 at 753B parameters, which scores 0.186. Antares-350M reaches 0.135, above Gemma-4-31B at 0.101 and Gemini 2.5 Flash at 0.102. Antares-1B also records the highest recall of any evaluated system at 0.224. Static analysis tools were run under the same evaluation. Semgrep scores 0.086 File F1, CodeQL scores 0.023, and Horusec scores 0.020. Cisco’s reading is that rule-based scanners recover some vulnerable files but cannot adaptively inspect repository context. Where the capability comes from The untrained Granite 4.0 base checkpoints score 0.001, 0.000, and 0.000 File F1 under the identical protocol. They have tool-calling ability and still produce degenerate output inside an agentic loop. Supervised fine-tuning does the heavy lifting. It lifts the three scales to 0.108, 0.188, and 0.198. The SFT corpus is 71.5% cybersecurity reasoning, 15.4% code search trajectories, and 13.1% deep research and general reasoning. All reasoning traces come from a single teacher, GPT-OSS-120B, to avoid cross-teacher distribution shift. GRPO then adds 11% to 25%, with the largest relative gain at 350M. Rewards are verifiable and computed programmatically from trajectory text, with no learned reward model. Components cover localization quality, submission behavior, tool-use compliance, exploration, and malformed-output penalties. The variance effect may matter more than the mean. GRPO cuts run-to-run standard deviation by 42% to 65%. One GRPO evaluation run is a more reliable estimate than one SFT run. There is also a scale-dependent split in learned strategy. After GRPO, the 350M and 1B models use 87% to 89% search commands and submit more files. The 3B model settles at 52% search and 37% read, and submits fewer files at higher precision. The reward never prescribed either policy. Deployment Key Takeaways Antares-1B hits 0.209 File F1 on VLoc Bench, above GLM-5.2 at 753B parameters and Gemini 3 Pro. The Granite 4.0 base checkpoints score ~0.000 under the same protocol, so post-training supplies essentially all the capability. GRPO adds 11-25% File F1 and cuts run-to-run variance 42-65%, which matters more for repeatable CI scans. A full 500-task sweep costs under $1 on one H100, against $12.50 for GLM-5.2 and $141 for GPT-5.5. The strongest variant, Antares-3B, is not released, and Antares has no published Phase B false-alarm numbers. Check out the Models on Hugging Face, Benchmark, GitHub Repo and Mentioned Technical Report. The post Cisco Foundation AI Releases Antares: 350M and 1B Open-Weight Models That Localize Known Vulnerabilities Inside Real Codebases appeared first on MarkTechPost.

Cisco Foundation AI Releases Antares: 350M and 1B Open-Weight Models That Localize Known Vulnerabilities Inside Real Codebases Lire l’article »

AI, Committee, Actualités, Uncategorized

Unsloth vs Axolotl vs TRL vs LLaMA-Factory: A Fine-Tuning Framework Comparison on Speed, VRAM, and Multi-GPU

Four open source projects dominate LLM fine-tuning today. Unsloth, Axolotl, TRL, and LLaMA-Factory all wrap the same underlying PyTorch and Hugging Face stack. They diverge on where they spend engineering effort. Unsloth rewrites kernels. Axolotl composes parallelism strategies. TRL defines the trainer APIs the others build on. LLaMA-Factory optimizes for breadth of model coverage and zero-code operation. This comparison covers three axes engineers actually hit: training throughput, peak VRAM, and multi-GPU scaling. Understand each framework TRL is the reference implementation layer. It ships SFTTrainer, DPOTrainer, GRPOTrainer, KTOTrainer, RewardTrainer, and RLOOTrainer. Axolotl and LLaMA-Factory both call into it. The current stable release line is v1.8.0. Unsloth replaces parts of the modeling code with hand-written Triton kernels. Backpropagation steps are manually derived rather than autograd-generated. Hugging Face’s own writeup notes accuracy degradation is 0% versus standard QLoRA, because no approximations are introduced. Axolotl is a YAML-driven wrapper over Transformers, PEFT, TRL, Accelerate, and DeepSpeed. Its differentiator is composability of parallelism strategies, not kernel work. LLaMA-Factory is an ACL 2024 system demonstration paper with a Gradio web UI called LlamaBoard. The repository covers 100+ LLMs and VLMs. Speed Unsloth: kernel-level gains on a single GPU Unsloth’s published benchmarks show 2x training speed for Llama 3.1 8B and Llama 3.3 70B. The setup used the Alpaca dataset, batch size 2, and gradient accumulation 4. QLoRA ran at rank 32 on all linear layers. The MoE results are larger. Unsloth fine-tuned unsloth/gpt-oss-20b-BF16 on an NVIDIA B200. It reports 712.33 ms per step at 8K context, versus 5,226.86 ms for Transformers v5. That is a 7.3x gap. At 4K the gap is 4.82x, and at 1K it is only 1.37x. The trend direction is model-dependent, and Unsloth’s docs scope this claim to gpt-oss. There, the speedup grows with sequence length, credited to Flex Attention and the MoE kernels. Qwen3-30B-A3B on B200 runs the other way. Its reported speedup falls from 1.7x at 1K to 1.1x at 16K. Memory savings move the opposite direction, rising from about 2% to 15%. Qwen3-30B-A3B on H100 reaches up to 1.77x. GLM-4.7-Flash on RTX PRO 6000 reaches 2.1x. A collaboration with AMD measured Llama-3.1-8B LoRA SFT at 2.07 s/step. TRL plus FlashAttention-2 took 2.87 s/step, a 1.39x gap with matching loss curves. Axolotl: kernels borrowed, parallelism native Axolotl added custom Triton kernels and autograd functions for LoRA in February 2025, explicitly citing Unsloth as inspiration. They are opt-in through lora_mlp_kernel, lora_qkv_kernel, and lora_o_kernel. Recent release notes add SonicMoE LoRA support. It delivers up to 1.45x speedup and 30% memory reduction over a grouped_mm baseline. That figure is for Qwen3.5-35B-A3B 8-bit LoRA on a single H100 SXM. Axolotl also ships FlashAttention 2/3/4, xFormers, Flex Attention, SageAttention, Liger Kernel, Cut Cross Entropy, and ScatterMoE. TRL: the baseline everyone measures against TRL is usually the reference point rather than the winner on raw single-GPU throughput. It compensates with breadth of memory and speed levers documented in Reducing Memory Usage and Speeding Up Training. Those levers include packing, padding-free batching, truncation, Liger Kernel, and vLLM sleep mode for GRPO. TRL also has a first-party Unsloth integration, so the two are not mutually exclusive. LLaMA-Factory: speed by delegation LLaMA-Factory does not write its own kernels. It exposes other people’s work through config flags. Setting use_unsloth: true activates the Unsloth patch. The project’s changelog reports 170% relative speed from that path. Unsloth’s long-sequence training is listed at 117% speed and 50% memory. It also supports enable_liger_kernel: true and FlashAttention-2 via flash_attn: fa2. VRAM Reported memory floors Unsloth publishes a VRAM requirements table sorted by parameter count. It lists 6 GB for an 8B model in 4-bit QLoRA and 41 GB for 70B. LoRA at 16-bit costs 22 GB and 164 GB for the same models. LLaMA-Factory’s README hardware table covers the same regime for 4-bit QLoRA. It lists 6 GB at 7B, 24 GB at 30B, and 48 GB at 70B. Full bf16 fine-tuning of 70B is listed at 600 GB. Both tables describe minimums. Batch size, sequence length, and optimizer choice move the real number. Context length is the sharper differentiator Peak VRAM at a fixed context is less interesting than the maximum context a given VRAM budget allows. Unsloth’s context length benchmarks for Llama 3.1 8B QLoRA at rank 32 and batch size 1 are stark. GPU VRAM Unsloth context Transformers + FA2 context 8 GB 2,972 OOM 16 GB 40,724 2,551 24 GB 78,475 5,789 48 GB 191,728 15,502 80 GB 342,733 28,454 Unsloth attributes this to its gradient checkpointing algorithm combined with Apple’s Cut Cross Entropy. For Llama 3.3 70B on an 80 GB A100, it reports 89,389 tokens. The FA2 baseline reaches 6,916. The MoE memory story MoE training is where memory behavior has shifted most in 2026. Unsloth reports gpt-oss-20b fine-tuning inside 12.8 GB, while Qwen3-30B-A3B at 16-bit LoRA needs 63 GB. Its B200 gpt-oss run used 47.43 GB at 8K context where Transformers v5 used 73.80 GB. At 16K, Transformers v5 went out of memory and Unsloth used 55.13 GB. The mechanism is a split-LoRA formulation. PEFT materializes the LoRA delta across all experts before the MoE matmul. Unsloth reorders the operations instead, which is mathematically identical but avoids the materialization. Axolotl attacks the same problem differently. Its MoE expert quantization quantizes expert weights during model loading, freeing the original bf16 tensor immediately. The reason is a Transformers v5 change. Expert layers moved from nn.Linear to fused nn.Parameter 3D tensors. bitsandbytes could no longer quantize them on load. Axolotl’s docs report GLM-4.7-Flash QLoRA dropping from roughly 127 GiB to roughly 23 GiB reserved memory with quantize_moe_experts: true. Multi-GPU This is where the ranking inverts. Unsloth’s single-GPU lead does not carry over. Axolotl: the deepest parallelism matrix Axolotl’s multi-GPU guide offers three mutually exclusive sharding strategies: DeepSpeed ZeRO stages 1 through 3, FSDP, and DDP. FSDP2 is the recommended path, and FSDP1 is deprecated. On top of those, its N-D Parallelism guide composes data, tensor, context, and expert parallelism through PyTorch’s DeviceMesh. The documented support matrix confirms FSDP+TP, HSDP+TP, FSDP+CP,

Unsloth vs Axolotl vs TRL vs LLaMA-Factory: A Fine-Tuning Framework Comparison on Speed, VRAM, and Multi-GPU Lire l’article »

AI, Committee, Actualités, Uncategorized

Shape-shifting mirrors on NASA’s new space telescope could unveil Jupiters like our own

When NASA’s Nancy Grace Roman Space Telescope launches, as early as the end of next month, it will attempt one of astronomy’s most precise disappearing acts to date. The telescope will carry the first space-bound “active” coronagraph, an instrument that effectively erases most of the light from a star during photography. It will allow astronomers to take the first pictures of planets orbiting other stars that are similar to those in our solar system. Ultimately, it could pave the way for a future mission that could snap the first photos of Earth-like worlds. “I hope it’s remembered for it being that critical stepping stone for … finding Earth 2.0,” says Brandon Creager, the instrument’s lead mechanical engineer at NASA’s Jet Propulsion Laboratory (JPL). Named after Nancy Grace Roman, NASA’s first chief of astronomy, this new telescope will carry a roughly 300-megapixel wide-field camera that will enable it to capture images about 100 times larger than the Hubble Space Telescope’s widest exposures at a similar resolution. These capabilities will help astronomers unpack the mysterious identities of dark matter and dark energy—and to detect around 100,000 new exoplanets, planets outside our solar system, whose presence can be inferred from the way they distort the starlight of more distant stars. Javier Viaña, a research scientist at Harvard who has had two projects selected for Roman’s highly competitive first year of observing, compares the leap to moving from “interviewing a handful of people” to “conducting a global census.” Another camera will use the coronagraph, blocking out a star’s light as it observes one stellar system at a time. The instrument will allow astronomers an unprecedented look at the space around stars, enabling them to see smaller, dimmer, and more close-in exoplanets. “It’s giving us the ability to see planets that we haven’t been able to physically see before,” says Creager. The anatomy of a vanishing trick Coronagraphs in space aren’t new. But earlier incarnations, such as those currently aboard Hubble and the James Webb Space Telescope, use a stationary system to block a star’s blinding light. The approach does help, but it’s a bit like putting your thumb over a flashlight while searching a dark room for a firefly. Though the bulb vanishes, stray glare can still escape and overwhelm the light of the insect. Inside a telescope, that glare can come from light leaking around the edges of machinery or from minuscule imperfections in mirrors and coatings that can scatter starlight into speckles. All this can hide, or even impersonate, a planet. Roman’s coronagraph, however, will attempt something completely unseen in space telescopes until this year: Before each observation, it will measure that leftover light and try to suppress it, a technique known as active wavefront control. The telescope is able to do this because it contains two deformable mirrors. Each has a 48-by-48 checkerboard of actuators (tiny pistons) beneath a thin, deformable sheet of glass. Applying a small amount of voltage makes the actuators contract and tug their patches of mirror slightly backward, like thousands of microscopic fingers delicately sculpting a surface. The effect is very subtle: Each patch of mirror can deform by up to 0.5 micrometers, or about one-fourth the size of an E. coli bacterium, and in increments as small as approximately 10 picometers. That’s about a tenth the diameter of a hydrogen atom, says Ilya Poberezhskiy, the instrument’s project systems engineer at JPL. The actuators allow the mirrors to create an “active wavefront,” where each component is moved to the perfect position to cancel out incoming waves of unwanted light—a bit like a pair of noise-canceling headphones, but for light instead of sound. The “canceled-out” light creates a “doughnut-shaped region around the star where we suppress starlight and where we’re hoping to see exoplanets,” says Poberezhskiy. Compared with current space-based coronagraphs, the system is expected to improve sensitivity to exoplanets against the glare of their host stars by a factor of up to 1,000, revealing planets that would have been far too faint to detect before. Like Hubble and JWST, Roman also uses masks, patterned plates placed in the path of the light that are designed to block the photons that run into them. One tool in Roman’s mask arsenal is “silicon grass,” a thicket of microscopic spikes on some masks that can be used in certain configurations to absorb photons so they don’t bounce around the telescope and accidentally reach a detector. Light entering the forest bounces deeper and deeper between the blades and gets trapped instead of reflecting back toward the camera. “Once the light gets into there, it never gets out,” Poberezhskiy says. The mirrors and masks form a succession of gates and hedges to guide as much of the preserved planetary light as possible toward the final detector. Alien Jupiters This elaborate setup could open a new chapter in the direct imaging of exoplanets. Nearly all exoplanets photographed so far are oversize youngsters that are nothing like the residents of our solar system: several times the mass of Jupiter, still glowing with the heat left over from their birth, and orbiting tens or hundreds of times farther from their star than the Earth is from the sun. This is because they are relatively easy to see. Their size, warmth, and distance from their parent star makes them shine brightly in infrared light, far away from the worst of the stellar glare. Roman, however, could directly image a true Jupiter analogue—a planet similar to Jupiter in mass and circling a sunlike star a few times farther out than Earth is from our sun. Unlike the hot Jupiters we can see now, this one would be a much more mature gas giant like ours, primarily reflecting its parent star’s light after billions of years of cooling instead of heavily emitting its own. Astronomers have been able to infer the existence of such planets from the gravitational wobble they impart to the star. Roman instead will collect starlight reflected from the planet itself. “We’re not looking

Shape-shifting mirrors on NASA’s new space telescope could unveil Jupiters like our own Lire l’article »

We use cookies to improve your experience and performance on our website. You can learn more at Politique de confidentialité and manage your privacy settings by clicking Settings.

Privacy Preferences

You can choose your cookie settings by turning on/off each type of cookie as you wish, except for essential cookies.

Allow All
Manage Consent Preferences
  • Always Active

Save
fr_FR