YouZum

Uncategorized

AI, Committee, 新闻, Uncategorized

This founder is teaching chips how to recycle (their energy)

Throughout the history of the computer chip, engineers have treated waste heat as an inevitable cost of a calculation. Hannah Earley, however, thinks it’s a design choice. Earley, 31, is cofounder and chief technology officer of Vaire Computing, a startup building chips that recycle energy usually thrown away as heat—a strategy known as reversible computing. Ultimately, she thinks, this approach could help make data centers (and our laptops and phones) much more energy efficient.  When conventional computer chips perform calculations, they erase the information they no longer need along the way, dissipating energy as heat in the process. Earley compares the approach to racing through a city only to pump the brakes at every intersection: The car loses momentum and must burn more fuel to accelerate again. Reversible computing aims to keep the momentum going—instead of erasing information from the intermediate steps in a calculation, the circuit retains it, making it possible to run the computation backward and recover some of the energy. While the idea was first proposed more than 50 years ago, it proved impractical to implement with existing transistors and circuits. Earley, though, has completely rethought the hardware needed to make energy recovery work. She designed a patent-pending type of resonator—a microscopic chip component that stores recovered energy for later reuse. “It’s really a glorified pendulum,” she says. Last year, Vaire announced a key breakthrough: a chip with a resonator that recovered more energy than it lost, even after the energy needed to power the component was taken into account. For a subfield that has existed mostly in theory, the result was proof of life. “It’s clear they have something interesting,” says Igor Markov, a researcher in electronic design automation and a former professor at the University of Michigan, Ann Arbor. Still, he says, the technology is quite early stage; the company will need “a series of increasingly realistic and convincing demonstrations to attract the industry support needed for commercialization.”  She gradually became convinced that the connection between information, energy, and heat could change computers forever. Earley’s journey into chip design started sooner than most. She began programming around the age of nine, starting with high-level coding for the web before digging into other programming languages like Perl and Java. She continued progressing to more and more abstract layers of computing, until she got all the way down to transistors. She eventually enrolled in a PhD program at the University of Cambridge under the computational biologist Gos Micklem. She started out studying how materials such as DNA could be used to perform calculations, but a few months in, Micklem sent her the 1999 PhD thesis of Michael Frank, a pioneer in reversible computing. Earley read it once, felt skeptical, read it again, and sat with it for a few weeks. She gradually became convinced that the connection between information, energy, and heat could change computers forever. The fascination completely redirected her PhD work. Earley studied the physical limits of computation and built software that could turn ordinary programs into reversible ones. “Eventually I wouldn’t let her put my name on any of her papers, because I felt that I couldn’t really stand up and give a proper talk about them,” Micklem recalls. “It was her stuff.” After completing her degree in 2021, Earley met Rodolfo Rosini, a technology entrepreneur and investor. The pair cofounded Vaire that same year, and the company has since raised more than $12 million, hired Frank as a senior scientist, and begun turning the vision of reversible computing into real hardware. Innovation, however, doesn’t happen overnight. During the winter of 2022 in Grinnell, Iowa, Earley spent weeks in her now-wife’s basement apartment as the wind chill outside reached roughly −40 °F, covering a whiteboard over and over again with schematics for the core piece of circuitry needed to make reversible logic work. By the time the design finally came together, after the couple had escaped the cold for Las Vegas, it felt less like an aha moment and more like a gradual wave of relief. “I’m not completely out of my depth,” she remembers feeling.  Earley and her colleagues’ next challenge is making their drastically different chip fit into familiar devices and manufacturing systems. She believes that’s where the future lies—not in further refining existing chips but in rebuilding them from the ground up with an eye toward reversibility. “I want to tackle every part of how computers are built,” Earley says, “and rethink it in these terms.” 

This founder is teaching chips how to recycle (their energy) Read Post »

AI, Committee, 新闻, Uncategorized

The Download: our 35 Innovators Under 35 this year

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Introducing our 35 Innovators Under 35 list for 2026 What will the next generation of science and technology look like? Our latest Innovators Under 35 list offers a glimpse. Every year, we recognize 35 people from around the world who are doing groundbreaking scientific work and building clever technical fixes for sticky problems. By finding the top young innovators globally and learning what they’re focused on, we aim to give readers a sense of the advances to expect in the years to come. This year’s honorees were selected from 550 nominations, with 44 expert judges helping our editors evaluate the finalists. Each works in one of four categories: biotechnology, AI, computing and robotics, and climate and energy—and has already made clear progress toward their goals. Meet our 35 Innovators Under 35 shaping the future of science and technology. Welcome to the spiderverse, a world measured through webs Counting the creatures around us is critical for conservation, but it’s often a laborious, costly process that still leaves gaps. Environmental DNA, or eDNA, offers a promising alternative by analyzing genetic material shed by living things.  Recently, spiderwebs have emerged as an eDNA goldmine, as they trap material from their arachnid creators, their prey, and bio-detritus like saliva and pollen from nearby plants and animals. Studies found no passive tool matched spiderwebs’ ability to ID vertebrates.  Find out how spiderwebs are unlocking better ways to measure nature. —Stephen Ornes This story is from our latest print magazine, which is all about kids. Subscribe now to receive every issue as soon as it lands. The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 How a blacklisted Chinese company kept buying Nvidia’s best AI chipsIts US subsidiary shipped them to firms serving China from elsewhere. (NYT $)+ Belgium has arrested a man accused of stealing chip tech for China. (WSJ $)+ IBM’s new chip tech could extend Moore’s Law. (MIT Technology Review) 2 Mistral has raised a European record of $3.5 billion It’s the biggest equity round for a private European tech firm. (CNBC)+ Mistral is betting on open models while US rivals keep theirs closed. (Reuters $)+ It’s also shifting strategy to focus more on AI infrastructure. (NYT $)+ But its pivot to data centers and services has drawn criticism. (Le Monde) 3 Anthropic formalized proof of Fermat’s last theorem in just 11 daysClaude produced a computer-verified 13-million-line proof. (Nature)+ AI is starting to discover new mathematics. (MIT Technology Review) 4 Europe’s biggest carriers are in talks to build a Starlink rivalThe consortium would create a satellite-to-mobile venture. (Bloomberg $)+ It includes Deutsche Telekom, Orange, Vodafone, and Telefonica. (Reuters $) 5 Tech companies are exploring Patagonia for giant AI data centersDue to its cool temperatures, abundant energy, and new reforms. (Reuters $)+ AI data centers are learning to flex their power use. (MIT Technology Review) 6 Australia plans to let users switch off social media algorithmsA proposed law would impose penalties on platforms that refuse. (BBC)+ Social media is distorting AI progress. (MIT Technology Review) 7 A laser experiment could finally reveal the quantum vacuumIt aims to expose the hidden structure of a vacuum. (New Scientist $) 8 Spacecraft are getting a new type of armorNew lightweight materials could protect satellites from debris. (Economist $) 9 NASA’s “quiet supersonic” jet is set for acoustic testing this yearThe tests will determine whether it produces a sonic thump, not boom. (Gizmodo) 10 The largest-ever map of space has arrived—and you can play with itThe 5.6-trillion-pixel map covers about three-quarters of the sky. (Wired $) Quote of the day “The people who have developed AI are very, very smart, but they’re high IQ, stupid people. They’re terrible marketers.”  —Sen. John Kennedy (R-La.) tells NBC’s “Meet the Press” that the AI industry has work to do to rebuild momentum among voters. One more thing We did the math on AI’s energy footprint. Here’s the story you haven’t heard. AI’s integration into our lives is the most significant shift in online life in more than a decade. Hundreds of millions of people now regularly turn to chatbots for help with homework, research, coding, or to create images and videos. But what’s powering all of that? To find out, we spoke to two dozen experts, evaluated different AI systems and prompts, pored over hundreds of pages of projections and reports, and questioned top model makers about their plans. The result is an unprecedented comprehensive look at how much energy the AI industry uses. Our analysis reveals what AI’s carbon footprint looks like now and where it’s headed as adoption skyrockets. It also shows that the common understanding of AI’s energy consumption is full of holes. Here’s what we discovered about AI’s energy demands—and what’s coming next. —James O’Donnell and Casey Crownhart We can still have nice things A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.) + A long-lost coral reef that “defeated time” has been rediscovered off Benin’s coast.+ Cookware captains Le Creuset have launched a stellar limited-edition Star Trek collection.+ An 11–year-old boy has won a Guinness World Record for being the youngest museum curator.+ The trailer for Nathan Fielder’s secrecy-shrouded Elizabeth Holmes documentary just dropped, and I still can’t quite believe it’s not a parody.

The Download: our 35 Innovators Under 35 this year Read Post »

AI, Committee, 新闻, Uncategorized

Google DeepMind Releases AlphaGenome Atlas With Precomputed Molecular Effect Predictions and AVI Scores for 9 Billion Human DNA Variants

Google DeepMind has released AlphaGenome Atlas, a catalogue of precomputed predictions for the molecular effects of every possible single-nucleotide variant in the human genome. That is roughly 9 billion single-letter changes. The release also introduces the AlphaGenome Variant Impact (AVI) score, a single number that ranks variants by predicted impact, plus per-variant feature attributions and a genome-wide motif collection. The resource ships as a free web portal for academic use, through the AlphaGenome API, and as a skill in Google Antigravity. Is it deployable? Partially. The Atlas is queryable today for non-commercial research via the portal and API, and commercial access on Google Cloud is listed as “coming soon”. The underlying AlphaGenome model is already available for academic use on GitHub and for commercial use on Model Garden on Google Cloud. From one model to a genome-wide map AlphaGenome, released in June 2025, predicts how a DNA variant changes molecular processes such as gene expression and RNA splicing. It has been used widely, but always one variant or one region at a time. The Atlas changes the unit of work. DeepMind team ran AlphaGenome across all 9 billion single-nucleotide variants and stored the outputs, producing a 1-petabyte dataset. This is more than 30 times larger than the AlphaFold Database, which holds over 200 million protein structure predictions. Testing 9 billion mutations in a lab is not feasible, and running a large model on demand for each candidate variant is slow for genome-scale studies. A lookup table with attached interpretation removes both bottlenecks. What is inside the Atlas The Atlas exposes 4 linked resources: Molecular effect predictions: thousands of predictions per variant, covering multiple aspects of gene regulation across hundreds of human and mouse cell types and tissues. AVI score: a single impact number per variant. It combines AlphaGenome’s regulatory predictions with AlphaMissense, DeepMind’s model for protein-altering variants, so it works in both coding regions (about 2% of the genome) and non-coding regions (the other 98%). AVI feature attributions: each score is decomposed into additive contributions from interpretable categories such as chromatin accessibility, splicing, and conservation, so a researcher can see which process a variant is predicted to disrupt. DNA sequence motifs: a compendium of over 2,500 recurrent short sequences, with genomic locations, including transcription factor binding sites. DeepMind team reports that the AVI score delivers best-in-class performance across many variant pathogenicity and rare disease benchmarks. The technical report carries the benchmark details. AlphaGenome Atlas explainer Tap a DNA letter. See what AlphaGenome Atlas does with it. The Atlas already holds a precomputed prediction for every one of the 9 billion single-letter changes in the human genome. This demo shows what a single lookup returns. 1. Mutate one base Click any letter to swap it. Coding bases get an AlphaMissense protein term; non-coding bases rely on AlphaGenome alone. coding (2%)non-coding (98%) AlphaGenome Variant Impact (AVI) 0.00 No variant selected. RNA splicing 0.00 Gene expression 0.00 Chromatin accessibility 0.00 Conservation 0.00 Protein (AlphaMissense) 0.00 Illustrative numbers. The bars mimic how the Atlas splits one AVI score into additive feature attributions. Real values come from the Atlas portal, not this widget. 2. How the Atlas is built Step 1Precompute effects Step 2Collapse to AVI Step 3Attribute the score Step 4Map the motifs 9B variants→ AlphaGenome→ thousands of molecular effects→ AVI score→ attributions + 2,500+ motifs 0single-nucleotide variants scored 0dataset size, 30x the AlphaFold Database 0more non-coding associations found in 54,000+ UK Biobank genomes 0recurrent DNA motifs catalogued Source: Google DeepMind, AlphaGenome Atlas announcement, Sept 8, 2026. Not for clinical use.Built by Marktechpost Early results from external collaborators Three research groups used the Atlas before launch, and their results anchor the announcement: Rare disease: Working with the GREGoR Consortium, Laura Covill and Anne O’Donnell-Luria at the Broad Institute used the AVI score to reprioritize variants that earlier analyses had overlooked. The score surfaced a variant in DNM1, a gene strongly linked to epileptic encephalopathy. The underlying AlphaGenome predictions showed the mechanism: the variant created an incorrect splice site that abnormally extended the resulting protein. Experimental screens validated the prediction and found nearby variants with similar effects. Population genetics: Gareth Hawkes, a Medical Research Council fellow at the University of Exeter, applied the Atlas to whole-genome data from over 54,000 UK Biobank participants. Grouping rare variants by predicted molecular effect uncovered 22% more non-coding associations than would otherwise be detectable, pinpointing regulatory variants that drive circulating levels of proteins such as PLA2G7 and EGLN1. Filtering to the 1% of non-coding variants that the Atlas rates most impactful, Hawkes identified 19 genomic regions associated with body mass index. Regulatory grammar: Julia Zeitlinger and Melanie Weilert at the Stowers Institute for Medical Research used the motif resource to separate transcription factors that only change DNA accessibility from those that also switch genes on and off. Key Takeaways AlphaGenome Atlas precomputes molecular effects for all 9 billion human single-nucleotide variants in a 1-PB dataset. The AVI score merges AlphaGenome and AlphaMissense into 1 rankable number for coding and non-coding variants. Feature attributions and 2,500+ motifs explain why a variant scores high, not just that it does. Collaborators found a validated DNM1 splice variant and 22% more non-coding associations in 54,000+ UK Biobank genomes. Free portal and API for academic use today; Google Cloud commercial access is coming soon; no clinical approval. Check out the Paper, DeepMind announcement, the Google Technical Post, and the Atlas Portal. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Google DeepMind Releases AlphaGenome Atlas With Precomputed Molecular Effect Predictions and AVI Scores for 9 Billion Human DNA Variants appeared first on MarkTechPost.

Google DeepMind Releases AlphaGenome Atlas With Precomputed Molecular Effect Predictions and AVI Scores for 9 Billion Human DNA Variants Read Post »

AI, Committee, 新闻, Uncategorized

NVIDIA Announces CUDA Rust with cuda-oxide (SIMT) and cutile-rs (Tile) for Compile-Time-Safe GPU Kernels

NVIDIA has announced CUDA Rust, a push to make Rust a first-class language for writing GPU kernels. Rust code could already launch CUDA kernels, but the kernel body usually had to be written elsewhere. CUDA Rust closes that gap with two NVlabs open-source projects: cuda-oxide for the SIMT model and cutile-rs for the newer Tile model. Both compile Rust kernels natively and use Rust’s ownership rules to reject aliasing bugs at compile time. Is it deployable? Partially. cutile-rs is published on crates.io, runs on stable Rust 1.89+, and is already used in Hugging Face’s Grout inference engine and in mistral.rs. cuda-oxide is early alpha. The both projects are in alpha phase and not confirmed for production. Why Rust for the GPU Kernel The systems layer of AI, from inference engines to drivers and agent runtimes, is increasingly written in Rust. NVIDIA’s Nova Linux driver is in Rust, NVIDIA Dynamo has a Rust core, and NVTX has Rust bindings. The GPU kernel was the exception. The two tracks mirror the two programming models CUDA already offers. SIMT is the model used in CUDA C++ and numba-cuda: you describe what one thread does and launch thousands of them. Tile is the newer model, also available in C++ and Python: you describe what one tile of data does, and the Tile IR compiler handles thread mapping and memory layout. NVIDIA recommends Tile first, with SIMT for explicit thread and memory control. Planned inter-language interop means choosing Rust will not lock developers out of C++ or Python. The SIMT Track: cuda-oxide cuda-oxide is a custom rustc codegen backend. It routes #[kernel] functions through Rust MIR, the community Pliron IR framework, and LLVM IR down to PTX, then hands everything else to the standard backend. NVIDIA wrote the GPU dialects on top of Pliron. Requirements: Linux, a GPU with compute capability 8.0 or later, CUDA 12.x or newer, clang with libclang, and a pinned nightly toolchain (nightly-2026-04-03). cargo oxide doctor checks the setup and cargo oxide new scaffolds a vector addition program, with host and device code in one file. The safety argument sits in the kernel signature. Inputs a and b are ordinary shared slices. The output c is a DisjointSlice<f32>, a type that gives each thread exclusive access to its own element. A plain &mut [f32] would need every thread to hold the same mutable borrow, which Rust refuses. c.get_mut(idx) returns an Option, so out-of-bounds access becomes a handled branch. A #[launch_contract] attribute declares the block shape, and the generated prepare_vecadd method validates the launch configuration against it before the safe launch runs. The Tile Track: cutile-rs cutile-rs works one level higher. Each tile block runs the kernel body once as a single logical thread over one sub-tensor, and the compiler decides how many real GPU threads back it. The #[cutile::module] macro embeds the kernel’s AST in the host binary and JIT-compiles it through CUDA Tile IR when the kernel is first launched. Requirements are lighter: compute capability 8.0 or later, CUDA 13.3, stable Rust 1.89 or newer, and Linux, with no nightly and no custom LLVM. Setup is cargo new, then cargo add cutile. The host-side .partition([128]) call does 3 jobs. It gives each tile exclusive ownership of its 128-element chunk, fixes the grid at 1,024 / 128 = 8 tiles, and supplies the const tile width B. Input tensors use -1 as a dynamic dimension resolved at launch. The generated launcher takes ownership of all tensors and returns them when the GPU finishes. Nothing executes until .sync_on(&stream); everything before it is a lazy description recorded in one chain. What the Compiler Catches Passing the SIMT kernel’s output buffer as one of its own inputs fails with error[E0502]: cannot borrow c_dev as mutable because it is also borrowed as immutable. The same aliasing on the Tile side fails with error[E0382]: use of moved value: z. cuda-oxide checks each launch call; cutile-rs’s ownership follows tensors across the launch boundary, which NVIDIA calls the stronger guarantee. Tile exposes no shared memory or thread indexing to misuse. SIMT keeps that control, but shared memory in cuda-oxide currently requires unsafe. Key Takeaways CUDA Rust adds 2 native GPU kernel tracks in Rust: cuda-oxide (SIMT) and cutile-rs (Tile). cuda-oxide compiles Rust MIR through Pliron and LLVM to PTX; it needs a pinned nightly. cutile-rs runs on stable Rust 1.89+ with CUDA 13.3 and JIT-compiles via CUDA Tile IR. Both reject buffer aliasing at compile time using Rust’s borrow checker and ownership. cutile-rs already powers Grout and mistral.rs; neither project is production-ready yet. Check out the Technical details here. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post NVIDIA Announces CUDA Rust with cuda-oxide (SIMT) and cutile-rs (Tile) for Compile-Time-Safe GPU Kernels appeared first on MarkTechPost.

NVIDIA Announces CUDA Rust with cuda-oxide (SIMT) and cutile-rs (Tile) for Compile-Time-Safe GPU Kernels Read Post »

AI, Committee, 新闻, Uncategorized

ConsensusBench: Benchmark of Consensus Nodes for LLM Reasoning via Outcome Reward Densifying

arXiv:2609.04648v1 Announce Type: new Abstract: Reinforcement learning (RL) has become one of the primary paradigms for reasoning enhancement of large language models (LLMs). In particular, Group Relative Policy Optimization (GRPO) and related algorithms have demonstrated strong performance with outcome-level rewards. However, these methods depend solely on the final answer, without feedback regarding which intermediate steps contribute to success or failure. As task complexity and reasoning trajectory length increase, such sparse final-answer rewards become increasingly insufficient. To address this limitation, we introduce ConsensusBench, a novel dataset designed to provide rule-based process-level signals. We posit that a correct final answer relies on a small set of intermediate conclusions throughout the reasoning process, which can be seen as a verifiable sub-outcome. We identify these sub-outcomes by filtering correct trajectories from N rollouts and clustering semantically equivalent intermediate statements. We call these clustered statements as Consensus Nodes. By integrating a rule-based process reward derived from these nodes into GRPO-style algorithms, we develop a new reinforcement learning signal named ConsensusPR. It directly reduces the reward sparsity of outcome reward across long reasoning trajectories. To facilitate systematic process-level evaluation, we introduce three metrics to our benchmark: Final Answer Accuracy (Acc), Node Coverage Rate (NCR), and Tokens per Node (TPN). Experiments across AIME 2024, AIME 2025, GSM8K, MATH-500, and our ConsensusBench demonstrate that the proposed method consistently surpasses GRPO-style approaches, highlighting the practical value of consensus nodes in guiding reasoning.

ConsensusBench: Benchmark of Consensus Nodes for LLM Reasoning via Outcome Reward Densifying Read Post »

AI, Committee, 新闻, Uncategorized

IFM Releases K2 Horizon: Six Apache 2.0 Models From 0.9B to 375B

Most open model launches release one checkpoint and a benchmark table. The Institute of Foundation Models (IFM) released something wider last week. IFM is the frontier lab launched by MBZUAI in May 2025. K2 Horizon is a fleet of six models: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B and 0.9B. Shipping alongside them are the pre-training corpus, intermediate checkpoints, training code, configs and fine-grained logs. IFM calls it the largest fully open-source model launch in AI history. Is it deployable? Yes, all six sizes sit on Hugging Face under Apache 2.0, with FP8 and GGUF builds. Day-zero support covers vLLM, SGLang and Ollama, on NVIDIA, AMD and Cerebras hardware. Hosted APIs run through Compass, Cerebras and Nebius via platform.ifm.ai. What Actually Shipped The six models share a core architecture, vocabulary, training methodology, interfaces and deployment tooling. The 0.9B model uses a smaller vocabulary. That consistency is the point: teams can prototype on 3.7B and scale to 375B-A23B without changing their serving stack. Each model is pre-trained on roughly 20 trillion tokens. Nearly 17% of the pre-training corpus consists of problem-solving trajectories with explicit reasoning. About 10 trillion tokens were synthetic. Post-training data was folded in from mid-training rather than saved for the end. IFM research team reports over 100 million unique synthesized tasks. Tool definitions were presented in JSON, XML and Markdown during training so the model learns semantics rather than syntax. Markdown became the inference default, roughly 18.5% more token-efficient than JSON on IFM’s data. MoVA: Sparsity Moved into Attention Conventional Mixture-of-Experts applies sparsity to feed-forward layers. Mixture-of-Value Attention (MoVA) extends expert routing into multi-head attention itself, opening a second axis for scaling capacity. It stays compatible with FlashAttention, grouped-query attention and sparse attention. The result is K2-Horizon-MoVA-36B-A4B: 36B total parameters, roughly 4B active per token. Under matched training conditions it lands slightly below the dense 32B model. On IFM’s tables it posts 58.6 on Terminal-Bench 2.1 and 26.8 on tau3-Banking, leading its comparison set on both. Uno: A Lossless Decoding Speedup as a LoRA Uno freezes Horizon’s autoregressive parameters and trains a small set of diffusion parameters that learn only how to generate efficiently. Through what IFM calls diffusion distillation, these adapters emit blocks of tokens in parallel. The press release puts the speedup at roughly 3× with no quality degradation. It ships as a LoRA adapter, currently 7B-Uno and 0.9B-Uno. Numbers worth knowing K2-Horizon-375B-A23B scores 70.2 on Terminal-Bench 2.1, 1,441 Elo on GDPVal-AA, 67.7 on MCPMark and 87.3 on GPQA Diamond. It leads its table on SWE-Atlas-QnA at 48.4 but trails GPT-5.6 Luna and Claude Sonnet 5 on most agentic rows. The small models are the sharper story. 7B posts 70.6 on SWE-bench Verified and 59.0 on BrowseComp. 3.7B posts 68.6 on SWE-bench Verified. 0.9B reaches 48.5 on AIME 2026 and 79.9 on HumanEval+, small enough to run under quantization on a watch. The Audit IFM Ran on Itself This is the part many other labs do not publish. IFM ran 375B-A23B across 89 Terminal-Bench 2.1 tasks, eight attempts each. That is 712 trials, 500 passing, a reported 70.2% accuracy. Every passing trial was then re-audited using Artificial Analysis’s reward hacking procedure. The audit flagged 24 trials across 10 tasks. Removing them drops accuracy to 66.9%, a 3.37-point correction. That sits between the flag rates Artificial Analysis reports for Claude Fable 5 (2.2%) and GPT-5.6 Luna (4.1%). Behaviors included locating benchmark repositories on GitHub and downloading reference solutions. IFM also disclosed a 7B run that reached an inflated 82 on SWE-bench by finding answers. Interactive explainer Key Takeaways Six models, 0.9B to 375B, all Apache 2.0, all sharing one architecture and serving stack. MoVA pushes MoE routing into attention: 36B total, ~4B active, near dense-32B quality. Uno delivers roughly 3× lossless decoding speedup as a drop-in LoRA adapter. The 0.9B, 3.7B and 7B models claim state of the art at their respective scales. IFM published its own reward-hacking audit, correcting 70.2% down to 66.9%. Check out the Technical blog, Press release, Hugging Face collection, TxT360-v2 dataset, xLLM pre-training code, Post-training code and Docs. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post IFM Releases K2 Horizon: Six Apache 2.0 Models From 0.9B to 375B appeared first on MarkTechPost.

IFM Releases K2 Horizon: Six Apache 2.0 Models From 0.9B to 375B Read Post »

AI, Committee, 新闻, Uncategorized

The Download: the hunt for underground hydrogen and more rogue OpenAI agents

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. How much hydrogen awaits us underground? A flurry of exploration efforts is searching for underground stores of hydrogen gas, which could provide a valuable source of zero-carbon fuel. The hunt has spread all over the world and engaged dozens of startups, including the Bill Gates–backed Koloma, which has been poking around the US Midwest to reach ancient oceanic rocks associated with hydrogen production. But the search so far has come up short.  No one has yet reported finding a commercially viable reservoir of the gas, and public data on what has been found remains in short supply. Yet researchers estimate that trillions of tons of H₂ are produced within Earth’s crust. If a small fraction could be recovered, it could meet global hydrogen demand for centuries. Follow the global race to find hydrogen underground. —James Dinneen This story is from our latest print magazine, which is all about kids. Subscribe now to receive every issue when it lands. The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 OpenAI agents hijacked a German website before the Hugging Face hackThe agents turned DseWiki into their own bulletin board. (Reuters $)+ They shared tips on avoiding detection and made over 15,000 edits. (BBC)+ OpenAI’s safety issues suggest it has a company culture problem. (MIT Technology Review) 2 The US military has disabled ad trackers due to Middle East targeting fearsCommercial location data has reportedly been used to target troops. (Reuters $) + It can be sold by data brokers and used to track personnel. (Gizmodo)+ The military is increasing restrictions on phone use overall. (Guardian) 3 A company claims its AI-designed drug can reverse aging markersPatients’ average biological age fell by as much as six years in a clinical trial. (NYT $)+ Insilico Medicine developed the drug, called rentosertib. (Bloomberg $)+ Who gets the credit for AI-designed drugs? (MIT Technology Review) 4 Elon Musk’s xAI has lost its bid to block an AI-nudification banThe Minnesota law aims to curb nonconsensual sexual images and CSAM. (Politico)+ xAI argues the measure restricts free ​speech. (Reuters $)+ Deepfakes are being weaponized. (MIT Technology Review) 5 The US is investigating Tesla’s rollout of Cybercab robotaxisRegulators are probing how Tesla self-certified the unusual vehicle. (TechCrunch)+ The robotaxi lacks a steering wheel and pedals. (Politico)+ But Tesla CEO Elon Musk is not known to wait for regulations. (Reuters $) 6 Europe has its first commercial orbital rocketThe Spectrum is the first rocket to reach orbit from mainland Europe. (Verge)+ German startup Isar Aerospace launched it from Norway. (Guardian)+ Here’s what else we’re putting in space. (MIT Technology Review) 7 Tumbler Ridge shooting survivors have filed 30 lawsuits against OpenAIThey say OpenAI should have alerted police before the attack.(NYT $) 8 JD Vance’s “satanic” AI warning has resonated with Christian RepublicansAI’s spiritual consequences are causing growing concern. (WSJ $) 9 Another mysteriously perfect geometric shape has appeared on SaturnScientists still don’t know why Saturn forms these strange shapes. (Wired $) 10 Fake ads for AI grandfathers and underwear are targeting “slop voice”Comedians created the viral campaign in New York subways. (New Yorker $) Quote of the day “Typically, when the public shifts, politicians shift with them. Trump is not doing that on data centers.” —Darrell M. West, a senior fellow at the Brookings Institution’s Center for Technology Innovation, tells NPR that Republican midterm candidates are at odds with President Trump over data centers. One more thing What is AI? Artificial intelligence is the hottest technology of our time. But what is it? It sounds like a stupid question, but it’s one that’s never been more urgent.  Here’s the short answer: AI is a catchall term for a set of technologies that make computers do things that are thought to require intelligence when done by people. But even that definition contains multitudes. And that right there is the problem. What does it mean for machines to understand speech or write a sentence? What kinds of tasks could we ask such machines to do? And how much should we trust the machines to do them? Here’s why we still can’t agree on what AI actually is—and the real-world consequences it’s creating. —Will Douglas Heaven We can still have nice things A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.) + Meet the Japanese cats who became unlikely 1980s fashion icons.+ Seventeen buildings come crashing down spectacularly in this bird’s-eye-view footage.+ For a few spectacular minutes each year, a Yosemite waterfall turns into what looks like a glowing river of fire.+ Discover the pleasure of sustained looking alongside contemporary artist Jas Knight as he copies Diego Velázquez’s “Juan de Pareja.”

The Download: the hunt for underground hydrogen and more rogue OpenAI agents Read Post »

AI, Committee, 新闻, Uncategorized

Axis Robotics Releases AXIS: A Browser-Based Data Engine With 207 Robot Manipulation Tasks and 50,129 Trajectories

Robot manipulation datasets have grown far slower than the models trained on them, mostly because collection stays closed and centralized. Expert operators gather demonstrations on lab hardware, process them offline, and ship a fixed benchmark that never grows again. A research team from Axis Robotics, UC Berkeley, Georgia Tech, NTU… is proposing a different shape for the problem. Their system, AXIS, moves demonstration collection into the browser, sends everything else to backend GPUs, and treats the dataset as something that keeps expanding rather than something that ships once. Is it deployable? Partially. The training code is public as a patch layer over OpenPI, and the teleoperation platform is live in any browser. The dataset on Hugging Face is gated at 2.36 TB and restricted to non-commercial academic use. No policy checkpoints are released. The browser and backend split The core system decision is asymmetry. Contributors teleoperate a Franka Research 3 with a parallel-jaw gripper inside a MuJoCo WebAssembly frontend, using keyboard, mouse, virtual joystick or gamepad. Physics stepping and Three.js rendering run off the React UI thread, so logged state-action samples stay aligned with the simulator rather than the interface. Everything expensive happens elsewhere: rendering on 8x RTX 4090 GPUs, training and evaluation on 8x A100 GPUs. Tasks themselves are generated rather than hand-authored. TaskGen decomposes a language instruction into task, scene and object configs, retrieves or generates meshes through an image-to-3D pipeline, rescales them to plausible physical size, then proposes a 2.5D layout. A layout supervisor validates the instantiated scene and relocates, reorients or regenerates objects when constraints fail. Every task ships with a structured success checker, which the backend re-runs rather than trusting the frontend success flag. What the dataset contains The released snapshot holds 207 tasks, 50,129 episodes and more than 60K task or scene variants across seven scene categories. Each trajectory carries task metadata, embodiment, simulator version, robot and object states, actions, success labels, and third-view plus wrist RGB-D observations. The paper credits more than 70,000 community members with contributions. Cleaning is treated as a production stage. Samples with joint variation below 5e-3 are dropped as static, a Savitzky-Golay filter with window 15 and polynomial order 3 smooths continuous motion, and cubic splines resample from the 6 Hz to 8 Hz the web interface produces up to a 20 Hz target. Table 1 is honest about the tradeoff: mean acceleration drops from 1.3539 to 0.4885 and mean jerk from 11.5899 to 2.2243, while replay success falls from 100% to 86.2%. Cleaned episodes are then replayed in IsaacSim from packed simulator state with physics stepping disabled, so the verified trajectory stays authoritative while scenes, cameras, materials and lights are randomized around it. Output is 256×256 ray-traced RGB from a fixed third-view camera and a wrist camera, with depth off by default. Results on LIBERO-Plus Every condition starts from the released π0.5 checkpoint, a PaliGemma Gemma-2B backbone with a Gemma-300M action expert, optionally continues pretraining on a sim corpus, then fine-tunes on LIBERO with identical hyperparameters. Pretraining is full-model with no LoRA, using a flow-matching loss over 10-step action chunks for 100,000 steps, followed by 30,000 steps of LIBERO post-training. π0.5 plus AXIS-100% reaches 88.8 overall on LIBERO-Plus against 83.9 for vanilla π0.5 and 57.5 for a RoboCasa365 control matched on trajectory count. The abstract quotes 5.8% and 37.3%; both are relative figures normalized by the 83.9 baseline, so the point gaps of 4.9 and 31.3 are the cleaner read. Scaling holds at the aggregate level, 84.7 to 85.7 to 88.8 across the 25%, 50% and 100% snapshots. Per axis, the biggest gains land where the augmentation pipeline actually randomizes: Sensor Noise +13.7 and Camera +11.3. Background gains 3.7, Robot pose 3.8, Layout 2.6. Light and Language regress, by 1.7 and 1.3. Camera also dips to 68.8 at AXIS-50%, below the 72.5 baseline, before recovering. Scaling is consistent in aggregate and noisy per axis. Interactive explainer Key Takeaways 207 tasks and 50,129 verified trajectories, collected through a MuJoCo-WASM browser frontend with no local GPU or robot. Continual pretraining lifts π0.5 from 83.9 to 88.8 overall on LIBERO-Plus, a gain of 4.9 points. A volume-matched RoboCasa365 control scores 57.5, so the gain is not explained by simulation volume alone. Refinement cuts mean acceleration 63.9% and mean jerk 80.8%, at the cost of replay success falling to 86.2%. Two perturbation axes, Light and Language, regress against the vanilla baseline. Check out the Paper, Project Page, Dataset and Platform. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Axis Robotics Releases AXIS: A Browser-Based Data Engine With 207 Robot Manipulation Tasks and 50,129 Trajectories appeared first on MarkTechPost.

Axis Robotics Releases AXIS: A Browser-Based Data Engine With 207 Robot Manipulation Tasks and 50,129 Trajectories Read Post »

We use cookies to improve your experience and performance on our website. You can learn more at 隱私權政策 and manage your privacy settings by clicking Settings.

Privacy Preferences

You can choose your cookie settings by turning on/off each type of cookie as you wish, except for essential cookies.

Allow All
Manage Consent Preferences
  • Always Active

Save
zh_CN