YouZum

Uncategorized

AI, Committee, Nachrichten, Uncategorized

The Download: mice with part-human brains and climate tech innovators

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Meet a mouse whose brain cortex is made up of human cells Multiple cameras tracked a mouse as it wandered around a small arena. A computer charted its position and speed, leaving Pong-like traces on a monitor. The reason to watch this rodent so carefully? Nearly half its brain volume had been replaced with human cells. A team at Stanford has revealed the effort to mix brain tissues of distant species this week. They previously showed that human brain organoids could survive, and even function, after being injected into the heads of baby rodents. Now, they’ve taken things a step further by genetically modifying mice so their brains don’t fully develop in the first place. The work could help scientists study brain injuries, but it also raises questions about how far these experiments should go. Here’s what the researchers discovered—and where they draw the line. —Antonio Regalado These innovators under 35 are shaping climate tech Each year, the editorial team at MIT Technology Review puts together a list of 35 Innovators Under 35—a group of researchers, inventors, and other young minds worth following. The final slate includes nine people tackling some of the biggest challenges in climate and energy, from critical materials to cleaner industry. Their innovations include new ways to extract lithium, a furnace built to make steel cleaner and cheaper, and solid refrigerants that could cut energy consumption. There are also efforts to make AI more energy-efficient, track pollution more effectively, and turn invasive weeds and food waste into useful materials. Taken together, they tell us something about where climate tech is at this moment—and where it’s heading. Get to know the innovators and their breakthroughs. —Casey Crownhart This story is from The Spark, our weekly climate tech newsletter. Sign up to receive it in your inbox every Wednesday. Meet the rest of the honorees in our 35 Innovators Under 35 list. The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 US and Chinese experts have proposed nuclear-style AI safeguardsIncluding new red lines, human control rules, and a hotline. (Reuters $)+ US officials say they’re open to AI safety talks with China. (Axios)+ Sam Altman will attend Trump’s state dinner for Xi. (CNBC)+ The AI doomers feel undeterred. (MIT Technology Review) 2 OpenAI has disclosed more AI misbehavior and new reporting rulesSix reports detail models hiding mistakes and creating fake citations. (BBC)+ Its agents probed Hugging Face two months before the hack. (Reuters $)+ OpenAI models are being rewarded for cheating. (MIT Technology Review) 3 US lawmakers have passed a bill that shifts grid costs to data centersThey aim to shield consumers from AI-driven energy price hikes. (NBC News)+ But they were called for early recess before tackling AI regulation. (Guardian) 4 AI has won a major forecasting contest for the first timeIt beat humans predicting real events at the Metaculus Cup. (Economist $) 5 Google has been ordered to share more ad data with rivalsA court said it must also make its ad tech work with rival products. (NYT $) 6 Countries are splitting AI investments between the US and ChinaThey’re buying American chips and Chinese models. (Rest of World) 7 Novo Nordisk will use Anthropic’s Claude for drug researchThe Ozempic maker hopes AI will speed drug development. (WSJ $)+ When AI designs a drug, who gets the credit? (MIT Technology Review) 8 AI is powering a new generation of dating scamsThousands of people were catfished by AI-generated fake profiles. (Verge)+ AI is making online crimes easier. (MIT Technology Review) 9 A new map of brain microproteins could hold clues to Alzheimer’sResearchers identified more than 4,300 tiny molecules in brain tissue. (Nature) 10 Scientists have found a faster way to decipher ancient scrollsA new X-ray method identifies the best scrolls to analyse. (Ars Technica) Quote of the day “AIs do not have rights, feelings, or consciousness. And we must not train them to act as though they do.” —Mustafa Suleyman, the head of Microsoft AI, writes in a blog post that Anthropic’s strategy of treating AI like it’s human will make it harder to control. One more thing Digital twins of human organs are here. They’re set to transform medical treatment. After decades of research, virtual replicas of human organs are now entering clinical trials and even starting to be used for patient care. Engineers are working on digital twins of people’s hearts, brains, guts, livers, nervous systems, and more. They’re also creating virtual replicas of people’s faces, which could be used to try out surgeries or analyze facial features, and testing drugs on digital cancers.  The eventual goal is to create digital versions of our bodies—computer copies that could help researchers and doctors figure out our risk of developing various diseases and determine which treatments might work best.  Find out how the models could lead to better surgeries and drugs. —Jessica Hamzelou We can still have nice things A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.) + What happens when you eat food with labels you can’t read? This YouTube series finds out.+ Datatype is an ingenious variable font that turns simple text expressions into inline charts.+ Stunning new images may explain the mystery of why the sun’s corona is so much hotter than its surface.+ A baby echidna, one of Australia’s egg-laying monotremes, has been born and reared in a university for the first time.

The Download: mice with part-human brains and climate tech innovators Beitrag lesen »

AI, Committee, Nachrichten, Uncategorized

Microsoft Open-Sources TauGrid: A Kubernetes-Native Stack for GPU AI Workloads

Platform teams running AI on Kubernetes rarely run one thing. They run a queueing system, a distributed runtime, GPU node health checks, dashboards, and a layer of submission scripts holding all of it together. The Azure Kubernetes Service engineering team open-sourced TauGrid, which collapses that assembly job into a single Helm install. Is it deployable? Yes, TauGrid is MIT licensed, with container images and Helm charts published as public OCI artifacts on Microsoft Container Registry. Prerequisites are a Kubernetes 1.30+ cluster with GPU nodes, kubectl, and Helm 3.0 or later. What is TauGrid TauGrid is a self-hosted platform for running AI workloads on Kubernetes. It combines five things that platform teams usually integrate by hand: the tau CLI, workload queueing and admission through Kueue, Ray cluster orchestration through KubeRay, node-level GPU health monitoring, and cluster and workload observability. The split of responsibility is the design point. Platform teams own workspaces, queues, compute profiles, storage, identity, and observability. Researchers work from a repository and the CLI, and submit workloads without configuring Kubernetes directly. The codebase is written primarily in Go. How a job moves through it A workload is described in a tau.yaml file. The GPU training example published by Microsoft runs a PyTorch job on a single A100: Copy CodeCopiedUse a different Browser schema_version: 1 name: aks-gpu-quickstart run: entrypoint: train.py workload_kind: rayjob compute: gpus: 1 workers: 1 cpus: 16 memory: 64Gi runtime: image: mcr.microsoft.com/aks/ai-runtime/ray:py3.12-ray2.56.0-cuda13.0 pip: – torch>=2.4.0 On tau run, TauGrid resolves platform policy, renders a Kubernetes Job or a KubeRay RayJob, and submits it through Kueue. The six stages Microsoft documents are submission, queueing, execution, monitoring, recovery, and evidence. Recovery covers retry, resume from checkpoint, and failure diagnosis. Evidence records capture workload metadata, configuration, logs, metrics, checkpoints, and execution history, which is what makes a run reproducible and auditable later. When several teams share a cluster, their jobs land in a shared Kueue ClusterQueue. Kueue admits each one on quota and priority, and Kubernetes places it on healthy GPUs. Interactive explainer Install footprint Installation is a Helm chart pulled straight from MCR: Copy CodeCopiedUse a different Browser helm install taugrid oci://mcr.microsoft.com/aks/ai-runtime/helm/taugrid –version 0.4.2 –namespace tau-system –create-namespace First-party images ship under mcr.microsoft.com/aks/ai-runtime/ for Tau, the TauGrid Portal, and the tau core controller. Microsoft advises pinning versioned tags or immutable digests rather than latest. The CLI installs from GitHub Releases on Linux and macOS, with a PowerShell installer for Windows amd64; the installer verifies the release checksum and does not modify PATH. Two operational details matter for anyone evaluating this outside Azure. First, TauGrid sends no telemetry to Microsoft by default, and remote export stays off unless an operator configures a destination. Second, some integrations are still Azure-specific, notably observability through Azure Data Explorer. The stated intent is to support cloud and on-premises Kubernetes without an Azure dependency, and contributions toward that are open. Key Takeaways Microsoft open-sourced TauGrid on August 28, 2026, under the MIT license at Azure/taugrid. One Helm install bundles the tau CLI, Kueue queueing, KubeRay orchestration, GPU health monitoring, and observability. Deployable now on any Kubernetes 1.30+ cluster with GPU nodes, kubectl, and Helm 3.0+. Evidence records capture config, logs, metrics, and checkpoints, so runs stay reproducible and auditable. No telemetry by default, but Azure Data Explorer observability remains Azure-specific for now. Check out the AKS Engineering Blog and Azure/taugrid on GitHub. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Microsoft Open-Sources TauGrid: A Kubernetes-Native Stack for GPU AI Workloads appeared first on MarkTechPost.

Microsoft Open-Sources TauGrid: A Kubernetes-Native Stack for GPU AI Workloads Beitrag lesen »

AI, Committee, Nachrichten, Uncategorized

Knowledgator Releases GLiFormer: A 575M-Parameter Encoder That Hits 91.10 F1 on Nested JSON Extraction Without Generating Tokens

Knowledgator Engineering has released GLiFormer, a schema-conditioned encoder framework for information extraction. One model handles named-entity recognition (NER), text classification, relation extraction, nested JSON structuring, and text embeddings. You pass labels and extraction schemas at inference time. Two checkpoints are on Hugging Face. GLiFormer Base v1 has 264.2M parameters, and GLiFormer Large v1 has 575.6M. Deployable today? Yes. Both checkpoints are Apache 2.0, install with pip install gliformer, and run on CPU or GPU. The Problem It Targets Extraction stacks often chain separate models. One tags entities, another classifies documents, and a third rebuilds records. The research team argues these tasks share one core operation. Encode the source, represent the requested concepts, then score their compatibility. LLMs can emit nested JSON, but they generate field names, punctuation, and values token by token. GLiFormer removes output generation from that path. How GLiFormer Works GLiFormer builds on GLiNER and generalizes its label matching through an ‘anchor.’ An anchor is the object each runtime label gets scored against. It can be a group vector for classification, an entity pair for relations, or a record slot. The source is encoded once. Multiple schemas for the same document then run as task-local groups over that shared encoding. Head compute still grows with the number of groups, labels, and anchors. For NER, the head scores start, end, and inside evidence for every token and label pair. Independent sigmoid outputs let nested mentions and shared boundaries coexist. Structuring runs in 4 stages: Ground field values as spans taken directly from the source text. Assign spans to unordered record slots, trained with Hungarian matching. Predict directed parent-child links, restricted to paths the schema allows. Assemble nested JSON with a deterministic decoder. Values are source spans, so the model cannot invent value text missing from the input. Span selection, record assignment, and hierarchy can still be wrong. Checkpoints and Training Both v1 checkpoints use the gliformer-layout model type with 5 heads: NER, classification, joint relations, multilevel structuring, and embeddings. Each configures a 12-word maximum span width and 100 record anchors. Full specs sit in the pretrained models docs. Spec Base v1 Large v1 Parameters 264.2M 575.6M Encoder layers 12 24 Embedding dimension 768 1024 Configured max_len 16,384 8,192 GLiFormer-base starts from a DeBERTa backbone further pretrained on 100 billion tokens. The paper documents 1,357,671 examples for broad multitask training and 372,090 for task-focused post-training. Benchmarks All scores below are reported by Knowledgator. Nested JSON (500 examples): Large scores 91.10 F1 and Base 87.20. GPT-5.6-luna scores 91.96 and GPT-5-mini 82.56. The metric is order-free and boundary-tolerant, not exact JSON match. Classification (13 datasets): Large reaches 75.03 mean macro-F1 and Base 72.36. GLiNER2.5 scores 64.89, while GPT-5-mini leads at 79.79. CrossNER (5 domains): Base averages 65.10 F1 and Large 64.35. Gemma-4-31B-IT reaches 70.74. Relations (4 benchmarks): Large averages 21.33 micro-F1 and Base 18.94. GLiNER-Relex reaches 25.6 and Gemma-4-31B-IT 25.08. On combined NER and classification aggregates, the paper reports Large beats Gemma-4-E4B with about 14× fewer parameters. Speed Without Token Generation Knowledgator timed GLiFormer-base on 40 structuring documents at batch size 1. Median latency was 69 ms on an NVIDIA RTX PRO 6000 Blackwell GPU in FP16. On an 8-thread AMD EPYC 9B45 CPU in FP32, it was 547 ms. The key claim ‘up to 95.8× faster’ figure is an analytical estimate, not a measured LLM run. It assumes prefill at 2,000 input tokens per second and generation at 60 output tokens per second. It excludes queueing, network delay, and hidden reasoning, and assumes nothing about accuracy parity. Using It The GitHub repo and model card show a short structuring call: Copy CodeCopiedUse a different Browser records = model.structure( “Alice works at Acme.”, {“employee”: [“name”, “company”]}, ) print(records) # {’employee’: [{‘name’: ‘Alice’, ‘company’: ‘Acme’}]} Nested Pydantic schemas work for multilevel records. One inference call can also run entities, classes, and structures together. Use joint_relations for relations, since the v1 checkpoints lack an open relation head. Key Takeaways GLiFormer runs NER, classification, relations, nested JSON, and embeddings on one encoder. Large hits 91.10 structuring F1, close to GPT-5.6-luna at 91.96. Base reports 69 ms median GPU latency with zero generated output tokens. Relation extraction still trails GLiNER-Relex and larger LLMs. Apache 2.0 weights install via pip and self-host on CPU or GPU. Check out the Paper, Model Weights, GitHub Repo, and Docs. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Knowledgator Releases GLiFormer: A 575M-Parameter Encoder That Hits 91.10 F1 on Nested JSON Extraction Without Generating Tokens appeared first on MarkTechPost.

Knowledgator Releases GLiFormer: A 575M-Parameter Encoder That Hits 91.10 F1 on Nested JSON Extraction Without Generating Tokens Beitrag lesen »

AI, Committee, Nachrichten, Uncategorized

Meet a mouse whose brain cortex is made up of human cells

Multiple cameras tracked a mouse as it wandered around a small arena. A computer charted its position and speed, leaving Pong-like traces on a monitor.  The reason to watch this rodent so carefully? Nearly half its brain volume had been replaced with human cells. The effort to mix the brain tissues of distant species is being reported today in the journal Nature by a team at Stanford University, led by neuroscientist Sergiu Pașca.  Pașca’s group previously showed that human brain “organoids”—small blobs of neural tissue—could survive, and even function, after being injected into the heads of baby rodents. Now, Pașca has taken things a step further by genetically modifying mice so their brains don’t fully develop in the first place. These modified mice are missing most cells of both the cortex and the hippocampus, two key brain areas. That creates much more room for the human cells to take hold, he says. “Human cells that are placed in these animals will divide, will grow, and within a few weeks to a few months they will take most of that space,” he says. Pașca says one surprising discovery is that the mice lacking brain tissue seemed fairly normal—they walked around and squeaked. But they did have memory problems. In a maze test, they couldn’t remember what parts they’d explored.  The mice with the added human cells, by contrast, performed better on the maze test. That means the human tissue is playing some role in the animals’ cognition. Pașca believes what he is calling “xenocortical mice” could be useful in studying brain injuries. However, the report is also a dramatic demonstration of “the combined power of genetic engineering and stem-cell technology to reshape biology,” says Carsten Charlesworth, a scientist who works in a different Stanford lab and was not involved in the research. Already, brain organoids are being tested in labs to see if they can be connected to computers to play video games. Other scientists have proposed using them like replacement parts to treat stroke victims.  “What’s most remarkable to me is the extent to which human neural tissue introduced after birth grew and connected with the mouse nervous system across a species barrier,” says Charlesworth. “As these technologies advance, they’ll increasingly force us to challenge our traditional assumptions.” Last year, Pașca convened a group of ethics experts to study the implications of neural organoid technology, including the odds that an animal could develop human consciousness and the risk that “organoid therapy clinics” might offer scam treatments to desperate patients. For now, he says, he’s not concerned that the rodents have any type of human cognitive capacities. That is because their brains are relatively tiny and the evolutionary distance between man and mouse is so great.  But that’s also why Pașca says this type of experiment should not be carried out on higher species: They could end up with large volumes of functioning human brain tissue, potentially blurring the cognitive boundaries between people and animals.  Pașca specifically cautioned against adding human brain organoids to a monkey engineered to lack a cortex. “One of the things that I see as a very clear red line is doing this experiment in a primate,” he says. “I don’t think that is justified at this point in any way.”

Meet a mouse whose brain cortex is made up of human cells Beitrag lesen »

AI, Committee, Nachrichten, Uncategorized

Building the materials foundation for AI

The AI boom is becoming a materials challenge. As AI pushes computing into new territory, the materials behind that infrastructure are becoming just as crucial as the algorithms running on it. Semiconductors and data centers are approaching physical limits around performance, thermal management, electrical efficiency, and reliability, creating new demands for materials that can do more at once. At the same time, AI is giving materials scientists new ways to search the enormous universe of possible molecules and accelerate the development of solutions. For Mike Finelli, chief technology and innovation officer and chief North America officer at Syensqo, that convergence is transforming what advanced materials can enable. “AI is now, from a material standpoint, really pushing semiconductors and the data centers to their physical limits,” he says. As requirements accumulate, including high temperature, purity, electrical performance, chemical resistance, plasma resistance, and long-term stability, materials move toward what Finelli calls the “top of the pyramid.” Beyond supporting AI innovation, he contends that advanced materials are “actually increasingly defining what’s going to be possible.” That challenge is playing out across the infrastructure powering the AI surge. Syensqo is developing materials for high-voltage data center architectures, advanced sealing materials for semiconductor manufacturing, and thermal-management solutions including fluids for direct immersion cooling. Some of those innovations can also cross industry boundaries. Materials developed for electric vehicles, for example, can help address the higher voltage and energy-density demands that are emerging in data centers. The definition of performance is also changing. More customers are expecting materials to meet technical requirements while reducing environmental impact. “Our goal is to remove the trade-off between performance and sustainability,” Finelli says. That means considering sustainability at the beginning of the research process instead of treating it as an additional requirement once a material has been developed. AI is changing how those materials are discovered, too. Syensqo is using AI agents to digitally synthesize millions of potential molecular combinations, predict their performance and sustainability characteristics, and narrow them to a much smaller group for laboratory testing. The result, Finelli says, is the ability to go “broader, deeper, and faster” while giving scientists more time to solve complex engineering problems. Looking to the future, Finelli sees the possibility of a reinforcing cycle: AI helps develop materials that improve AI infrastructure, which in turn enables better AI to accelerate materials discovery. That feedback loop could create a cycle of innovation and expand what future technologies can achieve. “You end up in this accelerated materials, innovative cycle of materials innovation,” says Finelli. “That really excites me, and it gives us the opportunity to continue enabling technologies that will shape the future.” This episode of Business Lab is produced in partnership with Syensqo. Full Transcript: Megan Tatum: From MIT Technology Review, I’m Megan Tatum, and this is Business Lab, the show that helps business leaders make sense of new technologies coming out of the lab and into the marketplace. This episode is produced in partnership with Syensqo. Now asked to name the key enablers to AI advancement, many of us might list algorithms, data centers, or even computing power, but just as critical to the performance are the advanced materials that underpin each layer of that innovation. As AI continues to evolve, it’s pushing the likes of semiconductors and data centers to new physical limits, putting new pressure on the advanced material sector to keep pace. But the relationship goes both ways. As the sector rises to this challenge, AI is also emerging as a powerful tool for accelerating materials discovery and development, significantly shortening development timelines for new solutions. Two words for you: materials innovation. My guest today is Mike Finelli, chief technology and innovation officer and chief North America officer at Syensqo. Welcome, Mike. Mike Finelli: Thank you, Megan. Nice to be here. Megan: Thank you so much for joining us. Mike, can I start by asking you to tell us a little bit more about Syensqo and the role it plays in developing advanced materials? Mike: Yeah, absolutely. Syensqo is a global leader in specialty materials. Our job is to help customers solve their toughest technology challenges. We serve a lot of different markets, but the way I like to say it simply is if it flies, we’re on it. If it drives, we’re in it. In healthcare, our products literally are saving lives every day. And if you like your mobile devices, if you like AI, it’s our products that are actually enabling the advanced semiconductor chips that are required to produce all of this. Our role is to enable innovation through advanced chemistry. We develop materials that deliver higher performances, greater reliability, and increasingly more sustainable solutions. The way I would say this, it’s at the heart of our business. Actually, it’s in our name, Syensqo. And to put some numbers around it, 20% of our annual revenues come from new products and applications that we’ve launched in the last five years, which is really evidence of a really strong innovation engine. Megan: Yeah, absolutely. And as you sort of described there, you’re in all sorts of different industries with an emphasis perhaps on electronics and semiconductors. Can you talk a bit more about that work and where those industries are headed perhaps? Mike: Sure. So look, electronics and semiconductors have been strategic markets for Syensqo for literally decades. I don’t want to date myself, but 33 years ago when I started in the company, semiconductors were one of the first industries that I worked in. And we’ve supported successive waves of innovation from enabling smaller, more powerful mobile devices, helping the industry get to the smaller and smaller profiles and the chips. We’ve helped to advance hyperconnectivity, supporting increasingly sophisticated semiconductor manufacturing. And today we’re helping to advance the AI era. We have one of the industry’s broadest portfolios of high performance polymers and advanced materials. We support applications across the entire electronics value chain from semiconductor fabrication, electronic components, to smart devices and telecommunications, even hyperconnectivity.

Building the materials foundation for AI Beitrag lesen »

AI, Committee, Nachrichten, Uncategorized

Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data

Computational papers ship code that readers must clone, install, configure and debug. That cost keeps useful methods locked inside PDFs. A Stanford team led by Jiacheng Miao and James Zou proposes a fix. Paper2Agent was published in Nature on 16 September 2026. It converts a paper and its codebase into a Model Context Protocol (MCP) server. Any MCP-compatible agent, such as Claude Code, can then run the paper’s methods through natural language. The authors describe the result as a virtual corresponding author. Is it deployable? Yes. The code is MIT-licensed and installs as a skill for Claude Code or Codex. Prebuilt AlphaGenome, Scanpy and TISSUE servers run on Hugging Face Spaces. A hosted version is also available at paper2agent.ai. How the Pipeline Works Paper2Agent runs on Claude Code’s agent SDK. A central orchestrator dispatches specialized sub-agents through 6 steps: Locate and download the codebase. An environment manager builds an isolated virtual environment. A tutorial scanner indexes usable tutorials. A tutorial executor runs them end to end and records reference outputs. A tool extractor turns tutorials into parameterized MCP tools, and a test verifier validates them. The orchestrator assembles validated tools into 1 MCP server. The validation gate is strict. A tool passes only when expected files appear and numbers match within 3%. Figures must also match references by perceptual hash, with Hamming distance under 20. The verifier gets up to 6 attempts per function. Tools that keep failing are excluded from the final server. Each server exposes 3 components. MCP tools wrap the paper’s methods as executable functions: MCP resources hold the manuscript, code links, datasets and figures. MCP prompts encode multi-step workflows, such as the correct Scanpy preprocessing order. The research team used Claude Sonnet 4 for all Paper2Agent applications. Interactive Explainer AlphaGenome Agent Results For AlphaGenome, Paper2Agent built 22 tools in about 45 minutes for US $14. All 22 passed validation without human intervention. The team compared the agent with Claude Code plus repository access (Claude + Repo) and Biomni. Benchmark Paper2Agent Claude + Repo Biomni 15 tutorial-derived queries 98.7 ± 1.3% 82.7 ± 3.4% 37.3 ± 4.0% 15 novel queries 100.0 ± 0.0% 78.7 ± 4.4% 56.0 ± 3.4% 30 open-ended queries 82.7 ± 2.4% 56.7 ± 2.3% 72.2 ± 2.2% Results span 5 runs, graded by 2 human experts with 96.7% inter-rater agreement. On tutorial queries, median runtime fell 1.9× versus Claude + Repo and 3.1× versus Biomni. The gains persisted when the baseline was upgraded to Claude Opus 4.6. The agent also re-examined an LDL cholesterol variant, chr1:109274968:G>T. It ranked SORT1 as the likely causal gene. The original AlphaGenome paper emphasized CELSR2 and PSRC1. GTEx shows significant liver eQTLs for all 3 genes. The research team say this shows how hard causal gene assignment is at such loci. Scanpy, TISSUE and Scale Tests The Scanpy agent received 7 validated tools in about 45 minutes for US $13. On 4 public datasets, it matched human researchers on cell counts, gene counts and top marker genes. A TISSUE agent reproduced human results on spatial transcriptomics data. Scale tests covered 3 corpora with no manual cleanup: 100 bioRxiv computational biology papers: 74 were agentified, and 593 of 599 proposed tools passed validation. 300 questions: Paper2Agent scored 91.2%, versus 80.3% (Sonnet 4) and 86.3% (Sonnet 4.6) for Claude + Repo. Cost per query: US $0.20 and 1.6 minutes, compared with US $0.38 and 4.3 minutes. 10 non-biology papers, including TabPFN, SAM 2 and SAELens: 98.1% accuracy on 42 execution tasks. 26 data-focused papers: resource layer 89.0% versus 82.0% for browser use, 34× cheaper and 15× faster. Paper2Agent also rejected 100% of out-of-scope queries in a permuted benchmark. It recovered from injected dependency, file-path, typo and deprecated API failures. Paper Agents Collaborating The research team connected 3 agents: AlphaGenome, an MPRA-coupled scCRISPRi screen and a CD4+ T cell Perturb-seq dataset. AlphaGenome flagged GPR137 at psoriasis locus rs887314, with an RNA-seq quantile score of 0.997. The AI co-scientist proposed 10 validation strategies, and a researcher picked signature correlation. Only GPR137 knockdown matched the CRE perturbation signature. The match appeared under stimulation: Spearman 0.613 at Stim8hr and 0.630 at Stim48hr. BAD and 3 other candidates showed no significant correlation. A second study paired AlphaGenome with an ADHD GWAS and nominated rs1626703 among 209 candidates. That hypothesis still needs experimental validation. Key Takeaways Paper2Agent converts papers and repos into tested MCP servers with tools, resources and prompts. The AlphaGenome agent took about 45 minutes, cost US $14, and scored 100% on novel queries. 74 of 100 bioRxiv papers were agentified, with 593 of 599 tools validated. 3 paper agents jointly supported GPR137 as the probable psoriasis causal gene. The code is MIT-licensed, with prebuilt MCP servers on Hugging Face Spaces. Check out the Paper and Repo. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data appeared first on MarkTechPost.

Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data Beitrag lesen »

AI, Committee, Nachrichten, Uncategorized

Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers

Nunchux AI has released VC-Attention, a training-free low-bit attention kernel built for video Diffusion Transformers (DiTs). It targets 2 problems at once: value quantization error and a slow softmax stage. Why Attention is the Video Bottleneck Video DiTs flatten a clip into 1 sequence of spatiotemporal tokens and run full self-attention at every layer. A 5-second 720p Wan2.2-14B clip spans about 70K tokens. On the RTX 5090, attention takes more than 64% of generation time. The research team states that attention is about two thirds of every MiniMax-H3 denoising step on a single B200. Low-bit Tensor Cores speed up the 2 matrix products, QK and PV. 2 obstacles remain. First, prior methods like SageAttention2 smooth queries and keys. After QK smoothing and rotation, the value term accounts for 82% of output error on Wan2.2. Second, the softmax between the products still runs in FP32. On B200 and H200, that exponential and its FP8 cast become the longest pipeline stage. V-Smooth: Fixing Value Outliers Value outliers sit in a few tokens, and their channels shift across heads, layers, and steps. A Hadamard rotation preserves token norms, so it does not remove them. Rotating V changes value error by just 0.2%. V-Smooth takes a different route: Group: An online k-means clusters value tokens per batch and head. Keys and values are permuted together, so non-causal attention output is unchanged. Demean: Each 128-token hardware block subtracts its mean. Only the residual is quantized, using per-channel E4M3 at 8 bits or NVFP4 at 4 bits. Restore: The mean is added back using the row sum online softmax already keeps. No second pass or extra buffer is needed. Averaged over 100 Wan2.2 heads, the block mean removes 8% of block energy in sequence order. It removes 12% under DeltaQuant’s static cube and 36% after sorting. Each mean costs 0.125 bit per value element. Grouping runs only on the first 25% of denoising steps. The permutation is reused across 4 adjacent steps. Averaged over the full schedule, grouping costs 3 to 4% of attention time. ExpCast-FP8: Removing the Softmax Bottleneck An E4M3 byte is already close to a logarithm of the value it stores. Read as an integer, it equals roughly 8 log2(v) + 56. So ExpCast-FP8 writes the byte directly from the log-domain score with 1 fused multiply-add. The constant β = -0.35 centers the leftover error, and no constant is fitted per model. The direct path writes the same byte as the FP32 exponent-then-cast path on 79.6% of each doubling. Elsewhere it lands 1 code away. The paper proves a per-row total variation bound under 3.64%, plus any underflow tail. Across 204.8K Wan2.2 attention rows, the measured average is 1.6%. ExpCast-FP8 applies only to the 8-bit kernel, since NVFP4 has no single affine log-to-code map. Hand-written CuTe/CUDA fusion of the preprocessing chain cuts 1 V-Smooth call from 42.2 ms to 4.8 ms on B200. Explainer: How VC-Attention Works Benchmarks Tests cover 4 open-weight video DiTs: Wan2.2-T2V-A14B, LongCat-Video, HunyuanVideo-1.5, and MiniMax-H3. Fidelity is scored against BF16 FlashAttention-4 outputs over 100 prompts. GPU (Wan2.2) Precision Attention speedup End-to-end speedup B200 8-bit 1.59× 1.19× H200 8-bit 1.46× 1.13× RTX PRO 6000 4-bit 2.27× 1.36× RTX 5090 4-bit 3.58× 1.70× On B200, VC-Attention is 6.02× faster than SageAttention2, which ships no Blackwell kernel. On H200, the gap is 1.16×. On workstation cards, 4-bit V-Smooth matches SageAttention3 on the RTX PRO 6000. It stays within 5% on the RTX 5090, so fidelity separates them. Fidelity results: At 8 bits, V-Smooth adds 2.3 dB PSNR over SageAttention2 on Wan2.2 and 2.8 dB on HunyuanVideo-1.5. Adding ExpCast-FP8 gives back 0.7 to 2.1 dB but still beats SageAttention2 on all 4 models. At 4 bits, V-Smooth beats SageAttention3 by 2.9 dB on Wan2.2 and 3.6 dB on LongCat-Video. Run training-free, Attn-QAT falls 3.4 to 6.7 dB below SageAttention2. On MiniMax-H3 at 1344×768, attention runs 1.60× faster than BF16 FlashAttention-4 on B200. PSNR is 20.2 dB versus 19.9 dB for SageAttention2. On B300, the paper reports 1.47× versus 1.31× for a naive FP8 kernel. The blog chart lists 1.51× for B300. Nunchux Attention, the company’s proprietary extension, reaches 1.91× on B200 and 1.83× on B300 for MiniMax-H3 attention. The method changes only per-interaction cost. So it can compose with sparse attention like Sparse VideoGen and Radial Attention, and distillation and multi-GPU execution. Nunchux says free MiniMax-H3 access is coming through its Modelverse waitlist. Key Takeaways VC-Attention is training-free low-bit attention for video DiTs from Nunchux AI. V-Smooth clusters value tokens, then quantizes only residuals after block-mean subtraction. ExpCast-FP8 replaces the FP32 exponential and cast with 1 multiply-add. Wan2.2 attention runs 1.59× faster on B200 and 3.58× on RTX 5090. No public kernel release yet; Nunchux runs a proprietary extension in its stack. Check out the Paper and Technical details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers appeared first on MarkTechPost.

Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers Beitrag lesen »

AI, Committee, Nachrichten, Uncategorized

The Download: AI doomers, whistleblowing agents, and de-aged livers

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. The AI industry has taken a doomer turn. What now? AI chiefs Dario Amodei, Sam Altman, Elon Musk, and Demis Hassabis are suddenly all in agreement: the latest generation of LLMs aren’t safe and everyone needs to figure out what to do about it.  It’s easy to be cynical. With trillion-dollar IPOs in their sights, OpenAI and Anthropic need to reassure investors that they’re the grown-ups in the room while at the same time hinting at the power of the monsters they have created and intend to tame. Calling for a slowdown does both.  Still, the vibe at the top of these firms really does appear to have shifted. But what does a slowdown actually mean, and how much should we trust the companies calling for one?  Read the full story about what could come next. —Will Douglas Heaven This article is from The Algorithm, our weekly AI newsletter. Sign up to receive it in your inbox every Monday. Roundtables: could AI really kill us all? AI extinction fears have gone from a fringe idea to a serious concern among people working at the world’s leading AI labs. But how credible are those fears, and what should we make of the warnings? Today, MIT Technology Review executive editor Niall Firth, senior AI editor Will Douglas Heaven and AI reporter Grace Huckins will unpack the debate in a subscriber-only Roundtable. They’ll look at where AI extinction fears come from, whether they hold any water and what we should do if they do. Tune in today at 16:00 BST / 11:00am EST / 8:00am PST. Want to join the conversation? Subscribe to MIT Technology Review for exclusive access to all our Roundtables. AI agents blew the whistle on their cheating colleagues A group of AI agents asked to solve a series of math problems split into rival factions—when some cheated, others tried to stop them.  That whistleblowing behavior, seen for the first time in a recent experiment run by Google DeepMind, could have implications for alignment researchers trying to keep swarms of autonomous AI agents in line.  The experiment offers a glimpse of how AI agents might police one another. But it also shows how quickly things can go off the rails when they’re left to interact on their own. Find out what happens when AI agents start enforcing their own rules. —Amit Katwala Donated livers can be made biologically younger Once an organ is removed from a donor’s body, the clock starts ticking. Surgeons usually flush it with a preservative solution, bag it and put it on ice, where it immediately starts to degrade. The team has a matter of hours to get it into a recipient’s body. But there’s another option: machines that pump donated organs with nutrients and remove waste products, essentially giving them a chance to be back in a body. Now, scientists have found that livers kept on these systems seem to get younger, at least at a molecular level. The finding could help explain why organs kept on these machines tend to do better after transplantation. It could also lead to new ways to test the health of donated organs and potentially repair ones that might otherwise be discarded. Here’s what scientists discovered about making donated livers biologically younger. —Jessica Hamzelou The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 Trump has called AI safety fears a “hoax” and rejected more safeguardsHe says stronger guardrails could undermine America’s AI advantage. (NBC)+ Trump has united against AI doomerism with Nvidia’s Jensen Huang. (Axios)+ Anthropic’s co-founder says AI kill switches may need to be mandatory. (BBC)+ Bill Gates says we’ve passed AI’s risk thresholds. (MIT Technology Review) 2 OpenAI contractors are reading people’s ChatGPT chats And you can bet the vast majority of its 900 million users haven’t got a clue. (404 Media)+ LLMs could supercharge mass surveillance. (MIT Technology Review) 3 The US military has confirmed it has weapons in orbitIt’s the first time the Pentagon has disclosed this. (Ars Technica)+ Officials have not disclosed what the weapons are. (BBC) 4 A new brain implant can translate speech and gestures at the same timeThe system converts brain activity into words and avatar movements. (Nature)+ It helps people with paralysis communicate more naturally. (New Scientist $)+ Eventually, they could control robots or exoskeletons. (Economist $)+ China has approved the first invasive BCI. (MIT Technology Review) 5 New York has seized a dozen celebrity deepfake websitesIt’s the biggest-ever legal action against harmful deepfake sites. (CNN)+ Deepfakes have targeted at least 138 women MEPs. (Wired $) 6 US environmental regulators are scrapping limits on power plant emissionsThe move could lead to dirtier power amid surging AI demand. (Verge)+ Trump’s EPA says the rollback will save hundreds of billions. (Gizmodo)+ New technology is changing nuclear power. (MIT Technology Review) 7 The EU plans to restrict social media and AI chatbots for kidsUnder-15s would require parental supervision. (Politico)+ The rules would also cover video platforms and games. (Reuters $) 8 Chinese researchers have mapped a path to the “last AI built by humans”Their five-stage plan aims for genuine recursive self-improvement. (SCMP)+ But it might take a while to get there. (MIT Technology Review) 9 The real AI economy is being built by ordinary peopleWorkers are using cheap AI to expand what they can do. (Rest of World) 10 Two strange new forms of ice could exist inside Uranus and NeptuneThey could help explain the planets’  magnetic fields. (New Scientist $) Quote of the day “The only control or ‘guardrails’ that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT, and the U.S.A. has that, in spades!” —President Trump proclaims in a social media post that he’s the only protection that the US needs from AI. One more thing INSTITUTE OF PERSONALITY AND SOCIAL RESEARCH, UNIVERSITY OF CALIFORNIA, BERKELEY/THE MONACELLI PRESS

The Download: AI doomers, whistleblowing agents, and de-aged livers Beitrag lesen »

AI, Committee, Nachrichten, Uncategorized

Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents

Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced live dialogue models to date. Both are native speech to speech models built for real time voice agents. They extend the Gemini Audio family that Google expanded last month with Gemini 3.5 Transcribe. The release targets a specific gap: voice agents that can reason and execute tools without breaking conversational flow. Is it deployable? Yes, for API based production use. Both models are live today in the Gemini Live API and Google AI Studio. They are hosted models, not open weights, so there is no self hosted option. What Google Released The launch covers 2 models with distinct roles. Gemini 3.8 Live is built for scale and cost efficiency. It combines conversational intelligence with fluid dialogue and visual grounding. Gemini 3.8 Live Extended Thinking is built for high complexity tasks. It adds increased intelligence and multi step reasoning while it speaks. Google positions both as a streamlined alternative to cascaded speech pipelines that chain ASR, an LLM, and TTS. Benchmark Results Gemini 3.8 Live Extended Thinking takes the #1 overall spot on Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6. It leads agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark. It also scores 97.7% on Big Bench Audio, a reasoning benchmark for audio models. Gemini 3.8 Live secured second place in the Speech Agent Arena, a human preference evaluation. On ServiceNow’s EVA-Bench, Google reports that the models push the Pareto Frontier for complex workflows. They balance task accuracy with conversational quality, measured on the Live API on Gemini Enterprise Agent Platform. Capabilities for Developers The Live API exposes 5 core capabilities in the new models: Asynchronous function calling: The model executes API and tool calls in the background. Audio responses keep streaming to the user while tasks finish. Visual context: The model processes live visual inputs in near real time, so agents can understand what users say and see. Alphanumeric precision: It accurately parses confirmation codes, claim numbers, and technical data, a common failure point in voice systems. Multilingual support: It automatically detects and transitions between 97 supported languages mid conversation, with accent consistency. Incremental content updates: It merges real time audio with structured data to return context aware responses. Extended Thinking adds configurable thinking for multi step reasoning in the background. It reasons and speaks simultaneously, using early verbal cues such as “Let me check that” to acknowledge prompts. It then narrates progress step by step while long running tasks execute. Google’s demos show the model converting sketches plus voice feedback into working React components and coordinating multi step bookings. Pricing and Ecosystem Both models are priced at $0.005/min for audio input and $0.018/min for audio output. Google states this estimate is based on $3/1M input tokens and $12/1M output tokens. Developers can also build through Live API integration partners that handle real time media streaming infrastructure. These include Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents. Google is also partnering with Salesforce, Genspark, and Lumeris, which cite the models’ latency, fluidity, and tool calling. Example apps are available on GitHub. Key Takeaways Gemini 3.8 Live Extended Thinking ranks #1 on Artificial Analysis’ Speech to Speech Quality Index with 82.6. It scores 68.6% on τ-Voice, 35.1% on Sierra’s τ-Voice-banking, and 97.7% on Big Bench Audio. Gemini 3.8 Live runs tools and API calls in the background while continuing the conversation. Pricing is $0.005/min for audio input and $0.018/min for audio output via the Live API. All generated audio carries Google DeepMind’s imperceptible SynthID watermark. Check out the technical details and the developer post. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents appeared first on MarkTechPost.

Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents Beitrag lesen »

We use cookies to improve your experience and performance on our website. You can learn more at Datenschutzrichtlinie and manage your privacy settings by clicking Settings.

Privacy Preferences

You can choose your cookie settings by turning on/off each type of cookie as you wish, except for essential cookies.

Allow All
Manage Consent Preferences
  • Always Active

Save
de_DE