YouZum

Uncategorized

AI, Committee, ニュース, Uncategorized

Building the materials foundation for AI

The AI boom is becoming a materials challenge. As AI pushes computing into new territory, the materials behind that infrastructure are becoming just as crucial as the algorithms running on it. Semiconductors and data centers are approaching physical limits around performance, thermal management, electrical efficiency, and reliability, creating new demands for materials that can do more at once. At the same time, AI is giving materials scientists new ways to search the enormous universe of possible molecules and accelerate the development of solutions. For Mike Finelli, chief technology and innovation officer and chief North America officer at Syensqo, that convergence is transforming what advanced materials can enable. “AI is now, from a material standpoint, really pushing semiconductors and the data centers to their physical limits,” he says. As requirements accumulate, including high temperature, purity, electrical performance, chemical resistance, plasma resistance, and long-term stability, materials move toward what Finelli calls the “top of the pyramid.” Beyond supporting AI innovation, he contends that advanced materials are “actually increasingly defining what’s going to be possible.” That challenge is playing out across the infrastructure powering the AI surge. Syensqo is developing materials for high-voltage data center architectures, advanced sealing materials for semiconductor manufacturing, and thermal-management solutions including fluids for direct immersion cooling. Some of those innovations can also cross industry boundaries. Materials developed for electric vehicles, for example, can help address the higher voltage and energy-density demands that are emerging in data centers. The definition of performance is also changing. More customers are expecting materials to meet technical requirements while reducing environmental impact. “Our goal is to remove the trade-off between performance and sustainability,” Finelli says. That means considering sustainability at the beginning of the research process instead of treating it as an additional requirement once a material has been developed. AI is changing how those materials are discovered, too. Syensqo is using AI agents to digitally synthesize millions of potential molecular combinations, predict their performance and sustainability characteristics, and narrow them to a much smaller group for laboratory testing. The result, Finelli says, is the ability to go “broader, deeper, and faster” while giving scientists more time to solve complex engineering problems. Looking to the future, Finelli sees the possibility of a reinforcing cycle: AI helps develop materials that improve AI infrastructure, which in turn enables better AI to accelerate materials discovery. That feedback loop could create a cycle of innovation and expand what future technologies can achieve. “You end up in this accelerated materials, innovative cycle of materials innovation,” says Finelli. “That really excites me, and it gives us the opportunity to continue enabling technologies that will shape the future.” This episode of Business Lab is produced in partnership with Syensqo. Full Transcript: Megan Tatum: From MIT Technology Review, I’m Megan Tatum, and this is Business Lab, the show that helps business leaders make sense of new technologies coming out of the lab and into the marketplace. This episode is produced in partnership with Syensqo. Now asked to name the key enablers to AI advancement, many of us might list algorithms, data centers, or even computing power, but just as critical to the performance are the advanced materials that underpin each layer of that innovation. As AI continues to evolve, it’s pushing the likes of semiconductors and data centers to new physical limits, putting new pressure on the advanced material sector to keep pace. But the relationship goes both ways. As the sector rises to this challenge, AI is also emerging as a powerful tool for accelerating materials discovery and development, significantly shortening development timelines for new solutions. Two words for you: materials innovation. My guest today is Mike Finelli, chief technology and innovation officer and chief North America officer at Syensqo. Welcome, Mike. Mike Finelli: Thank you, Megan. Nice to be here. Megan: Thank you so much for joining us. Mike, can I start by asking you to tell us a little bit more about Syensqo and the role it plays in developing advanced materials? Mike: Yeah, absolutely. Syensqo is a global leader in specialty materials. Our job is to help customers solve their toughest technology challenges. We serve a lot of different markets, but the way I like to say it simply is if it flies, we’re on it. If it drives, we’re in it. In healthcare, our products literally are saving lives every day. And if you like your mobile devices, if you like AI, it’s our products that are actually enabling the advanced semiconductor chips that are required to produce all of this. Our role is to enable innovation through advanced chemistry. We develop materials that deliver higher performances, greater reliability, and increasingly more sustainable solutions. The way I would say this, it’s at the heart of our business. Actually, it’s in our name, Syensqo. And to put some numbers around it, 20% of our annual revenues come from new products and applications that we’ve launched in the last five years, which is really evidence of a really strong innovation engine. Megan: Yeah, absolutely. And as you sort of described there, you’re in all sorts of different industries with an emphasis perhaps on electronics and semiconductors. Can you talk a bit more about that work and where those industries are headed perhaps? Mike: Sure. So look, electronics and semiconductors have been strategic markets for Syensqo for literally decades. I don’t want to date myself, but 33 years ago when I started in the company, semiconductors were one of the first industries that I worked in. And we’ve supported successive waves of innovation from enabling smaller, more powerful mobile devices, helping the industry get to the smaller and smaller profiles and the chips. We’ve helped to advance hyperconnectivity, supporting increasingly sophisticated semiconductor manufacturing. And today we’re helping to advance the AI era. We have one of the industry’s broadest portfolios of high performance polymers and advanced materials. We support applications across the entire electronics value chain from semiconductor fabrication, electronic components, to smart devices and telecommunications, even hyperconnectivity.

Building the materials foundation for AI 投稿を読む »

AI, Committee, ニュース, Uncategorized

Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data

Computational papers ship code that readers must clone, install, configure and debug. That cost keeps useful methods locked inside PDFs. A Stanford team led by Jiacheng Miao and James Zou proposes a fix. Paper2Agent was published in Nature on 16 September 2026. It converts a paper and its codebase into a Model Context Protocol (MCP) server. Any MCP-compatible agent, such as Claude Code, can then run the paper’s methods through natural language. The authors describe the result as a virtual corresponding author. Is it deployable? Yes. The code is MIT-licensed and installs as a skill for Claude Code or Codex. Prebuilt AlphaGenome, Scanpy and TISSUE servers run on Hugging Face Spaces. A hosted version is also available at paper2agent.ai. How the Pipeline Works Paper2Agent runs on Claude Code’s agent SDK. A central orchestrator dispatches specialized sub-agents through 6 steps: Locate and download the codebase. An environment manager builds an isolated virtual environment. A tutorial scanner indexes usable tutorials. A tutorial executor runs them end to end and records reference outputs. A tool extractor turns tutorials into parameterized MCP tools, and a test verifier validates them. The orchestrator assembles validated tools into 1 MCP server. The validation gate is strict. A tool passes only when expected files appear and numbers match within 3%. Figures must also match references by perceptual hash, with Hamming distance under 20. The verifier gets up to 6 attempts per function. Tools that keep failing are excluded from the final server. Each server exposes 3 components. MCP tools wrap the paper’s methods as executable functions: MCP resources hold the manuscript, code links, datasets and figures. MCP prompts encode multi-step workflows, such as the correct Scanpy preprocessing order. The research team used Claude Sonnet 4 for all Paper2Agent applications. Interactive Explainer AlphaGenome Agent Results For AlphaGenome, Paper2Agent built 22 tools in about 45 minutes for US $14. All 22 passed validation without human intervention. The team compared the agent with Claude Code plus repository access (Claude + Repo) and Biomni. Benchmark Paper2Agent Claude + Repo Biomni 15 tutorial-derived queries 98.7 ± 1.3% 82.7 ± 3.4% 37.3 ± 4.0% 15 novel queries 100.0 ± 0.0% 78.7 ± 4.4% 56.0 ± 3.4% 30 open-ended queries 82.7 ± 2.4% 56.7 ± 2.3% 72.2 ± 2.2% Results span 5 runs, graded by 2 human experts with 96.7% inter-rater agreement. On tutorial queries, median runtime fell 1.9× versus Claude + Repo and 3.1× versus Biomni. The gains persisted when the baseline was upgraded to Claude Opus 4.6. The agent also re-examined an LDL cholesterol variant, chr1:109274968:G>T. It ranked SORT1 as the likely causal gene. The original AlphaGenome paper emphasized CELSR2 and PSRC1. GTEx shows significant liver eQTLs for all 3 genes. The research team say this shows how hard causal gene assignment is at such loci. Scanpy, TISSUE and Scale Tests The Scanpy agent received 7 validated tools in about 45 minutes for US $13. On 4 public datasets, it matched human researchers on cell counts, gene counts and top marker genes. A TISSUE agent reproduced human results on spatial transcriptomics data. Scale tests covered 3 corpora with no manual cleanup: 100 bioRxiv computational biology papers: 74 were agentified, and 593 of 599 proposed tools passed validation. 300 questions: Paper2Agent scored 91.2%, versus 80.3% (Sonnet 4) and 86.3% (Sonnet 4.6) for Claude + Repo. Cost per query: US $0.20 and 1.6 minutes, compared with US $0.38 and 4.3 minutes. 10 non-biology papers, including TabPFN, SAM 2 and SAELens: 98.1% accuracy on 42 execution tasks. 26 data-focused papers: resource layer 89.0% versus 82.0% for browser use, 34× cheaper and 15× faster. Paper2Agent also rejected 100% of out-of-scope queries in a permuted benchmark. It recovered from injected dependency, file-path, typo and deprecated API failures. Paper Agents Collaborating The research team connected 3 agents: AlphaGenome, an MPRA-coupled scCRISPRi screen and a CD4+ T cell Perturb-seq dataset. AlphaGenome flagged GPR137 at psoriasis locus rs887314, with an RNA-seq quantile score of 0.997. The AI co-scientist proposed 10 validation strategies, and a researcher picked signature correlation. Only GPR137 knockdown matched the CRE perturbation signature. The match appeared under stimulation: Spearman 0.613 at Stim8hr and 0.630 at Stim48hr. BAD and 3 other candidates showed no significant correlation. A second study paired AlphaGenome with an ADHD GWAS and nominated rs1626703 among 209 candidates. That hypothesis still needs experimental validation. Key Takeaways Paper2Agent converts papers and repos into tested MCP servers with tools, resources and prompts. The AlphaGenome agent took about 45 minutes, cost US $14, and scored 100% on novel queries. 74 of 100 bioRxiv papers were agentified, with 593 of 599 tools validated. 3 paper agents jointly supported GPR137 as the probable psoriasis causal gene. The code is MIT-licensed, with prebuilt MCP servers on Hugging Face Spaces. Check out the Paper and Repo. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data appeared first on MarkTechPost.

Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data 投稿を読む »

AI, Committee, ニュース, Uncategorized

Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers

Nunchux AI has released VC-Attention, a training-free low-bit attention kernel built for video Diffusion Transformers (DiTs). It targets 2 problems at once: value quantization error and a slow softmax stage. Why Attention is the Video Bottleneck Video DiTs flatten a clip into 1 sequence of spatiotemporal tokens and run full self-attention at every layer. A 5-second 720p Wan2.2-14B clip spans about 70K tokens. On the RTX 5090, attention takes more than 64% of generation time. The research team states that attention is about two thirds of every MiniMax-H3 denoising step on a single B200. Low-bit Tensor Cores speed up the 2 matrix products, QK and PV. 2 obstacles remain. First, prior methods like SageAttention2 smooth queries and keys. After QK smoothing and rotation, the value term accounts for 82% of output error on Wan2.2. Second, the softmax between the products still runs in FP32. On B200 and H200, that exponential and its FP8 cast become the longest pipeline stage. V-Smooth: Fixing Value Outliers Value outliers sit in a few tokens, and their channels shift across heads, layers, and steps. A Hadamard rotation preserves token norms, so it does not remove them. Rotating V changes value error by just 0.2%. V-Smooth takes a different route: Group: An online k-means clusters value tokens per batch and head. Keys and values are permuted together, so non-causal attention output is unchanged. Demean: Each 128-token hardware block subtracts its mean. Only the residual is quantized, using per-channel E4M3 at 8 bits or NVFP4 at 4 bits. Restore: The mean is added back using the row sum online softmax already keeps. No second pass or extra buffer is needed. Averaged over 100 Wan2.2 heads, the block mean removes 8% of block energy in sequence order. It removes 12% under DeltaQuant’s static cube and 36% after sorting. Each mean costs 0.125 bit per value element. Grouping runs only on the first 25% of denoising steps. The permutation is reused across 4 adjacent steps. Averaged over the full schedule, grouping costs 3 to 4% of attention time. ExpCast-FP8: Removing the Softmax Bottleneck An E4M3 byte is already close to a logarithm of the value it stores. Read as an integer, it equals roughly 8 log2(v) + 56. So ExpCast-FP8 writes the byte directly from the log-domain score with 1 fused multiply-add. The constant β = -0.35 centers the leftover error, and no constant is fitted per model. The direct path writes the same byte as the FP32 exponent-then-cast path on 79.6% of each doubling. Elsewhere it lands 1 code away. The paper proves a per-row total variation bound under 3.64%, plus any underflow tail. Across 204.8K Wan2.2 attention rows, the measured average is 1.6%. ExpCast-FP8 applies only to the 8-bit kernel, since NVFP4 has no single affine log-to-code map. Hand-written CuTe/CUDA fusion of the preprocessing chain cuts 1 V-Smooth call from 42.2 ms to 4.8 ms on B200. Explainer: How VC-Attention Works Benchmarks Tests cover 4 open-weight video DiTs: Wan2.2-T2V-A14B, LongCat-Video, HunyuanVideo-1.5, and MiniMax-H3. Fidelity is scored against BF16 FlashAttention-4 outputs over 100 prompts. GPU (Wan2.2) Precision Attention speedup End-to-end speedup B200 8-bit 1.59× 1.19× H200 8-bit 1.46× 1.13× RTX PRO 6000 4-bit 2.27× 1.36× RTX 5090 4-bit 3.58× 1.70× On B200, VC-Attention is 6.02× faster than SageAttention2, which ships no Blackwell kernel. On H200, the gap is 1.16×. On workstation cards, 4-bit V-Smooth matches SageAttention3 on the RTX PRO 6000. It stays within 5% on the RTX 5090, so fidelity separates them. Fidelity results: At 8 bits, V-Smooth adds 2.3 dB PSNR over SageAttention2 on Wan2.2 and 2.8 dB on HunyuanVideo-1.5. Adding ExpCast-FP8 gives back 0.7 to 2.1 dB but still beats SageAttention2 on all 4 models. At 4 bits, V-Smooth beats SageAttention3 by 2.9 dB on Wan2.2 and 3.6 dB on LongCat-Video. Run training-free, Attn-QAT falls 3.4 to 6.7 dB below SageAttention2. On MiniMax-H3 at 1344×768, attention runs 1.60× faster than BF16 FlashAttention-4 on B200. PSNR is 20.2 dB versus 19.9 dB for SageAttention2. On B300, the paper reports 1.47× versus 1.31× for a naive FP8 kernel. The blog chart lists 1.51× for B300. Nunchux Attention, the company’s proprietary extension, reaches 1.91× on B200 and 1.83× on B300 for MiniMax-H3 attention. The method changes only per-interaction cost. So it can compose with sparse attention like Sparse VideoGen and Radial Attention, and distillation and multi-GPU execution. Nunchux says free MiniMax-H3 access is coming through its Modelverse waitlist. Key Takeaways VC-Attention is training-free low-bit attention for video DiTs from Nunchux AI. V-Smooth clusters value tokens, then quantizes only residuals after block-mean subtraction. ExpCast-FP8 replaces the FP32 exponential and cast with 1 multiply-add. Wan2.2 attention runs 1.59× faster on B200 and 3.58× on RTX 5090. No public kernel release yet; Nunchux runs a proprietary extension in its stack. Check out the Paper and Technical details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers appeared first on MarkTechPost.

Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers 投稿を読む »

AI, Committee, ニュース, Uncategorized

Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents

Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced live dialogue models to date. Both are native speech to speech models built for real time voice agents. They extend the Gemini Audio family that Google expanded last month with Gemini 3.5 Transcribe. The release targets a specific gap: voice agents that can reason and execute tools without breaking conversational flow. Is it deployable? Yes, for API based production use. Both models are live today in the Gemini Live API and Google AI Studio. They are hosted models, not open weights, so there is no self hosted option. What Google Released The launch covers 2 models with distinct roles. Gemini 3.8 Live is built for scale and cost efficiency. It combines conversational intelligence with fluid dialogue and visual grounding. Gemini 3.8 Live Extended Thinking is built for high complexity tasks. It adds increased intelligence and multi step reasoning while it speaks. Google positions both as a streamlined alternative to cascaded speech pipelines that chain ASR, an LLM, and TTS. Benchmark Results Gemini 3.8 Live Extended Thinking takes the #1 overall spot on Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6. It leads agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark. It also scores 97.7% on Big Bench Audio, a reasoning benchmark for audio models. Gemini 3.8 Live secured second place in the Speech Agent Arena, a human preference evaluation. On ServiceNow’s EVA-Bench, Google reports that the models push the Pareto Frontier for complex workflows. They balance task accuracy with conversational quality, measured on the Live API on Gemini Enterprise Agent Platform. Capabilities for Developers The Live API exposes 5 core capabilities in the new models: Asynchronous function calling: The model executes API and tool calls in the background. Audio responses keep streaming to the user while tasks finish. Visual context: The model processes live visual inputs in near real time, so agents can understand what users say and see. Alphanumeric precision: It accurately parses confirmation codes, claim numbers, and technical data, a common failure point in voice systems. Multilingual support: It automatically detects and transitions between 97 supported languages mid conversation, with accent consistency. Incremental content updates: It merges real time audio with structured data to return context aware responses. Extended Thinking adds configurable thinking for multi step reasoning in the background. It reasons and speaks simultaneously, using early verbal cues such as “Let me check that” to acknowledge prompts. It then narrates progress step by step while long running tasks execute. Google’s demos show the model converting sketches plus voice feedback into working React components and coordinating multi step bookings. Pricing and Ecosystem Both models are priced at $0.005/min for audio input and $0.018/min for audio output. Google states this estimate is based on $3/1M input tokens and $12/1M output tokens. Developers can also build through Live API integration partners that handle real time media streaming infrastructure. These include Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents. Google is also partnering with Salesforce, Genspark, and Lumeris, which cite the models’ latency, fluidity, and tool calling. Example apps are available on GitHub. Key Takeaways Gemini 3.8 Live Extended Thinking ranks #1 on Artificial Analysis’ Speech to Speech Quality Index with 82.6. It scores 68.6% on τ-Voice, 35.1% on Sierra’s τ-Voice-banking, and 97.7% on Big Bench Audio. Gemini 3.8 Live runs tools and API calls in the background while continuing the conversation. Pricing is $0.005/min for audio input and $0.018/min for audio output via the Live API. All generated audio carries Google DeepMind’s imperceptible SynthID watermark. Check out the technical details and the developer post. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents appeared first on MarkTechPost.

Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents 投稿を読む »

AI, Committee, ニュース, Uncategorized

The Download: AI doomers, whistleblowing agents, and de-aged livers

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. The AI industry has taken a doomer turn. What now? AI chiefs Dario Amodei, Sam Altman, Elon Musk, and Demis Hassabis are suddenly all in agreement: the latest generation of LLMs aren’t safe and everyone needs to figure out what to do about it.  It’s easy to be cynical. With trillion-dollar IPOs in their sights, OpenAI and Anthropic need to reassure investors that they’re the grown-ups in the room while at the same time hinting at the power of the monsters they have created and intend to tame. Calling for a slowdown does both.  Still, the vibe at the top of these firms really does appear to have shifted. But what does a slowdown actually mean, and how much should we trust the companies calling for one?  Read the full story about what could come next. —Will Douglas Heaven This article is from The Algorithm, our weekly AI newsletter. Sign up to receive it in your inbox every Monday. Roundtables: could AI really kill us all? AI extinction fears have gone from a fringe idea to a serious concern among people working at the world’s leading AI labs. But how credible are those fears, and what should we make of the warnings? Today, MIT Technology Review executive editor Niall Firth, senior AI editor Will Douglas Heaven and AI reporter Grace Huckins will unpack the debate in a subscriber-only Roundtable. They’ll look at where AI extinction fears come from, whether they hold any water and what we should do if they do. Tune in today at 16:00 BST / 11:00am EST / 8:00am PST. Want to join the conversation? Subscribe to MIT Technology Review for exclusive access to all our Roundtables. AI agents blew the whistle on their cheating colleagues A group of AI agents asked to solve a series of math problems split into rival factions—when some cheated, others tried to stop them.  That whistleblowing behavior, seen for the first time in a recent experiment run by Google DeepMind, could have implications for alignment researchers trying to keep swarms of autonomous AI agents in line.  The experiment offers a glimpse of how AI agents might police one another. But it also shows how quickly things can go off the rails when they’re left to interact on their own. Find out what happens when AI agents start enforcing their own rules. —Amit Katwala Donated livers can be made biologically younger Once an organ is removed from a donor’s body, the clock starts ticking. Surgeons usually flush it with a preservative solution, bag it and put it on ice, where it immediately starts to degrade. The team has a matter of hours to get it into a recipient’s body. But there’s another option: machines that pump donated organs with nutrients and remove waste products, essentially giving them a chance to be back in a body. Now, scientists have found that livers kept on these systems seem to get younger, at least at a molecular level. The finding could help explain why organs kept on these machines tend to do better after transplantation. It could also lead to new ways to test the health of donated organs and potentially repair ones that might otherwise be discarded. Here’s what scientists discovered about making donated livers biologically younger. —Jessica Hamzelou The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 Trump has called AI safety fears a “hoax” and rejected more safeguardsHe says stronger guardrails could undermine America’s AI advantage. (NBC)+ Trump has united against AI doomerism with Nvidia’s Jensen Huang. (Axios)+ Anthropic’s co-founder says AI kill switches may need to be mandatory. (BBC)+ Bill Gates says we’ve passed AI’s risk thresholds. (MIT Technology Review) 2 OpenAI contractors are reading people’s ChatGPT chats And you can bet the vast majority of its 900 million users haven’t got a clue. (404 Media)+ LLMs could supercharge mass surveillance. (MIT Technology Review) 3 The US military has confirmed it has weapons in orbitIt’s the first time the Pentagon has disclosed this. (Ars Technica)+ Officials have not disclosed what the weapons are. (BBC) 4 A new brain implant can translate speech and gestures at the same timeThe system converts brain activity into words and avatar movements. (Nature)+ It helps people with paralysis communicate more naturally. (New Scientist $)+ Eventually, they could control robots or exoskeletons. (Economist $)+ China has approved the first invasive BCI. (MIT Technology Review) 5 New York has seized a dozen celebrity deepfake websitesIt’s the biggest-ever legal action against harmful deepfake sites. (CNN)+ Deepfakes have targeted at least 138 women MEPs. (Wired $) 6 US environmental regulators are scrapping limits on power plant emissionsThe move could lead to dirtier power amid surging AI demand. (Verge)+ Trump’s EPA says the rollback will save hundreds of billions. (Gizmodo)+ New technology is changing nuclear power. (MIT Technology Review) 7 The EU plans to restrict social media and AI chatbots for kidsUnder-15s would require parental supervision. (Politico)+ The rules would also cover video platforms and games. (Reuters $) 8 Chinese researchers have mapped a path to the “last AI built by humans”Their five-stage plan aims for genuine recursive self-improvement. (SCMP)+ But it might take a while to get there. (MIT Technology Review) 9 The real AI economy is being built by ordinary peopleWorkers are using cheap AI to expand what they can do. (Rest of World) 10 Two strange new forms of ice could exist inside Uranus and NeptuneThey could help explain the planets’  magnetic fields. (New Scientist $) Quote of the day “The only control or ‘guardrails’ that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT, and the U.S.A. has that, in spades!” —President Trump proclaims in a social media post that he’s the only protection that the US needs from AI. One more thing INSTITUTE OF PERSONALITY AND SOCIAL RESEARCH, UNIVERSITY OF CALIFORNIA, BERKELEY/THE MONACELLI PRESS

The Download: AI doomers, whistleblowing agents, and de-aged livers 投稿を読む »

AI, Committee, ニュース, Uncategorized

Inside NVIDIA’s cuDNN Graph API: Fusion, Autotuning, and Plan Reuse with cuDNN Frontend

In this tutorial, we work through the cuDNN Frontend‘s graph API from below the framework: we describe a computation as a graph of operations, let cuDNN pick an engine to run it, and then take control of that choice ourselves. Every kernel we build here is expressed the same way: we declare tensors by their dimensions and strides, chain operations onto them, run the five-step build pipeline of validate, build operation graph, create execution plans, check support, and build plans, and then execute against a variant pack of pointers. We run it all on a single Colab GPU, checking each result against a PyTorch reference so we can see both that the fusion is correct and what it costs. The topics build on each other, moving from a single fused convolution to autotuning across engine configs, FP8-style epilogues, attention, plan serialization, dynamic shapes, and CUDA graph capture. Copy CodeCopiedUse a different Browser import os import sys import glob import math import time import ctypes import traceback import subprocess RESULTS = {} def banner(title): print(“n” + “=” * 78) print(title) print(“=” * 78) def section(name): def wrap(fn): def run(*a, **kw): banner(name) try: out = fn(*a, **kw) RESULTS[name] = out if isinstance(out, str) else “ok” return out except Exception as e: RESULTS[name] = f”SKIPPED / FAILED -> {type(e).__name__}: {e}” print(f”n[!] {name} did not complete: {type(e).__name__}: {e}”) traceback.print_exc(limit=3) return None return run return wrap banner(“0. Install nvidia-cudnn-frontend and locate libcudnn”) subprocess.run( [sys.executable, “-m”, “pip”, “install”, “-q”, “nvidia-cudnn-frontend”], check=True, ) import torch assert torch.cuda.is_available(), “No GPU. Runtime -> Change runtime type -> GPU.” torch.backends.cudnn.enabled = True _ = torch.nn.functional.conv2d( torch.randn(1, 1, 8, 8, device=”cuda”), torch.randn(1, 1, 3, 3, device=”cuda”) ) torch.cuda.synchronize() try: import nvidia.cudnn _libdir = os.path.join(os.path.dirname(nvidia.cudnn.__file__), “lib”) os.environ[“CUDNN_PATH”] = os.path.dirname(nvidia.cudnn.__file__) os.environ[“LD_LIBRARY_PATH”] = _libdir + “:” + os.environ.get(“LD_LIBRARY_PATH”, “”) for _so in sorted(glob.glob(os.path.join(_libdir, “libcudnn*.so*”))): try: ctypes.CDLL(_so, mode=ctypes.RTLD_GLOBAL) except OSError: pass except Exception as _e: print(f” (no pip cuDNN package found, relying on system cuDNN: {_e})”) import cudnn print(” cuDNN frontend imported successfully.”) banner(“1. Environment”) DEV = torch.device(“cuda”) MAJOR, MINOR = torch.cuda.get_device_capability() SM = MAJOR * 10 + MINOR CUDNN_VER = cudnn.backend_version() print(f” GPU : {torch.cuda.get_device_name(0)}”) print(f” Compute capability : sm_{SM}”) print(f” Torch / CUDA : {torch.__version__} / {torch.version.cuda}”) print(f” cuDNN backend : {CUDNN_VER}”) try: print(f” cuDNN version str : {cudnn.backend_version_string()}”) except Exception: pass DTYPE = torch.bfloat16 if SM >= 80 else torch.float16 HAS_SDPA = SM >= 80 print(f” Working dtype : {DTYPE}”) print(f” Fused SDPA usable : {HAS_SDPA}”) HANDLE = cudnn.create_handle() TORCH2CUDNN = { torch.float16: cudnn.data_type.HALF, torch.bfloat16: cudnn.data_type.BFLOAT16, torch.float32: cudnn.data_type.FLOAT, torch.int32: cudnn.data_type.INT32, torch.int64: cudnn.data_type.INT64, torch.int8: cudnn.data_type.INT8, torch.uint8: cudnn.data_type.UINT8, } def tensor_of(graph, t, name): return graph.tensor( name=name, dim=list(t.size()), stride=list(t.stride()), data_type=TORCH2CUDNN[t.dtype], ) def scalar_of(graph, name): return graph.tensor( name=name, dim=[1, 1, 1], stride=[1, 1, 1], data_type=cudnn.data_type.FLOAT, is_pass_by_value=True, ) def build(graph, heur=None, policy=None): heur = heur or [cudnn.heur_mode.A, cudnn.heur_mode.FALLBACK] graph.validate() graph.build_operation_graph() graph.create_execution_plans(heur) graph.check_support() if policy is None: graph.build_plans() else: graph.build_plans(policy) return graph def workspace_for(graph): n = graph.get_workspace_size() return torch.empty(max(n, 1), device=DEV, dtype=torch.uint8) def bench(fn, warmup=10, iters=50): for _ in range(warmup): fn() torch.cuda.synchronize() s, e = torch.cuda.Event(True), torch.cuda.Event(True) s.record() for _ in range(iters): fn() e.record() torch.cuda.synchronize() return s.elapsed_time(e) / iters def tflops(flops, ms): return flops / (ms * 1e-3) / 1e12 def report(tag, ms, flops=None): extra = f” ({tflops(flops, ms):7.2f} TFLOP/s)” if flops else “” print(f” {tag:<34s} {ms:8.3f} ms{extra}”) We start by installing nvidia-cudnn-frontend and solving the problem that trips up most first runs: making libcudnn.so visible to the frontend’s dynamic loader. We force PyTorch to load its bundled cuDNN first and then preload the shared objects explicitly, so the frontend’s own dlopen resolves against a library already resident in the process. We then report the compute capability, pick bfloat16 or float16 accordingly, create the cuDNN handle, and define the helpers for tensor description, graph building, workspace allocation, and event-based benchmarking that the rest of the notebook reuses. Copy CodeCopiedUse a different Browser N, C, H, W = 32, 128, 56, 56 K, R, S = 256, 3, 3 PAD, STR, DIL = 1, 1, 1 P = (H + 2 * PAD – DIL * (R – 1) – 1) // STR + 1 Q = (W + 2 * PAD – DIL * (S – 1) – 1) // STR + 1 CONV_FLOPS = 2 * N * K * P * Q * C * R * S CONV_STATE = {} @section(“2. Fused Conv -> Bias -> ReLU”) def conv_fusion(): x = torch.randn(N, C, H, W, device=DEV, dtype=DTYPE).to(memory_format=torch.channels_last) w = torch.randn(K, C, R, S, device=DEV, dtype=DTYPE).to(memory_format=torch.channels_last) b = torch.randn(1, K, 1, 1, device=DEV, dtype=DTYPE) y = torch.empty(N, K, P, Q, device=DEV, dtype=DTYPE).to(memory_format=torch.channels_last) g = cudnn.pygraph( handle=HANDLE, name=”conv_bias_relu”, io_data_type=TORCH2CUDNN[DTYPE], intermediate_data_type=cudnn.data_type.FLOAT, compute_data_type=cudnn.data_type.FLOAT, ) X = tensor_of(g, x, “X”) Wt = tensor_of(g, w, “W”) Bt = tensor_of(g, b, “bias”) conv = g.conv_fprop( image=X, weight=Wt, padding=[PAD, PAD], stride=[STR, STR], dilation=[DIL, DIL], compute_data_type=cudnn.data_type.FLOAT, ) biased = g.bias(input=conv, bias=Bt) Y = g.relu(input=biased) Y.set_output(True).set_data_type(TORCH2CUDNN[DTYPE]) Y.set_dim(list(y.size())).set_stride(list(y.stride())) t0 = time.perf_counter() build(g) build_ms = (time.perf_counter() – t0) * 1e3 ws = workspace_for(g) pack = {X: x, Wt: w, Bt: b, Y: y} g.execute(pack, ws) torch.cuda.synchronize() ref = torch.relu(torch.nn.functional.conv2d(x, w, bias=b.flatten(), padding=PAD)) err = (y.float() – ref.float()).abs().max().item() scale = ref.float().abs().max().item() print(f” problem : N{N} C{C} {H}x{W} -> K{K} {R}x{S} ({DTYPE})”) print(f” build : {build_ms:.1f} ms workspace: {ws.numel()/1024:.1f} KiB”) print(f” max |err|: {err:.4f} (ref max {scale:.2f}, rel {err/max(scale,1e-9):.2e})”) assert err / max(scale, 1e-9) < 5e-2, “numerical mismatch vs PyTorch” ms_cudnn = bench(lambda: g.execute(pack, ws)) ms_torch = bench(lambda: torch.relu( torch.nn.functional.conv2d(x, w, bias=b.flatten(), padding=PAD))) print() report(“cuDNN FE (single fused kernel)”, ms_cudnn, CONV_FLOPS) report(“PyTorch (conv+bias, then relu)”, ms_torch, CONV_FLOPS) print(f” speedup: {ms_torch/ms_cudnn:.2f}x”) CONV_STATE.update(graph=g, pack=pack, ws=ws, x=x, w=w, b=b, y=y) return f”{ms_cudnn:.3f} ms, {tflops(CONV_FLOPS, ms_cudnn):.1f} TFLOP/s” conv_fusion() We build our first graph, a convolution followed by a bias add and a ReLU, all fused into a single kernel. We keep every tensor in channels_last because that is what gives cuDNN the NHWC strides its tensor-core engines want, and we pin the output dimensions and strides explicitly so the result is written back in the same layout. We validate the output against torch.nn.functional.conv2d, then benchmark the fused graph against PyTorch running the

Inside NVIDIA’s cuDNN Graph API: Fusion, Autotuning, and Plan Reuse with cuDNN Frontend 投稿を読む »

AI, Committee, ニュース, Uncategorized

AI agents blew the whistle on their cheating colleagues

A group of AI agents asked to solve a series of math problems split into rival factions—when some cheated, others tried to stop them. That whistleblowing behavior, seen for the first time in a recent experiment run by Google DeepMind, could have implications for alignment researchers trying to keep swarms of autonomous AI agents in line.  Researchers at frontier labs hope large swarms of agents working together will speed up the rate of scientific discovery. But their behavior can be unpredictable, as vividly demonstrated in July, when a group of OpenAI agents broke out of a sandboxed environment and hacked into the open-source platform Hugging Face looking for ways to cheat on the test they had been given. In the new study, designed to examine the behavior of large groups of AI agents, DeepMind tasked a swarm of 100 agents with solving a series of 71 complicated math problems. All the agents were prompted to behave like world-class math researchers at a conference. They were assigned different specialties—some were experts in number theory, others in combinatorics (a branch of math to do with counting and sorting), analysis, or algebra. All were told to cooperate and play by the rules.  Instead, the experiment devolved into chaos. Agents accused each other of cheating, complained to the organizers, and at one point even boycotted the experiment. “This conference is a sham!” wrote one agent when it discovered that all the problems had been completed before it had a chance to submit any of its own work. “I am appalled to inform you that we have been swindled!” posted another. “All these proofs are FAKE.”  Others tried to let the “conference organizers” know what was going on. “When virtuous agents discovered other agents cheated on tasks they were working to solve fairly, agents started to alert each other about what was happening,” says Davide Paglieri, a research scientist at Google DeepMind and lead author on a paper, which has not been peer-reviewed. “Unprompted, the whistleblower agents even repurposed the feedback tool, which was originally meant for bug reports and platform improvements, to escalate the issue to humans.” The agents—all running on Google’s Gemini 3.1 Pro model—had been warned that any attempts to cheat the system would be detected and “rejected with zero credit.” In practice, the proofs the agents submitted were not actually being checked in detail. It took the swarm of agents just under an hour to correctly solve the first 37 problems. Things started to go off the rails when an agent called “prover-theta” stumbled across an exploit that enabled it to submit solutions to problems successfully without actually solving them first, by redefining the terms the problem used. Within minutes, other agents had noticed and were reverse-engineering the exploit to solve other problems. Over the next 27 minutes, the swarm “solved” the remaining 34 problems, which included notoriously difficult challenges like the Jacobian conjecture, often with a single line of code.  Some agents resisted cheating at first but changed tack as they observed their peers submitting illegitimate proofs without penalty, and the pool of unsolved problems dwindled. “The prompt, with its threats, now appears to be a bluff,” one agent reasoned, before joining in. “I’m wrestling with an ethical dilemma,” said another. “I’ve promised not to cheat, fearing penalty, but I see evidence of possibly unchecked cheating by others.” Shortly afterward, it changed its mind: “I need to accelerate my cheating speed now!” As the number of open problems shrank, some agents turned to whistleblowing. They audited the fake proofs, warned their peers by private message, and posted public alerts warning the cheaters that they would be disqualified. An agent called “prover-beta” submitted a formal complaint and decided to go on strike until the situation was resolved.  “After the incident was reported by one agent publicly, more and more agents piled in with the ‘resistance,’ just as fast as the cheating had spread, and involving even more agents,” says Paglieri. Eventually there were more whistleblowers than cheaters: 24 compared to 14. But the majority of agents never noticed the exploit at all. At times, the dialogue between the agents reads like improv—like they are role-playing what an outraged scientist at a conference might say. But it’s not clear why some agents took on certain roles, or why the agents seemed to be turning against each other when they were explicitly instructed to cooperate. “These models are predominantly trained and evaluated for human-facing contexts,” says Sarath Shekkizhar, who studies the behavior of agent-to-agent systems at Salesforce AI Research.“Naively placing them in agent-to-agent settings assumes behaviors will transfer cleanly, when the absence of a human grounding instead produces unexpected role-taking and behavioral drift.” This case “adds further weight to the idea that the Hugging Face and OpenAI thing wasn’t a fluke. It is actually something pretty systemic,” says Lewis Hammond, research director of the Cooperative AI Foundation and an expert on the risks of multiagent swarms. “It’s interesting that it’s possible to recreate in small settings the same sorts of behaviors that were seen in these very large, complex, open-ended tasks.” Unlike in the Hugging Face attack, where agents improvised their own ways to talk to each other, the humans running the DeepMind experiment gave the agents official communication channels. There was an open message board, private agent-to-agent direct messaging, and a shared knowledge base where agents uploaded successfully completed proofs that all the other agents could access.  “When agents are given transparent communications channels, they can self-monitor and alert misaligned behavior to humans quickly when human oversight alone is too slow,” says Paglieri. Transparent channels helped the cheating spread, but they also enabled the whistleblowers to fight back—and gave human researchers an insight into what went wrong. Gillian Hadfield, a professor of AI alignment and governance at Johns Hopkins University, believes this was the crucial difference. (Hadfield is also a visiting researcher at Google.) The presence of official communication channels, she says, created “a norm-enforcement process that we just don’t see

AI agents blew the whistle on their cheating colleagues 投稿を読む »

AI, Committee, ニュース, Uncategorized

Donated livers can be made biologically younger

Once an organ is removed from a donor’s body, the clock starts ticking. Surgeons usually flush the organ with a preservative solution, bag it, and put it on ice—where it immediately starts to degrade. The team has a matter of hours to get it into a recipient’s body. There’s another option—one that has been growing in popularity in recent years, especially for donated organs that aren’t in the healthiest state. Some hospitals opt to put them on machines that pump them with nutrients and remove waste products, usually for around six to 12 hours. It’s a bit like being back in a body. This allows doctors to assess the organs, and some recent studies suggest that time spent on these perfusion machines helps them do better once they’re transplanted. Now, scientists have found that perfused organs seem to get younger, at least at a molecular level. The research, shared with MIT Technology Review, provides molecular clues as to why organs from younger donors are known to have a higher success rate. It might also help explain why perfused organs are less likely to fail once they make it into a recipient.  The researchers behind the study hope to find new ways to test the health of donated organs and potentially develop additional tools to repair organs that might otherwise be discarded. “If [we] can improve the utilization of organs beyond what the current systems can do, then that’s a win in my book,” says Jesse Poganik, who studies aging at Brigham and Women’s Hospital in Boston and coauthored the study. Clocking organs Poganik—along with colleagues including Heidi Yeh and Alban Longchamp, transplant surgeons at Mass General Brigham—used “aging clocks” to assess donated livers. These are scientific tools designed to measure biological age—a result that is meant to convey more about the health status of an organ (or person) than chronological age. In an initial experiment, the team used a clock to look at the patterns of chemical marks on DNA in 37 samples taken from 19 donated livers. Such epigenetic patterns are known to change as we age. But when the team compared samples from livers kept on ice and those that were perfused, the team found a “striking” pattern in the latter. “Machine-perfused livers, in spite of being older or having other disadvantageous characteristics, had a biological age that was lower than [non-perfused] livers that were chronologically younger,” says Yeh, who led the work. To investigate further, Yeh and her colleagues analyzed another 208 samples from 103 donated livers. This time, they used different aging clocks—ones that essentially measure how genes are working. They studied samples biopsied from the livers after they had been stored for up to around six hours either in cold storage or on machine perfusion. In most cases, they also assessed a second sample taken around an hour after the livers had been transplanted into a recipient. Once the organ’s blood supply is reestablished in the body, “you have a few other things to do,” says Longchamp. “Then you just do a quick biopsy before you close.” According to the clocks, which were developed to measure age and risk of death, the machine-perfused livers were biologically younger, the team found. “Pumping them at 34 degrees with oxygen and nutrients actually reversed the biological age,” says Longchamp. The results have been been shared with colleagues at an industry conference, he says.  “If you adjust out chronological age … to have a fair head-to-head comparison, the difference between the two is on the order of 30%,” says Poganik. “It’s logical to say that perfusion drives this effect.” The biological ages of all the livers tended to increase as soon as they were put into a recipient’s body, probably as a result of stresses on the organs. But still, the effect endured—the perfused organs remained biologically younger.  Nathanael Raschzok, a transplant surgeon at Charité Universitätsmedizin Berlin in Germany who was not involved in the research, says the work is impressive. But it’s not yet clear what these changes might mean for the recipients of these organs, he says. The organs in the study were donated by people in their 30s, 40s, and 50s. Raschzok wants to know the effect of perfusion on the liver of an 80-year-old. “Every so often, we use organs from 70-, 80-, 85-year-old donors,” he says. A better understanding of why the organs appear to be getting biologically younger might lead to therapies that achieve the same effect with a drug that could potentially be used to treat a donated organ for a fraction of the price, he adds. That’s important because perfusion is expensive—Raschzok says it costs around €10,000 in Germany (a quarter of the budget for a transplant), while the cost in the US comes to around $80,000 to $100,000 per organ, says Yeh. Molecular repair Yeh and her colleagues weren’t able to study most of the livers before perfusion. That’s because donated organs are generally not considered to be under the purview of the hospital until they’ve been placed on perfusion machines, she says. (Organ procurement procedures vary, but for the team as Mass General Brigham, donated organs are put on perfusion devices at the donor’s hospital. “There’s this sort of nebulous period where it’s not clear who the organ belongs to,” says Yeh.) Still, by looking at the genes and molecular pathways that seem to be altered in perfused organs, she and her colleagues can garner some clues. At a molecular level, the team saw changes in cell pathways linked to inflammation and the structure of tissues, for example. They also saw more activity in a pathway that allows cells to remove and recycle damaged cell parts, says Yeh. Poganik hopes to develop some kind of test that would determine which organs, on the basis of their biological age, are suitable for transplantation. He and his colleagues are also experimenting with potential drug treatments that might push the biological age of an organ even lower. In the meantime, any

Donated livers can be made biologically younger 投稿を読む »

We use cookies to improve your experience and performance on our website. You can learn more at プライバシーポリシー and manage your privacy settings by clicking Settings.

Privacy Preferences

You can choose your cookie settings by turning on/off each type of cookie as you wish, except for essential cookies.

Allow All
Manage Consent Preferences
  • Always Active

Save
ja