YouZum

Uncategorized

AI, Committee, ニュース, Uncategorized

The Download: a new Christian phone network, and debugging LLMs

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. A new US phone network for Christians aims to block porn and gender-related content A new US-wide cell phone network marketed to Christians is set to launch next week. It blocks porn using network-level controls that can’t be turned off—even by adult account owners. It’s also rolling out a filter on sexual content aimed at blocking material related to gender and trans issues, optional but turned on by default across all plans. The trouble is, many websites don’t fit neatly into one category. That leaves its maverick founder with broad, subjective control over what is allowed or banned. Read the full story. —James O’Donnell This startup’s new mechanistic interpretability tool lets you debug LLMs The San Francisco–based startup Goodfire has released a new tool, Silico, that lets researchers peer inside an AI model and adjust its parameters during training. It could give users more control over how this technology is built than was once thought possible. The goal is to make building AI models less like alchemy and more like a science. Using a technique called mechanistic interpretability, Silico maps the neurons and pathways inside a model and lets developers tweak them to reduce unwanted behaviors or steer outputs. By exposing the “knobs and dials,” Goodfire hopes to bring AI training closer to traditional software engineering. Read the full story. —Will Douglas Heaven With mass firing, Trump deals a fresh blow to American science This past week delivered another gut punch for science in the US. This time, the target was the National Science Foundation—a federal agency that funds major research projects to the tune of around $9 billion. On Friday, the 22 scientists overseeing those efforts were all fired. Since 2025, the NSF has faced budget cuts, grant terminations, and mass firings, with staff numbers down sharply and many ambitious projects grinding to a halt. The result is a major shift in how American science is funded and governed. Discover what it means, and what’s next. —Jessica Hamzelou This article first appeared in The Checkup, MIT Technology Review’s weekly biotech newsletter. To receive it in your inbox every Thursday, and read articles like this first, sign up here. China’s open-source bet: 10 Things That Matter in AI Right Now Silicon Valley AI companies follow a familiar playbook: keep the models behind an API and charge for access. China’s leading AI labs are playing a different game, releasing “open-weight” models that developers can download, adapt, and run on their own hardware. That approach went mainstream after DeepSeek open-sourced its R1 model, which matched top US systems at a fraction of the cost. It also won something subtler: goodwill with developers. A growing cohort of Chinese labs is now following the same blueprint. As AI shifts from hype to deployment, open-source models are making the future of AI more multipolar than Silicon Valley expected. Read the full story. —Caiwei Chen China’s open-source bet is one of the 10 Things That Matter in AI Right Now, our list of the biggest ideas, trends, and advances in AI today. We’re unpacking one item from the list each day here in The Download, so stay tuned. The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 Elon Musk has admitted that xAI trained Grok on OpenAI models“Distillation” is standard practice in AI, despite being legally dubious. (Wired $)+ The White House has accused Chinese firms using distillation of theft. (BBC)+ American labs are widely assumed to use similar techniques. (TechCrunch) 2 A “de-extinction” startup wants to resurrect a long-lost antelopeColossal Biosciences wants to bring back the bluebuck. (Axios)+ The company is using genomic editing to revive the animal. (Gizmodo)+ It previously claimed to have cloned red wolves. (MIT Technology Review) 3 ​​An OpenAI model outperformed ER doctors at diagnosing patientsBy analyzing health records data and information provided to physicians. (NPR)+ But it still must be proven in real-world clinical trials. (Vox) 4 Scientists are trying to power AI data centers with tiny nuclear reactorsThey could provide a new way to meet AI’s energy demands. (Gizmodo)+ We did the math on AI’s energy footprint. (MIT Technology Review) 5 Spotify has started verifying human artistsA new badge will distinguish them from AI. (The Guardian)+ Spotify has faced criticism for its handling of AI. (BBC) 6 The US is backing a Congolese railway to break China’s grip on critical mineralsThe old railroad is key to the race for critical metals in Africa. (Rest of World)+ The US is also searching for alternative sources. (MIT Technology Review) 7 Huawei is set to overtake Nvidia in China’s AI chip marketIt’s expected to capture the largest market share this year. (FT $) 8 Japan is building cardboard drones for the battlefieldThe flatpack designs are cheap, disposable, and built at scale. (404 Media) 9 The more young people use AI, the more they hate itResearch shows that Gen Z doesn’t trust GenAI. (The Verge) 10 A new organoid can menstruate—and show how tissue repairs itselfIt’s revealing how the uterus can shed without scarring. (Nature) Quote of the day “I suspect that there are a number of people who do not want to put the future of humanity in Mr Musk’s hands. But we’re not going to get into that.” —Judge Gonzalez Rogers rebukes attempts by Elon Musk’s lawyer to focus on AI’s existential risks as part of his lawsuit against OpenAI, the New York Times reports.  One More Thing TMY350 VIA WIKIMEDIA COMMONS This rare earth metal shows us the future of our planet’s resources The materials we need to power our world are shifting from fossil fuels to energy sources that don’t produce greenhouse gas emissions. Take neodymium, a rare earth metal used in powerful magnets that power everything from smartphones to wind turbines. Its story reveals many of the challenges we’ll likely face across the supply chain in

The Download: a new Christian phone network, and debugging LLMs 投稿を読む »

AI, Committee, ニュース, Uncategorized

Operationalizing AI for Scale and Sovereignty

Companies are taking control of their own data to tailor AI for their needs. The challenge lies in balancing ownership with the safe, trusted flow of high‑quality data needed to power reliable insights. This conversation from MIT Technology Review’s EmTech AI conference examines how AI factories unlock new levels of scale, sustainability, and governance—positioning data control as a strategic imperative for governments and enterprises. About the speakers Chris Davidson, Vice President, HPC & AI Customer Solutions, HPE Chris Davidson is Vice President of HPC & AI Customer Solutions at Hewlett Packard Enterprise. He leads HPE’s global strategy for AI Factory solutions and Sovereign AI, working with governments, enterprises, and research institutions to build secure, scalable national- and enterprise-grade AI capabilities. He also directs Product Management and Performance Engineering across HPE’s HPC and AI portfolio, including large-model training platforms and Cray exascale systems. His teams define product strategy, performance architecture, and deployment models that position HPE at the forefront of high-performance and AI computing. During his nine years at HPE, Chris has led key initiatives across Performance Engineering, AI Cloud, and Professional Services, shaping how HPE delivers optimized, cloud-native, and globally deployed high-performance systems. He previously held technical and leadership roles in the biotech and medical diagnostics sectors. Chris holds an M.B.A. in Entrepreneurship and Finance and a B.S. in Biology from Loyola University Chicago. Arjun Shankar, Division Director, National Center for Computational Science, Oak Ridge National Laboratory Mallikarjun (Arjun) Shankar is the Division Director for the National Center for Computational Science at the Oak Ridge National Laboratory. His research focuses on the interdisciplinary bridge between computer science and large-scale scientific discovery campaigns that rely on scalable computing and data science. He is a joint faculty appointee at the University of Tennessee’s Bredesen Center, a senior member of the IEEE and a senior member of the ACM.

Operationalizing AI for Scale and Sovereignty 投稿を読む »

AI, Committee, ニュース, Uncategorized

Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval Augmented Generation

arXiv:2505.21072v5 Announce Type: replace Abstract: Large Language Models (LLMs) enhanced with retrieval, an approach known as Retrieval-Augmented Generation (RAG), have achieved strong performance in open-domain question answering. However, RAG remains prone to hallucinations: factually incorrect outputs may arise from inaccuracies in the model’s internal knowledge and the retrieved context. Existing approaches to mitigating hallucinations often conflate factuality with faithfulness to the retrieved evidence, incorrectly labeling factually correct statements as hallucinations if they are not explicitly supported by the retrieval. In this paper, we introduce FRANQ, a new method for hallucination detection in RAG outputs. FRANQ applies distinct uncertainty quantification (UQ) techniques to estimate factuality, conditioning on whether a statement is faithful to the retrieved context. To evaluate FRANQ and competing UQ methods, we construct a new long-form question answering dataset annotated for both factuality and faithfulness, combining automated labeling with manual validation of challenging cases. Extensive experiments across multiple datasets, tasks, and LLMs show that FRANQ achieves more accurate detection of factual errors in RAG-generated responses compared to existing approaches.

Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval Augmented Generation 投稿を読む »

AI, Committee, ニュース, Uncategorized

IBM Releases Two Granite Speech 4.1 2B Models: Autoregressive ASR with Translation and Non-Autoregressive Editing for Fast Inference

IBM released two new open speech recognition models— Granite Speech 4.1 2B and Granite Speech 4.1 2B-NAR — and they make a compelling case for what a ~2B-parameter speech model can do. Both are available on Hugging Face under the Apache 2.0 license. The pair targets a specific problem that enterprise AI teams know well: most production-grade automatic speech recognition (ASR) systems either demand massive compute or sacrifice accuracy to stay within budget. IBM’s bet is that careful architecture decisions can let you have it both ways. What These Models Actually Do Granite Speech 4.1 2B is a compact and efficient speech-language model designed for multilingual automatic speech recognition (ASR) and bidirectional automatic speech translation (AST) covering English, French, German, Spanish, Portuguese, and Japanese. Its non-autoregressive counterpart, Granite Speech 4.1 2B-NAR, focuses exclusively on ASR — specifically targeting latency-sensitive deployments — and supports English, French, German, Spanish, and Portuguese, but not Japanese. That’s a meaningful distinction: teams that need Japanese transcription or any speech translation capability should reach for the standard autoregressive model. IBM also quietly released a third variant alongside these two. Granite Speech 4.1 2B-Plus adds speaker-attributed ASR and word-level timestamps for applications where knowing who said what — and exactly when — is a requirement. Word Error Rate (WER) is the primary metric for measuring transcription quality. Lower is better. A WER of 5% means roughly 5 out of every 100 words are wrong. On the Open ASR Leaderboard (as of April 2026), Granite Speech 4.1 2B scores a mean WER of 5.33. Drilling into benchmark detail — on LibriSpeech clean, the model achieves a WER of 1.33, and 2.5 on LibriSpeech other. The Architecture, Explained Both models share the same three-component design at a high level — a speech encoder, a modality adapter, and a language model — though the decoding mechanism diverges significantly. The first component is the speech encoder. The architecture uses 16 conformer blocks trained with Connectionist Temporal Classification (CTC) with two classification heads — one for graphemic (character-level) outputs and one for BPE units — using frame importance sampling to focus on informative parts of the audio. A Conformer is a neural network layer that combines convolutional layers (good at capturing local acoustic patterns) with attention mechanisms (good at capturing long-range dependencies). CTC is a training technique that lets the model learn from audio-text pairs without needing exact frame-level alignment. The second component is a speech-text modality adapter. A 2-layer window query transformer (Q-Former) operates on blocks of 15 1024-dimensional acoustic embeddings coming from the last conformer block, downsampling by a factor of 5 using 3 trainable queries per block and per layer — for a total temporal downsampling factor of 10 — resulting in a 10Hz acoustic embedding rate for the LLM. This adapter bridges the gap between continuous acoustic features and discrete text tokens, compressing the audio representation so the language model can process it efficiently. In the NAR model, the Q-Former has 160M parameters and downsamples the concatenated hidden representations from four encoder layers (layers 4, 8, 12, and 16). The third component is the language model. Granite Speech 4.1 2B uses an intermediate checkpoint of granite-4.0-1b-base with 128k context length, fine-tuned on all training corpora. In the NAR variant, this becomes a 1B-parameter bidirectional LLM editor — granite-4.0-1b-base with its causal attention mask removed to enable bidirectional context — adapted with LoRA at rank 128 applied to both attention and MLP layers. The Autoregressive vs. Non-Autoregressive Tradeoff This is where the two models diverge most sharply, and it has direct consequences for production deployment. In the standard Granite Speech 4.1 2B, text is generated autoregressively — one token at a time, each depending on every token before it. This produces accurate, stable transcripts with full support for AST, keyword-biased recognition, and punctuation, but is inherently sequential and slower at scale. Granite Speech 4.1 2B-NAR takes a fundamentally different approach. Rather than decoding tokens one at a time, it edits a CTC hypothesis in a single forward pass using a bidirectional LLM, achieving competitive accuracy with faster inference than autoregressive alternatives. This is the NLE (Non-autoregressive LLM-based Editing) architecture. Concretely: the CTC encoder produces a rough initial transcript, that hypothesis is interleaved with insertion slots, and then a bidirectional LLM predicts edits — copy, insert, delete, or replace — at all positions simultaneously in one pass. The NAR model measured an RTFx of approximately 1820 on a single H100 GPU using batched inference at batch size 128. RTFx (real-time factor multiplier) measures how many times faster than real time a model can process audio — an RTFx of 1820 means a one-hour audio file can be transcribed in under two seconds on that hardware. One practical constraint engineers should note: the NAR model requires flash_attention_2 for inference, since this backend supports sequence packing and respects the is_causal=False flag. Training Data and Infrastructure The two models were trained on different datasets. The standard model was trained on 174,000 hours of audio from public corpora for ASR and AST, as well as synthetic datasets tailored to support Japanese ASR, keyword-biased ASR, and speech translation. The NAR model was trained on approximately 130,000 hours of speech across five languages using publicly available datasets including CommonVoice 15, MLS, LibriSpeech, LibriHeavy, AMI, Granary VoxPopuli, Granary YODAS, Earnings-22, Fisher, CallHome, and SwitchBoard. The infrastructure gap between the two is equally telling. The standard model’s training was completed in 30 days — 26 days for the encoder and 4 days for the projector — on 8 H100 GPUs. The NAR model trained in just 3 days on 16 H100 GPUs (2 nodes) for 5 epochs — a much lighter training run, which reflects the architectural simplicity of editing over full autoregressive generation. Key Takeaways Here are 5 short key takeaways: IBM released two open ASR models — Granite Speech 4.1 2B (autoregressive) and Granite Speech 4.1 2B-NAR (non-autoregressive) — both ~2B parameters, and Apache 2.0 licensed. The standard model achieves a mean WER of 5.33

IBM Releases Two Granite Speech 4.1 2B Models: Autoregressive ASR with Translation and Non-Autoregressive Editing for Fast Inference 投稿を読む »

AI, Committee, ニュース, Uncategorized

Cursor Introduces a TypeScript SDK for Building Programmatic Coding Agents With Sandboxed Cloud VMs, Subagents, Hooks, and Token-Based Pricing

Cursor, the AI-powered code editor, is opening up the core technology behind its coding agents to developers everywhere. The Cursor team announced the public beta of the Cursor SDK — a TypeScript library that gives engineers programmatic access to the same runtime, harness, and models that power Cursor’s desktop app, CLI, and web interface. This signals a meaningful shift in how AI coding tools are being positioned: not just as interactive assistants sitting alongside a developer, but as deployable infrastructure that organizations can wire into their existing systems. From Interactive Tool to Programmable Infrastructure If you’ve used Cursor before, you know it as an IDE where you interact with an agent in real time — asking it to write functions, fix bugs, or explain code. The Cursor SDK changes the access model. Instead of a developer sitting at a keyboard, the agent can now be invoked programmatically: from a CI/CD pipeline trigger, a backend service, or embedded directly inside another product. Think of it this way: previously, you had to be “in” Cursor to use its agents. Now, you can call those same agents from anywhere in your stack with a few lines of TypeScript. Getting started is a single command: Copy CodeCopiedUse a different Browser npm install @cursor/sdk From there, you create an Agent instance, send it a task, and stream the response back — all in TypeScript. Here’s the minimal example from Cursor’s announcement: Copy CodeCopiedUse a different Browser import { Agent } from “@cursor/sdk”; const agent = await Agent.create({ apiKey: process.env.CURSOR_API_KEY!, model: { id: “composer-2” }, local: { cwd: process.cwd() }, }); const run = await agent.send(“Summarize what this repository does”); for await (const event of run.stream()) { console.log(event); } The Agent.create() call accepts an apiKey, a model field (where you specify which model to run), and either a local or cloud configuration depending on where you want execution to happen. Why Building Your Own Agent Stack is Hard Before diving into what the SDK offers, it’s worth understanding the problem it solves. Building fast, reliable, and capable coding agents that run safely against your data requires meaningful engineering effort: secure sandboxing, durable state and session management, environment setup, and context management. And when a new model ships, dev teams often have to rework their agent loops entirely just to take advantage of it. The Cursor SDK eliminates this complexity so teams can focus on building useful agents instead of maintaining the underlying infrastructure. https://cursor.com/blog/typescript-sdk The Agent Harness: What “Same Runtime” Actually Means SDK agents use the same harness that powers Cursor’s own products. ‘Harness’ here refers to the full set of supporting infrastructure that makes an agent effective beyond just the LLM call itself. In Cursor’s case, that includes: Intelligent context management — Codebase indexing, semantic search, and instant grep so agents retrieve the right code context before generating responses. This is critical because LLMs are only as good as the context they receive; poor retrieval leads to hallucinated or irrelevant outputs. MCP servers — Agents launched through the SDK can connect to external tools and data sources over stdio or HTTP, either via a .cursor/mcp.json config file or passed inline in the API call. MCP (Model Context Protocol) is an open standard for wiring tools into agent runtimes. Skills — Agents automatically pick up reusable behavior definitions from a .cursor/skills/ directory in the repository. Hooks — A .cursor/hooks.json file lets you observe, control, and extend the agent loop across cloud, self-hosted, and local runtimes — useful for logging, guardrails, or custom orchestration. Subagents — The main agent can delegate subtasks to named subagents with their own prompts and models via the Agent tool, enabling multi-agent workflows without custom orchestration code. Cloud Deployment: Persistent, Sandboxed, and Resumable One of the more practical features of the SDK is cloud execution. When configured to run in Cursor’s cloud, each agent gets its own dedicated VM with strong sandboxing, a clone of the target repository, and a fully configured development environment. Critically, the agent keeps running even if the initiating machine goes offline — the developer can reconnect and stream the conversation later. Cloud agents integrate with Cursor’s existing Agents Window and web app, so a task started programmatically via the SDK can be inspected or taken over manually inside the Cursor interface. When the agent finishes, it can open a PR, push a branch, or attach demos and screenshots — making them suitable for asynchronous, unattended workflows: Copy CodeCopiedUse a different Browser const agent = await Agent.create({ apiKey: process.env.CURSOR_API_KEY!, model: { id: “gpt-5.5” }, cloud: { repos: [{ url: “https://github.com/cursor/cookbook”, startingRef: “main” }], autoCreatePR: true, }, }); const run = await agent.send(“Fix the auth token expiry bug”); console.log(`Started ${run.id}`); // …check back in later, from anywhere: const result = await ( await Agent.getRun(run.id, { runtime: “cloud”, agentId: run.agentId }) ).wait(); console.log(result.git?.branches[0]?.prUrl); For dev teams with security requirements, the SDK also supports self-hosted workers, where both code and tool execution remain inside the organization’s own network. Model Flexibility and Composer 2 The SDK exposes every model supported in Cursor. Switching models is a single field change in the model parameter, letting teams route tasks to the best model for a given combination of cost and capability. Cursor’s own Composer 2 — described as a specialized coding model achieving frontier-level performance at a fraction of the cost of general-purpose models — is positioned as the default recommendation for most coding agent tasks. Getting Started To accelerate adoption, Cursor has published a public cookbook repository on GitHub with four starter projects: a minimal quickstart (a Node.js example that creates a local agent, sends one prompt, and streams the response), a web-based prototyping tool for scaffolding new projects in a sandboxed cloud environment, an agent-powered kanban board that automatically opens PRs when engineers drag a card, and a lightweight coding agent CLI for spawning Cursor agents from the terminal. Cursor has also released a Cursor SDK plugin in the Cursor Marketplace to help developers start building directly from within the editor.

Cursor Introduces a TypeScript SDK for Building Programmatic Coding Agents With Sandboxed Cloud VMs, Subagents, Hooks, and Token-Based Pricing 投稿を読む »

AI, Committee, ニュース, Uncategorized

The Download: the North Pole’s future and humanoid data

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Digging for clues about the North Pole’s past In the past, getting to the North Pole involved a treacherous trip through ice many meters thick. But last year, a research vessel encountered open water and thin ice, which created an easy passage. It provided a reminder of how quickly the Arctic is changing.  Now scientists are digging deep below the seabed to find out if the Arctic Ocean was ever ice-free—and what that could mean for the future of Earth’s northernmost waters. Here’s what they hope to discover. —Tim Kalvelage This story is from the latest issue of our print magazine, which is all about nature. Check out the full issue here, and subscribe to get the next one when it lands.  Humanoid data: 10 Things That Matter in AI Right Now I was recently invited to join an app that would pay me to film myself doing tasks like putting food in a bowl and microwaving it. Another site asked if I’d like to remotely control a robotic arm to help improve its dexterity. What on earth is happening? These examples are just part of a growing push by robotics companies to collect data on our movements for training humanoids. As the race for real-world data heats up, our everyday movements are being turned into training data. Read the full story. —James O’Donnell Humanoid data is one of our 10 Things That Matter in AI Right Now, a new look at the big ideas, trends, and technologies really worth your attention in the buzzy world of AI. The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 Google, Microsoft, Amazon, and Meta have all set AI spending recordsCollectively, they’re up 71% on the same quarter last year.  (NYT $)+ Microsoft, Google and Amazon reported big payoffs from the splurge. (FT $)+ But Meta’s shares slid after its plans spooked investors. (BBC)+ What even is the AI bubble? (MIT Technology Review) 2 The White House opposes Anthropic’s plan to expand Mythos accessIt’s concerned about the model’s cyber risks. (Bloomberg $)+ And worried that the government will lose compute access. (WSJ $)+ Anthropic is seeking funding at a valuation over $900 billion. (Bloomberg $) 3 Elon Musk has claimed OpenAI’s leaders “looted the nonprofit”During testimony, Musk said he “was a fool” for trusting them. (Gizmodo)+ But he had raised his own concerns about OpenAI’s non-profit status. (The Verge)+ The case could reshape the AI landscape. (MIT Technology Review) 4 Autonomous vehicles may be worseningAccording to emergency first-responders, glitches are increasing. (Wired) 5 OpenAI has abandoned much of its Stargate planIt will no longer develop its own data centers. (FT $)+ The project’s compute requirements have been questioned. (MIT Technology Review) 6 A convicted Harvard scientist is rebuilding a brain-computer lab in ChinaHe had previously been named the world’s top chemist. (Reuters $) + But was then convicted for lying about payments from China. (NYT $) 7 Families have sued OpenAI over a mass shooter’s use of ChatGPTThey say OpenAI provided a dangerously defective version of the chatbot. (NPR) 8 Apple is reportedly close to giving up on the Vision ProAfter the latest model flopped. (MacRumors)  9 Senators are interrogating US AI firms on safeguards against ChinaOver fears of IP theft. (Axios) 10 Friendly AI chatbots are more likely to be inaccurateA new study found kinder answers contained more mistakes. (BBC) Quote of the day “Never talk about goblins, gremlins, raccoons, trolls, ogres, pigeons, or other animals or creatures unless it is absolutely and unambiguously relevant to the user’s query.”  —OpenAI instructs Codex to avoid critter talk in a system prompt for the coding agent, Ars Technica reports. One More Thing ARTHUR MOUNT Is this the most energy-efficient way to build homes? When engineers began designing an ultra-efficient home in the 1970s, they realized the trick wasn’t generating energy in a greener way, but using less of it. They needed to make a better thermos, not a cheaper coffee maker. That idea helped inspire today’s passive-house standard: airtight buildings that can cut energy use by up to 90% through better windows, insulation, and ventilation. Although they’re often considered a cold-climate approach, passive houses actually have universal benefits. Find out what makes them so efficient. —Patrick Sisson We can still have nice things A place for comfort, fun and distraction to brighten up your day. (Got any ideas? Drop me a line.) + Finally, someone built a gaming PC inside a microwave that runs DOOM.+ Experience the rhythm of the city through this rapid-fire collage of urban photography.+ Get a dose of pure cuteness as these tiny snow leopard cubs leave their den for the first time.+ If you’re staring at a random assortment of groceries, SuperCook will find a recipe based on what’s already in your pantry.

The Download: the North Pole’s future and humanoid data 投稿を読む »

AI, Committee, ニュース, Uncategorized

smol-audio: A Colab-Friendly Notebook Collection for Fine-Tuning Whisper, Parakeet, Voxtral, Granite Speech, and Audio Flamingo 3

Audio AI has had a breakout year. Automatic speech recognition has gotten dramatically better with models like OpenAI’s Whisper variants, NVIDIA’s Parakeet, and Mistral’s Voxtral. Audio understanding stepped forward with models like NVIDIA’s Audio Flamingo 3. Dialogue-grade text-to-speech arrived via Nari Labs’ Dia-1.6B. And Meta shipped the Perception Encoder Audiovisual (PE-AV), a multimodal encoder capable of learning a shared embedding space across audio, video, and text. The frontier has never moved faster. The catch? The practical knowledge required to actually work with these models — how to fine-tune them, adapt them to new languages, or run efficient inference — is scattered across GitHub issues, research blogs, and private notebooks that never see the light of day. If you are an ML engineer who just wants to fine-tune Whisper on a new domain or run zero-shot video classification with PE-AV, you are often starting from scratch. That is the gap smol-audio is designed to close. What is smol-audio ? Released under the Apache-2.0 license by the Deep-unlearning team, smol-audio is a flat repository of self-contained Jupyter notebooks, each focused on a single practical audio AI task. Every notebook is designed to be opened directly in Google Colab, requires no local GPU setup, and is built entirely on the Hugging Face ecosystem — specifically transformers, datasets, peft, and accelerate. Most recipes fit within a 16 GB Colab runtime, which means a free or standard Colab tier is sufficient for the majority of tasks. The “flat repo” design is a deliberate choice. Rather than wrapping recipes inside a framework or hiding complexity behind convenience functions, smol-audio exposes every step. You can read the training loop, understand the data pipeline, and modify the configuration without reverse-engineering a library. For early-career engineers, that transparency is genuinely educational. ASR Fine-Tuning: Whisper, Parakeet, Voxtral, and Granite Speech The largest category in the repo today covers ASR fine-tuning across four distinct model families. Each requires meaningfully different handling. The Whisper notebook covers fine-tuning using transformers and datasets, making it straightforward to adapt the encoder-decoder architecture to a custom language or narrow domain. Whisper uses a sequence-to-sequence approach, generating transcripts token by token — familiar territory for anyone who has worked with language models. NVIDIA’s Parakeet uses a CTC (Connectionist Temporal Classification) architecture rather than a sequence-to-sequence setup. CTC is faster and lighter for inference but requires alignment between audio frames and output tokens rather than autoregressive decoding. The smol-audio notebook covers both full fine-tuning and LoRA (Low-Rank Adaptation) for Parakeet, which is important because full fine-tuning large CTC models can be memory-intensive. Mistral’s Voxtral is architecturally distinct from both Whisper and Parakeet. Rather than a traditional ASR encoder-decoder, Voxtral is built on a large language model backbone — Ministral 3B for Voxtral Mini and Mistral Small 3.1 24B for Voxtral Small — making it an LLM-based speech understanding model. The smol-audio notebook handles fine-tuning for ASR with prompt masking, supporting both full fine-tuning and LoRA. Prompt masking is important here precisely because of this LLM architecture: when a model accepts text prompts alongside audio input, you typically do not want to compute loss on the prompt tokens themselves — only on the generated transcription. Getting this wrong leads to degraded training dynamics, so having a working reference implementation saves significant debugging time. IBM’s Granite Speech gets its own notebook focused on Italian ASR using the YODAS-Granary dataset. This is a useful example beyond just the model: it demonstrates domain- and language-specific fine-tuning on a real multilingual speech corpus, a common production scenario. Audio Understanding with NVIDIA’s Audio Flamingo 3 Audio Flamingo 3, developed by NVIDIA, is a Large Audio Language Model (LALM) for reasoning and understanding across speech, sound, and music. The smol-audio notebook fine-tunes it specifically for the audio captioning task — generating a natural language description of an audio clip, which is useful for accessibility tooling, content indexing, and retrieval systems. The notebook covers both full fine-tuning and LoRA-based fine-tuning, giving practitioners the choice between maximum performance and memory efficiency. LoRA, for those newer to parameter-efficient fine-tuning, works by freezing the original model weights and injecting small trainable rank-decomposition matrices into specific layers. For large multimodal models like Audio Flamingo 3, LoRA can reduce GPU memory requirements by an order of magnitude compared to full fine-tuning, enabling iteration on commodity hardware. Dialogue TTS with Dia-1.6B The Dia-1.6B notebook covers dialogue-style text-to-speech, where the goal is not just synthesizing a single speaker but generating natural conversational exchanges. Dia is a 1.6-billion-parameter TTS model by Nari Labs capable of producing multi-speaker dialogue, making it relevant for anyone building voice agents, podcast generation tools, or conversational interfaces. Multimodal Inference with Meta’s PE-AV Perhaps the most forward-looking notebook in the current release covers inference with Meta’s Perception Encoder Audiovisual (PE-AV). PE-AV is a multimodal encoder that learns a single shared embedding space across audio, video, and text — enabling zero-shot video classification without any task-specific fine-tuning, and audiotext retrieval on benchmarks like AudioCaps. Because all three modalities map into the same embedding space, cross-modal queries such as retrieving an audio clip from a text description work via simple dot-product similarity. The notebook demonstrates how to run these inference pipelines directly, which is valuable because multimodal models with joint audio-visual-text encoders are architecturally more complex than single-modality models and typically require careful preprocessing of multiple input modalities. Check out the Repo here. Also, feel free to follow us on Twitter and don’t forget to join our 130k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post smol-audio: A Colab-Friendly Notebook Collection for Fine-Tuning Whisper, Parakeet, Voxtral, Granite Speech, and Audio Flamingo 3 appeared first on MarkTechPost.

smol-audio: A Colab-Friendly Notebook Collection for Fine-Tuning Whisper, Parakeet, Voxtral, Granite Speech, and Audio Flamingo 3 投稿を読む »

AI, Committee, ニュース, Uncategorized

Meta FAIR Releases NeuralSet: A Python Package for Neuro-AI That Supports fMRI, M/EEG, Spikes, and HuggingFace Embeddings

Researchers at Meta’s FAIR lab have released NeuralSet, a Python framework designed to eliminate one of the most persistent bottlenecks in Neuro-AI research: the painful, fragmented process of getting brain data into a deep learning pipeline. https://kingjr.github.io/files/neuralset.pdf The Problem: Neuroscience Data Is Stuck in the Pre-Deep-Learning Era Neuroscience already has excellent, battle-tested software. Tools like MNE-Python, EEGLAB, FieldTrip, Brainstorm, Nilearn, and fMRIPrep are the gold standard for signal processing across electrophysiology and neuroimaging. The trouble is that these tools were designed for a pre-deep-learning world: they rely on eager loading, assuming entire datasets fit into RAM, and they lack native abstractions to temporally align neural time series with high-dimensional embeddings from modern AI frameworks like HuggingFace Transformers. The result? Researchers spend enormous effort building ad-hoc pipelines that require manual data wrangling, manual caching, and complex backend configurations — just to get brain signals paired with, say, GPT-2 text embeddings for a single experiment. As public datasets on platforms like OpenNeuro now reach the terabyte scale, and experimental protocols increasingly incorporate continuous speech and video stimuli, this infrastructure gap is no longer just inconvenient — it is a scientific bottleneck. What NeuralSet Actually Does NeuralSet’s core design principle is structure–data decoupling. Instead of loading raw signals upfront, NeuralSet represents the logical structure of any experiment as lightweight, event-driven metadata — completely separate from the memory- and compute-intensive extraction of actual signals. The framework is organized around five core abstractions: Events, Extractors, Segments, Batch Data, and a Backend layer. In practice, everything in an experiment — an fMRI run, a word spoken during a task, a video stimulus — is modeled as an Event: a lightweight Python dictionary defined by a type, a start time, a duration, and a timeline (a unique identifier for a continuous recording session). A Study object assembles all events in an entire dataset into a single pandas DataFrame. Importantly, NeuralSet supports BIDS-compliant datasets, though it is not restricted to them. Because the DataFrame contains only lightweight metadata — not the raw signals themselves — engineers can filter, explore, and recombine massive datasets using standard pandas operations without loading a single byte of raw data into memory. Composable EventsTransform operations can then be chained to enrich or filter events — for example, annotating words with their sentence context, assigning cross-validation splits, or chunking long audio and video events into shorter segments. Multiple Study and Transform steps can also be composed together using a Chain, which creates a single reproducible, cacheable pipeline object. https://kingjr.github.io/files/neuralset.pdf Extractors: From Metadata to Tensors When it’s actually time to work with data, NeuralSet uses Extractors to bridge the gap between the metadata layer and numerical arrays required by machine learning models. For neural recordings, NeuralSet wraps the preprocessing stacks of domain-specific libraries directly: an FmriExtractor delegates to Nilearn for signal cleaning, spatial smoothing, and surface or atlas-based projection, while a MegExtractor or EegExtractor delegates to MNE-Python for filtering, re-referencing, and resampling. The same unified interface covers iEEG, fNIRS, EMG, and spike recordings — switching modalities requires only changing a configuration parameter, not rewriting a pipeline. For experimental stimuli, NeuralSet provides native integration with the HuggingFace ecosystem. A single HuggingFaceImage extractor can embed stimulus frames through DINOv2 or CLIP; analogous extractors exist for audio (Wav2Vec, Whisper), text (GPT-2, LLaMA), and video (VideoMAE). Critically, NeuralSet can expand a static embedding — say, a single vector per image — into a time series at an arbitrary frequency, so that stimulus representations are always temporally aligned with neural recordings. Extractors follow a three-phase execution model: configure (parameter validation at construction time), prepare (pre-compute and cache heavy outputs for all events), and extract (lazy retrieval from cache during model training). This means expensive computations — like running a large language model over every word in a corpus — are performed once and reused across experiments. The output of an Extractor for a single segment is Batch Data: a dictionary of tensors keyed by extractor name, along with the corresponding segments. Segmenter, DataLoader, and Cluster-Ready Infrastructure A Segmenter slices the events DataFrame into Segments — contiguous temporal windows representing single training examples — either on a sliding window grid or anchored to specific trigger events such as image or word onsets. The resulting SegmentDataset is a standard PyTorch Dataset, directly compatible with DataLoader, PyTorch Lightning, or any PyTorch-based framework. NeuralSet is built on the exca package, which handles deterministic, hash-based caching, full computational provenance, and hardware-agnostic execution. Changing a single preprocessing parameter invalidates only the affected downstream cache, leaving independent branches untouched. Full provenance is maintained, meaning any processed tensor can be traced back to the exact version of the raw data and the specific preprocessing chain used to generate it. Researchers can prototype on a single subject on their laptop, then dispatch 100 subjects to a SLURM-based HPC cluster by changing a single configuration flag — no infrastructure-specific code required. NeuralSet uses Pydantic to enforce strict schema validation at initialization time across every configurable object — Events, Studies, Extractors, Segmenters, and Transforms are all Pydantic BaseModel subclasses. This means a misconfigured parameter (for example, a negative filter frequency or an invalid BIDS directory path) raises a clear error immediately, before any job is submitted, rather than failing hours into a processing run. How It Stacks Up Against Existing Tools In the research paper, the research team presents a detailed comparison of NeuralSet against 18 existing neuroscience software packages across neural devices (fMRI, EEG, MEG, iEEG, spikes, and more), experimental task types (image, video, sound, text), and infrastructure features (Python support, memmap, batching, caching, cluster execution). NeuralSet is the only package in the comparison that achieves full support across all categories. Key Takeaways NeuralSet unifies brain data and AI in one pipeline. Researchers at Meta FAIR built NeuralSet to bridge the gap between diverse neural recordings (fMRI, M/EEG, spikes) and modern deep learning frameworks, delivering a single PyTorch-ready DataLoader for both. Structure–data decoupling eliminates memory bottlenecks. NeuralSet separates lightweight event metadata from heavy signal extraction, so AI devs and

Meta FAIR Releases NeuralSet: A Python Package for Neuro-AI That Supports fMRI, M/EEG, Spikes, and HuggingFace Embeddings 投稿を読む »

We use cookies to improve your experience and performance on our website. You can learn more at プライバシーポリシー and manage your privacy settings by clicking Settings.

Privacy Preferences

You can choose your cookie settings by turning on/off each type of cookie as you wish, except for essential cookies.

Allow All
Manage Consent Preferences
  • Always Active

Save
ja