YouZum

Uncategorized

AI, Committee, Notizie, Uncategorized

Inside Anduril and Meta’s quest to make smart glasses for warfare

The defense-tech company Anduril has shared new details about the augmented-reality headset for the military it’s prototyping with Meta, including a vision for ordering drone strikes via eye-tracking and voice commands. Quay Barnett, who leads the efforts as a vice president at Anduril following a career in the Army’s Special Operations Command, says his fundamental goal is to optimize “the human as a weapons system.” The vision is undoubtedly cyborg-inspired: Barnett wants drones and soldiers to see together, share information seamlessly, and make decisions as one.  Anduril actually has two such projects in the works. The first is the Army’s Soldier Born Mission Command, or SBMC, for which the company won a $159 million prototyping contract last year to work with Meta on augmented-reality glasses to attach to existing military helmets. But Anduril has also embarked on a self-funded side quest, announced in October, to design its own helmet and headset combo called EagleEye. This is something the military has not asked for, but Anduril insists it will prefer it and purchase it in the end. So far, both systems are years away. The Army isn’t expected to move its top choice for the SBMC program into production until 2028, if it picks one at all (the previous lead for the effort, Microsoft, was set to receive a $22 billion production contract that was ultimately cancelled when the glasses didn’t prove viable). But Barnett told MIT Technology Review about where both Anduril’s prototypes are headed. Depending on the situation, the glasses for either prototype will overlay certain information onto a soldier’s field of view. This might be as simple as a compass or as complex as an entire map of the area, information about where nearby drones are flying, or AI-driven recognition of a target like a truck.  The soldier would then speak to the interface in plain language—for example, to order an evacuation for someone who’s been injured or to plan a route taking into account which areas are off limits. A large language model—Anduril is in tests with Google’s Gemini, Meta’s Llama, and even Anthropic’s Claude, despite the company’s conflict with the Pentagon—will be used to help translate a soldier’s speech into commands the software can follow. And the engine for it all will be Anduril’s software Lattice, which incorporates data from lots of different military hardware into one picture. The Army announced in March that it would spend $20 billion to integrate Lattice with essentially its entire infrastructure. Barnett’s team is designing the headset to carry out multi-step tasks. A soldier might send a drone to surveil an area and instruct it to come back once it’s found something that looks like an artillery unit; then the system would recommend courses of action, like sending a nearby drone to strike, that would have to be approved by the normal chain of command. Leading the system through this, if all goes to plan, might not even require speech; the soldier could instead communicate through tracked eye movements and subtle taps. That’s the idea, anyway. It’s worked on early prototypes, Barnett says, but there aren’t yet versions ready for the Army to test at scale. The component parts began arriving in March. Because of federal military contracting rules, these parts—unlike Meta’s commercial smart glasses—required new supply chains that don’t rely on Chinese companies. It’s a lot for soldiers already bogged down in information overload, says Jonathan Wong, a former US Marine who works as a senior policy researcher at RAND on Army efforts to buy new tech. Both smart glasses projects aim to create a clean interface that presents only the right information at the right time. But it’s a product that soldiers will reject if it costs more of their attention than it saves. “How much mental bandwidth do you have to be both aware of your surroundings and to operate this technology in a way that makes you and your whole unit better?” he says. Wong recalls that as a platoon commander, for example, he had a radio that operated on three different channels at once. “The moment that two people were on different channels talking at the same time, I immediately couldn’t comprehend anything that either one of them was trying to tell me, and I was probably not aware of my own surroundings,” he says. “I think there are limits to what you can take in.” Ideally, Barnett says, smart glasses can ease that information overload. Anduril’s approach is to get creative with ways the user can access necessary information quickly. Voice commands and eye tracking are a piece of that strategy. But even if it’s all technically feasible, it might take years of field testing to know if the system is actually useful for soldiers, Wong says.  Such a system would mark a major escalation in how closely soldiers rely on imperfect AI systems. While computer vision models used to identify objects have long been employed by militaries, and chatbots have recently entered decision-making during the war in Iran, these technologies have not yet made their way to most frontline soldiers. A smart glasses system tasked with identifying threats and recommending strikes would introduce massive new risks of errors.  Anduril is not the only one competing to develop smart goggles for combat. Rivet, which specializes in wearable sensors for the military, received a $195 million prototyping contract the same time, and in March the Israeli defense-tech company Elbit received its own $120 million contract. This all comes after Microsoft lost its role leading the Army’s smart glasses effort, following a Pentagon audit that found the Army wasn’t properly testing the glasses, a mistake that could have wasted $22 billion. For both Anduril’s prototypes, the company is testing a new system for digital night vision, which uses electronic sensors and algorithms to boost low levels of light. It’s been a promised technology for decades but has tended to work too slowly for practical use and produce grainy images. Anduril says it has found improvements

Inside Anduril and Meta’s quest to make smart glasses for warfare Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

The Download: Musk v. Altman week 3, and Trump’s tech trading

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Musk v. Altman week 3: Musk and Altman traded blows over each other’s credibility. Now the jury will pick a side. In the final week of the Musk v. Altman trial, lawyers attacked the credibility of the two tech leaders. Sam Altman was accused of lying and self-dealing, while Elon Musk was portrayed as a power-seeker trying to control artificial general intelligence. The case unearthed new details about the two arch-rivals and OpenAI’s contested nonprofit status, as well as a golden trophy of a donkey’s ass awarded to an employee who challenged Musk. Read the full story on the explosive final week of the trial. —Michelle Kim Michelle Kim, who’s also a lawyer, has been in court throughout the Musk v. Altman trial. Read her coverage of week 1 and week 2, plus a Q&A on what it was like in the room.  The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 Trump traded hundreds of millions in tech stocks before favorable policy movesHe bought shares in Nvidia, AMD, and Arm ahead of policy boosts. (Quartz)+ And touted Palantir on Truth Social after buying its stock. (CNBC)+ His crypto venture and Iran’s top exchange tapped the same networks. (Reuters $) 2 SpaceX plans to list on the Nasdaq stock exchange as soon as June 12It wants to raise up to $75 billion at a $1.75 trillion valuation. (Reuters $)+ BlackRock may invest up to $10 billion in the offering. (The Information $)+ Cerebras’ blockbuster IPO has boosted hopes for the listing. (CNBC)+ Which is set to dwarf many of the biggest IPOs on ⁠record. (Reuters) 3 Chinese AI groups have pulled ahead of US rivals in video generationByteDance and Kuaishou’s models lead in realism and scale. (FT $)+ AI is fueling China’s short-drama boom. (MIT Technology Review)+ While its AI labs are betting big on open source. (MIT Technology Review) 4 Iran says it will charge Big Tech for using undersea internet cablesThe cables beneath the Strait of Hormuz carry vast digital traffic. (CNN)+ Tech bosses met at Uber HQ on Saturday to discuss Iran’s future. (404 Media) 5 Samsung has a “last chance” to stop a massive strike over AIOver 45,000 employees could walk out for 18 days this week. (CNBC) + They want a bigger share of the AI boom. (FT $)+ Samsung and its largest labor union will resume talks on Tuesday. (Reuters $) 6 Old oil and gas wells could become a new source of clean energyUS states plan to convert them into geothermal energy assets. (Wired $)+ A balcony solar boom is coming to the US. (MIT Technology Review) 7 The ChatGPT era has triggered a 30% surge in grades at a top universityGrades inflated in text-heavy courses but remained flat in others. (Axios)+ Princeton has changed its honor code because of AI cheating. (WSJ $)+ And real cheating rates may be far higher. (The Times $) 8 Ex-Google CEO Eric Schmidt was fiercely booed during an AI speechHis graduation speech praising AI agents sparked uproar. (The Verge)+ A populist backlash is building against AI. (MIT Technology Review) 9 Arm faces a US antitrust probe over its chip tech licensesRegulators are investigating whether it has an illegal monopoly. (Bloomberg $)+ Qualcomm has accused Arm of anticompetitive conduct. (Reuters $) 10 ArXiv will ban researchers who submit AI slopOffending authors face year-long bans from the pre-print server. (TechCrunch) Quote of the day “When someone offers you a seat on the rocket ship, you do not ask which seat. You just get on.”  —Ex-Google CEO Eric Schmidt extolls the virtues of AI agents in a graduation speech at the University of Arizona, prompting a chorus of boos. One More Thing WYSS INSTITUTE AT HARVARD UNIVERSITY Is this the end of animal testing? In a clean room in his lab, Sean Moore peers through a microscope at a bit of human intestinal tissue growing on a plastic chip. It’s one of 24 so-called “organs-on-chips” his team bought three years ago. The technology is designed to mimic human biology—and could reduce the need for animal testing. The appeal is not only ethical. Around 95% of drugs developed through animal research ultimately fail in people, and early studies suggest organ-on-a-chip systems may offer more accurate insights into how diseases behave and how drugs work. But the field still faces major technical and cost challenges before it can replace animal research. Find out how organ-on-chip technology could reshape drug testing. —Harriet Brown We can still have nice things A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.) + Listen to the captivating first recordings of whale songs from 1949.+ Meet the feline guardians of New York’s corner stores in this photo collection.+ A newly discovered floor plan allowed historians to pinpoint the location of Shakespeare’s only property in London.+ A music fan spent decades secretly recording 10,000 local shows. Now the entire collection is available online.

The Download: Musk v. Altman week 3, and Trump’s tech trading Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

Meet LiteLLM Agent Platform: A Kubernetes-Based, Self-Hosted Infrastructure Layer for Isolated Agent Sandboxes and Persistent Session Management in Production

Running AI agents in a local script is straightforward. Running them reliably in production across teams, across restarts, with isolated environments per context is a different problem entirely. BerriAI, the company behind the LiteLLM AI Gateway, is now open-sourcing a purpose-built answer to that problem: the LiteLLM Agent Platform. The platform is described as a simple, self-hosted infrastructure platform for running multiple agents in production. What Problem Does it Solve? It helps to understand what happens when you try to scale agents beyond a single process. Agents are stateful: they carry session history, tool call results, and intermediate reasoning across turns. If the container running your agent crashes, restarts, or gets replaced during a deployment, that session state is gone unless something is explicitly managing it. At the same time, different teams often need different runtime environments, different tools, different secrets, different access scopes which means you cannot throw all agents into one shared container. The platform manages two things: per-team and per-context sandboxes, and session continuity across pod restarts and upgrades. These two capabilities are the core infrastructure primitives the platform provides. Architecture and Technical Stack The platform is a standalone Next.js dashboard for LiteLLM v2 managed agents, covering sessions chat, agent CRUD, and live status. The codebase is primarily TypeScript (92.8%), with Shell scripts for provisioning, a Dockerfile for containerization, and CSS for the dashboard UI. The architecture separates concerns cleanly. A web process runs on port 3000 and serves the Next.js dashboard. A worker process handles async agent tasks. Postgres is used as the persistent backing store, and a schema migration runs as an init container on startup — so the database is always in the correct state before the application boots. For the sandbox layer — the isolated runtime environment where agents actually execute — sandboxes run on Kubernetes via the kubernetes-sigs/agent-sandbox CRD. Local development uses kind. If you are not already familiar with it: kind (Kubernetes in Docker) lets you spin up a full Kubernetes cluster locally using Docker containers as nodes, without needing a cloud provider. The agent-sandbox CRD (Custom Resource Definition) is a Kubernetes extension from kubernetes-sigs that the platform installs to manage the lifecycle of individual sandbox environments. The platform also includes a harness system under harnesses/opencode, which contains the configuration for running coding agents — such as Claude Code or OpenAI Codex — inside isolated sandboxes with a vault proxy for credential management. BerriAI team also maintains a separate litellm-agent-runtime repository, described as a coding-agent runtime that runs inside per-session VMs provisioned by a LiteLLM proxy, generic by design, with customization happening via harness configuration or a hydrate payload. One practical detail worth noting is how environment variables are handled across sandbox containers. Anything in .env prefixed with CONTAINER_ENV_ is injected into every sandbox container with the prefix stripped — for example, CONTAINER_ENV_GITHUB_TOKEN=ghp_… means the container sees GITHUB_TOKEN=ghp_… This gives teams a clean way to pass secrets into sandboxed agent sessions without modifying container images. https://github.com/BerriAI/litellm-agent-platform Getting Started The prerequisites for local development are Docker Desktop, kind, kubectl, helm, and a LiteLLM gateway. No cloud credentials are required to get started locally. The quickstart is two commands: Copy CodeCopiedUse a different Browser bin/kind-up.sh docker compose up bin/kind-up.sh is idempotent — it provisions a kind cluster named agent-sbx, installs the agent-sandbox controller, and loads the harness image. docker compose up boots Postgres, runs the schema migration, and starts the web process on port 3000 along with the worker. For production deployment, the recommended path is AWS EKS for the sandbox cluster and Render for the web and worker processes. bin/eks-up.sh provisions the EKS cluster, and a Render Blueprint provides a one-click deployment option. Relationship to the LiteLLM Gateway The Agent Platform is a layer on top of the existing LiteLLM ecosystem, not a replacement for it. LiteLLM’s core is a Python SDK and Proxy Server — an AI Gateway — that calls 100+ LLM APIs in OpenAI format, with cost tracking, guardrails, load balancing, and logging, supporting providers including Bedrock, Azure, OpenAI, VertexAI, Cohere, Anthropic, SageMaker, HuggingFace, vLLM, and NVIDIA NIM. The Agent Platform consumes a running LiteLLM gateway as a dependency and builds agent orchestration and session management infrastructure on top of it. Model routing, cost tracking, and rate limiting remain in the gateway layer. Sandbox isolation, session continuity, and the management dashboard are handled by the Agent Platform. Marktechpost’s Visual Explainer LiteLLM Agent Platform Self-Hosted Agent Infrastructure Guide Alpha Overview Concepts Architecture Prerequisites Quickstart Production 01 / 06 What is LiteLLM Agent Platform? BerriAI open-sourced this platform on May 8, 2026. It is a self-hosted infrastructure layer for running multiple AI agents in production, built on top of the LiteLLM AI Gateway. Self-Hosted Runs entirely on your own infrastructure. No data leaves your environment. Suited for regulated industries and teams with data residency requirements. Multi-Agent Designed to run multiple agents in parallel, with full isolation between teams and contexts using per-session sandboxes. Session Continuity Agent sessions persist across pod restarts and upgrades, so stateful work is not lost when containers are replaced. Open Source (MIT) Fully open source under the MIT license. Repo: github.com/BerriAI/litellm-agent-platform. File issues and contribute directly. Prerequisite Knowledge This guide assumes familiarity with Docker, basic command-line usage, and a general understanding of what an AI agent is (a model that calls tools and runs multi-step tasks). Kubernetes experience helps but is not required to follow along. 02 / 06 Key Concepts to Know First Before running the platform, understand these four building blocks. They appear throughout the setup and configuration. A LiteLLM Gateway The underlying AI Gateway that the Agent Platform depends on. It routes requests to 100+ LLM providers (OpenAI, Anthropic, Bedrock, VertexAI, etc.) using a unified OpenAI-format API. The Agent Platform does not include the gateway, you must have one running separately and point the platform at it. B Sandbox An isolated container environment where a single agent session executes. Each sandbox is independent, meaning one agent cannot access the filesystem, secrets, or state

Meet LiteLLM Agent Platform: A Kubernetes-Based, Self-Hosted Infrastructure Layer for Isolated Agent Sandboxes and Persistent Session Management in Production Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

A Coding Guide Implementing SHAP Explainability Workflows with Explainer Comparisons, Maskers, Interactions, Drift, and Black-Box Models

In this tutorial, we implement SHAP workflows as a practical framework for interpreting machine learning models beyond basic feature-importance plots. We start by training tree-based models and then compare different SHAP explainers, including Tree, Exact, Permutation, and Kernel methods, to understand how accuracy and runtime change across model-aware and model-agnostic approaches. We also examine how maskers affect explanations when features are correlated, how interaction values reveal pairwise feature effects, and how link functions alter interpretation between the log-odds and probability spaces. Also, we use Owen values, cohort testing, SHAP-based feature selection, drift monitoring, and custom black-box explanations to build a complete interpretability workflow that can run directly in Google Colab. Copy CodeCopiedUse a different Browser !pip install -q –upgrade shap xgboost transformers import warnings, time, numpy as np, pandas as pd, matplotlib.pyplot as plt from scipy import stats from scipy.cluster import hierarchy warnings.filterwarnings(“ignore”) import shap, xgboost as xgb from sklearn.datasets import fetch_california_housing, load_breast_cancer from sklearn.model_selection import train_test_split from sklearn.metrics import roc_auc_score, r2_score shap.initjs() np.random.seed(42) print(f”SHAP: {shap.__version__}n”) housing = fetch_california_housing() X = pd.DataFrame(housing.data, columns=housing.feature_names) y = pd.Series(housing.target, name=”MedHouseVal”) reg = xgb.XGBRegressor(n_estimators=300, max_depth=5, learning_rate=0.05, subsample=0.9, random_state=42, n_jobs=-1).fit(X_tr, y_tr) print(f”Housing regressor R² = {reg.score(X_te, y_te):.3f}”) def reg_predict(X): return reg.predict(np.asarray(X)) We install the required libraries and import the core tools for SHAP, XGBoost, statistics, visualization, and model evaluation. We load the California housing dataset and train an XGBoost regression model. We also define a clean prediction wrapper so that SHAP can explain the model without running into compatibility issues with bound model methods. Copy CodeCopiedUse a different Browser print(“n” + “=”*72) print(“PART 1: Explainer comparison — correctness & speed”) print(“=”*72) X_sample = X_te.iloc[:25] bg_small = shap.sample(X_tr, 50, random_state=42) def _wrap_kernel(expl, X, bg_mean): vals = expl.shap_values(X, nsamples=200, silent=True) return shap.Explanation(values=vals, base_values=np.full(len(X), bg_mean), data=X.values, feature_names=X.columns.tolist()) runs = {} t0 = time.time(); tree_expl = shap.TreeExplainer(reg); sv_tree = tree_expl(X_sample) runs[“Tree (exact, model-aware)”] = (sv_tree, time.time() – t0) t0 = time.time() sv_exact = shap.Explainer(reg_predict, bg_small, algorithm=”exact”)(X_sample) runs[“Exact (model-agnostic)”] = (sv_exact, time.time() – t0) t0 = time.time() sv_perm = shap.Explainer(reg_predict, bg_small, algorithm=”permutation”)(X_sample) runs[“Permutation”] = (sv_perm, time.time() – t0) t0 = time.time() ke = shap.KernelExplainer(reg_predict, shap.sample(X_tr, 50, random_state=42).values) sv_kern = _wrap_kernel(ke, X_sample, ke.expected_value) runs[“Kernel”] = (sv_kern, time.time() – t0) ref = sv_tree.values.flatten() print(f”n{‘Method’:30s} {‘time(s)’:>8s} {‘ρ vs Tree’:>10s} {‘max|Δ|’:>8s}”) for name, (sv, dt) in runs.items(): flat = sv.values.flatten() rho = np.corrcoef(ref, flat)[0, 1] err = np.abs(ref – flat).max() print(f”{name:30s} {dt:8.2f} {rho:10.4f} {err:8.4f}”) print(“nTakeaway: Tree is the only exact + fast option for tree ensembles.”) print(“Exact ≈ Permutation when permutation has enough samples; Kernel is noisier and slowest.”) print(“n” + “=”*72) print(“PART 2: Maskers — Independent vs Partition under correlation”) print(“=”*72) corr = X_tr.corr().abs() top_pair = corr.where(np.triu(np.ones_like(corr, dtype=bool), k=1)) .stack().sort_values(ascending=False).head(3) print(“Top correlated pairs (|ρ|):”) for (a, b), v in top_pair.items(): print(f” {a:10s} {b:10s} |ρ| = {v:.3f}”) masker_ind = shap.maskers.Independent(X_tr, max_samples=100) masker_part = shap.maskers.Partition(X_tr, max_samples=100) sv_ind = shap.Explainer(reg_predict, masker_ind)(X_sample) sv_part = shap.Explainer(reg_predict, masker_part)(X_sample) a, b = top_pair.index[0] print(f”nMean |φ| for top-correlated pair ({a}, {b}):”) print(f” Independent : {a}={np.abs(sv_ind[:,a].values).mean():.4f} {b}={np.abs(sv_ind[:,b].values).mean():.4f}”) print(f” Partition : {a}={np.abs(sv_part[:,a].values).mean():.4f} {b}={np.abs(sv_part[:,b].values).mean():.4f}”) print(“Partition redistributes credit across correlated features (on-manifold semantics).”) fig, axes = plt.subplots(1, 2, figsize=(13, 4)) plt.sca(axes[0]); shap.plots.bar(sv_ind, show=False); axes[0].set_title(“Independent masker”) plt.sca(axes[1]); shap.plots.bar(sv_part, show=False); axes[1].set_title(“Partition masker”) plt.tight_layout(); plt.show() We compare multiple SHAP explainers, including Tree, Exact, Permutation, and Kernel, on the same regression model and sample data. We measure each method by runtime, correlation with TreeExplainer, and maximum attribution difference to understand the trade-off between speed and approximation quality. We then study Independent and Partition maskers to see how correlated features receive different attribution credit under different masking assumptions. Copy CodeCopiedUse a different Browser print(“n” + “=”*72) print(“PART 3: Interaction decomposition”) print(“=”*72) inter = tree_expl.shap_interaction_values(X_te.iloc[:500]) inter_abs = np.abs(inter).mean(0) diag = np.diagonal(inter_abs).copy() off = inter_abs.copy(); np.fill_diagonal(off, 0) main_share = diag.sum() / (diag.sum() + off.sum()) print(f”Total attribution mass: {main_share*100:.1f}% main effects, ” f”{(1-main_share)*100:.1f}% interactions”) pairs = [(X.columns[i], X.columns[j], off[i, j]) for i in range(X.shape[1]) for j in range(i+1, X.shape[1])] pairs.sort(key=lambda t: -t[2]) print(“nTop 5 interaction pairs (mean |φ_ij|):”) for a, b, v in pairs[:5]: print(f” {a:10s} × {b:10s} → {v:.4f}”) fig, ax = plt.subplots(figsize=(7.5, 6)) im = ax.imshow(off, cmap=”viridis”) ax.set_xticks(range(X.shape[1])); ax.set_xticklabels(X.columns, rotation=45, ha=”right”) ax.set_yticks(range(X.shape[1])); ax.set_yticklabels(X.columns) plt.colorbar(im, label=”mean |φ_ij|”); plt.title(“Pairwise interaction strength”) plt.tight_layout(); plt.show() a, b, _ = pairs[0] i, j = X.columns.get_loc(a), X.columns.get_loc(b) xs = X_te.iloc[:500][a].values; cs = X_te.iloc[:500][b].values fig, axes = plt.subplots(1, 2, figsize=(13, 4), sharex=True) axes[0].scatter(xs, inter[:, i, i], c=cs, s=12, cmap=”coolwarm”) axes[0].set_title(f”Main effect of {a}”); axes[0].set_xlabel(a); axes[0].set_ylabel(“φ_{ii}”) sc = axes[1].scatter(xs, 2*inter[:, i, j], c=cs, s=12, cmap=”coolwarm”) axes[1].set_title(f”Interaction {a} × {b}”); axes[1].set_xlabel(a); axes[1].set_ylabel(“2·φ_{ij}”) plt.colorbar(sc, ax=axes[1], label=b); plt.tight_layout(); plt.show() print(“n” + “=”*72) print(“PART 4: Link functions — logit vs probability space”) print(“=”*72) cancer = load_breast_cancer() Xc = pd.DataFrame(cancer.data, columns=cancer.feature_names) yc = pd.Series(cancer.target) clf = xgb.XGBClassifier(n_estimators=300, max_depth=4, learning_rate=0.05, eval_metric=”logloss”, random_state=42).fit(Xc_tr, yc_tr) print(f”AUC = {roc_auc_score(yc_te, clf.predict_proba(Xc_te)[:,1]):.3f}”) expl_logit = shap.TreeExplainer(clf) sv_logit = expl_logit(Xc_te) expl_prob = shap.TreeExplainer(clf, Xc_tr.sample(100, random_state=42), model_output=”probability”) sv_prob = expl_prob(Xc_te) print(f”nSample 0 reconstruction (φ should sum to f – E[f]):”) print(f” log-odds : base + Σφ = {sv_logit.base_values[0] + sv_logit.values[0].sum():+.3f}”) print(f” prob : base + Σφ = {sv_prob.base_values[0] + sv_prob.values[0].sum():.3f} ” f”(model proba = {clf.predict_proba(Xc_te.iloc[[0]])[0,1]:.3f})”) fig, axes = plt.subplots(1, 2, figsize=(15, 5)) plt.sca(axes[0]); shap.plots.waterfall(sv_logit[0], max_display=8, show=False); axes[0].set_title(“Log-odds space”) plt.sca(axes[1]); shap.plots.waterfall(sv_prob[0], max_display=8, show=False); axes[1].set_title(“Probability space”) plt.tight_layout(); plt.show() We calculate SHAP interaction values to separate main feature effects from pairwise interaction effects in the housing model. We identify the strongest interaction pairs and visualize their attribution strength using heatmaps and scatter plots. We then move to a classification task and compare SHAP explanations in log-odds and probability spaces using a breast cancer classifier. Copy CodeCopiedUse a different Browser print(“n” + “=”*72) print(“PART 5: Owen values from a correlation-based feature hierarchy”) print(“=”*72) D = 1 – X_tr.corr().abs().values np.fill_diagonal(D, 0) condensed = D[np.triu_indices_from(D, k=1)] linkage = hierarchy.linkage(condensed, method=”average”) masker_owen = shap.maskers.Partition(X_tr, clustering=linkage, max_samples=100) sv_owen = shap.Explainer(reg_predict, masker_owen)(X_sample) fig, axes = plt.subplots(1, 2, figsize=(14, 4.5)) hierarchy.dendrogram(linkage, labels=X.columns.tolist(), ax=axes[0]) axes[0].set_title(“Feature hierarchy (1 − |ρ|)”) plt.sca(axes[1]); shap.plots.bar(sv_owen.abs.mean(0), show=False) axes[1].set_title(“Owen values (cluster-aware)”) plt.tight_layout(); plt.show() print(“n” + “=”*72) print(“PART 6: Cohort comparison with bootstrap CIs and hypothesis tests”) print(“=”*72) sv_all = tree_expl(X_te) q1, q3 = X_te[“MedInc”].quantile([0.25, 0.75]) low = (X_te[“MedInc”] <= q1).values high = (X_te[“MedInc”] >= q3).values def boot_ci(v, B=1000, seed=0):

A Coding Guide Implementing SHAP Explainability Workflows with Explainer Comparisons, Maskers, Interactions, Drift, and Black-Box Models Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

Nous Research Proposes Lighthouse Attention: A Training-Only Selection-Based Hierarchical Attention That Delivers 1.4–1.7× Pretraining Speedup at Long Context

Training large language models on long sequences has a well-known problem: attention is expensive. The scaled dot-product attention (SDPA) at the core of every transformer scales quadratically Θ(N²) in both compute and memory with sequence length N. FlashAttention addressed this through IO-aware tiling that avoids materializing the full N×N attention matrix in high-bandwidth memory, reducing the memory footprint significantly, but the underlying Θ(N²) compute scaling remains. Researchers at Nous Research have introduced a new method called Lighthouse Attention that addresses this bottleneck specifically at pretraining time, achieving a 1.40× to 1.69× end-to-end wall-clock speedup against a cuDNN-backed SDPA baseline, with matching or lower final training loss. The core problem with existing sparse attention methods To understand why Lighthouse works the way it does, it helps to know what existing sparse attention methods do. Most prior work like NSA, HISA, DSA, MoBA makes the same two design decisions. First, they pool only the key and value side while leaving queries at full resolution (asymmetric compression). Second, their selection logic lives inside a custom attention kernel, which means teams can’t reuse the optimized dense-attention kernels that modern GPU tensor cores are built around. There is also a concern specific to training that inference-only sparse methods don’t face. An inference-time sparse method is evaluated only against its dense backbone and it is at most as good as that backbone. A training-time sparse method faces a harder test: once training is done, will the resulting weights still produce a competent dense-attention model at inference? Lighthouse treats that question as its central correctness criterion. Lighthouse takes a different approach on both design decisions. It pools queries, keys, and values symmetrically across a multi-level pyramid, and it places selection entirely outside the attention kernel. After selection, the system gathers the chosen entries into a contiguous, dense sub-sequence and runs stock FlashAttention on it — the same kernel used by the dense baseline. https://arxiv.org/pdf/2605.06554 How the four-stage pipeline works A Lighthouse attention layer wraps around, but does not modify, scaled dot-product attention. The pipeline has four stages. In the first stage, average pooling constructs an L-level pyramid from Q, K, and V. With pooling factor p, level ℓ of the pyramid has N/p^ℓ tokens, each summarizing p^ℓ base positions. Crucially, the same pooling applies to all three projections, producing coherent (Q^(ℓ), K^(ℓ), V^(ℓ)) triples at every level. Total pyramid construction costs Θ(N) time and memory. In the second stage, a parameter-free scorer assigns each pyramid entry two scalar scores using per-head ℓ₂ norms: one as a query score (∥Q^(ℓ)_i∥₂) and one as a key score (∥K^(ℓ)_i∥₂). Coarser levels inherit scores from finer ones via max-pooling, so a coarse span picks up the importance of its strongest token. A fused chunked-bitonic top-K kernel then selects k entries jointly across all pyramid levels. One design detail worth noting: the coarsest pyramid level is always retained in full — it is cheap and guarantees at least one contributor at every base position; the remaining selection budget is spent on finer levels. Additionally, the chunked-bitonic design produces a stratified top-K rather than a strict global top-K: the score stream is partitioned into fixed-size chunks, each maintaining an in-register top-m buffer, so if the k globally highest-scoring entries clustered in one chunk, some would be replaced by lower-scoring entries from other chunks. The result is more balanced attention coverage across the sequence and avoids selection collapse onto a narrow span. The top-K step is discrete and non-differentiable — no straight-through estimator, no Gumbel softmax. Selection indices carry no gradient. Gradients flow only through the gathered Q, K, V entries into WQ, WK, WV, so the projections learn to produce values that are useful when selected rather than scores that are good at selecting. In the third stage, the selected entries are gathered into a contiguous sub-sequence of length S = N/p^(L−1) + (L−1)·p·k and passed to standard FlashAttention. At N = 1,000,000 with L = 4, p = 4, k = 4,096, S ≈ 65,000 — far smaller than N. A critical property of the gathering process is that it guarantees no “holes” or empty spaces in the assembled sub-sequence. This matters specifically because Lighthouse also compresses queries: a gap in the sequence would mean those missing tokens have no gradient path during the backward pass and could cause training instabilities. Asymmetric methods that leave queries at full resolution don’t face this problem, but Lighthouse’s symmetric design requires that the gathered sub-sequence remains fully dense. In the fourth stage, each output entry is scattered back to the p^ℓ base positions it represents via a deterministic integer-atomic scatter kernel, with a shift of p^ℓ − 1 to preserve causality. The per-position fan-in is bounded by L regardless of k. https://arxiv.org/pdf/2605.06554 Why symmetric pooling changes the compute Pooling queries alongside keys and values changes the computational character of the attention call from O(N Sd) to O(S² d) at training time. Because S ≪ N at long contexts, this is what produces the latency advantage. Benchmarked on a single NVIDIA B200 at 512K context (bfloat16, B=1, H=8, head dimension 128, L=3, p=4, sparsity ≈ 1:64), Lighthouse is 21× faster on the forward pass and 17.3× faster on the combined forward+backward pass relative to cuDNN-backed SDPA. From an asymptotic standpoint, setting L = logp(N/k) gives a gathered sub-sequence size of S = Θ(k log N), which makes the dense FlashAttention call cost Θ(k² log² N d) — polylogarithmic in N at fixed k. Combined with the linear-cost stages (pyramid construction, scoring, scatter-back), total per-layer compute is Θ(T d) at bounded k — the same asymptotic class as linear attention and SSMs — while preserving softmax attention’s recall properties on the selected sub-sequence. Inference is a different constraint. Autoregressive decoding presents one query at a time, which violates the assumption that all queries co-occur in one forward pass. Lighthouse is a training-only method, and the symmetric pooling design cannot be used directly at inference. The two-stage training recipe and recoverability The experimental setup used a 530M-parameter

Nous Research Proposes Lighthouse Attention: A Training-Only Selection-Based Hierarchical Attention That Delivers 1.4–1.7× Pretraining Speedup at Long Context Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

Vercel Labs Introduces Zero, a Systems Programming Language Designed So AI Agents Can Read, Repair, and Ship Native Programs

Most programming languages were designed for humans who read error messages, interpret warnings, and manually trace through stack output to fix bugs. AI agents do none of those things well. They work better with structured data: predictable tokens, stable codes, and machine-parseable repair hints. That gap is what Vercel Labs is trying to close by releasing Zero, an experimental systems language that is faster, smaller, and easier for agents to use and repair. What is Zero Language Zero is a systems programming language that sits in the same design space as C or Rust. It compiles to native executables, gives you explicit memory control, and targets low-level environments. What separates Zero from existing systems languages is that its compiler output and toolchain were designed from day one to be consumed by AI agents, not just human engineers. The Agent-First Toolchain The core problem Zero addresses is how agents interact with compiler feedback. In a typical development loop involving a coding agent, the agent writes code, the compiler emits an error as unstructured text, and the agent has to parse that text to determine what went wrong and how to fix it. This is fragile — error message formats change, messages are written for human readers, and there’s no built-in concept of a ‘repair action.’ Zero’s CLI emits structured JSON diagnostics by default. When you run zero check –json, the output looks like: Copy CodeCopiedUse a different Browser { “ok”: false, “diagnostics”: [{ “code”: “NAM003”, “message”: “unknown identifier”, “line”: 3, “repair”: { “id”: “declare-missing-symbol” } }] } Each diagnostic carries a stable code (e.g., NAM003), a human-readable message, a line reference, and a repair object with a typed repair ID. Humans read the message. Agents read the code and repair. The same CLI command surfaces both — there is no separate mode or secondary tool to run. The toolchain is unified into one binary: zero check, zero run, zero build, zero graph, zero size, zero routes, zero skills, zero explain, zero fix, and zero doctor are all subcommands of the same CLI. This matters for agentic workflows because agents don’t need to reason about which tool to invoke for which task. Two subcommands are particularly relevant to the repair loop. zero explain <diagnostic-code> returns a detailed explanation of a given diagnostic code, so an agent can look up NAM003 without parsing prose documentation. zero fix –plan –json <file-or-package> emits a structured fix plan — a machine-readable description of what changes to make to resolve a diagnostic — rather than requiring the agent to infer the fix from the error message alone. zero skills serves a different purpose: it provides version-matched agent guidance directly through the CLI. Running zero skills get zero –full returns focused workflows covering Zero syntax, diagnostics, builds, packages, standard library use, testing, and agent edit loops — all matched to the installed compiler version. This is notable because it means agents working with Zero don’t need to scrape external documentation that may be out of sync with the compiler they’re actually running. Explicit Effects and Capability-Based I/O One of Zero’s core design decisions is that effects are explicit in function signatures. If a function writes to standard output, accesses the filesystem, or makes a network call, it must declare that through a capability object. The canonical entry point in Zero looks like this: Copy CodeCopiedUse a different Browser pub fun main(world: World) -> Void raises { check world.out.write(“hello from zeron”) } The world: World parameter is the capability object that grants access to the outside world. A function that doesn’t receive World (or a capability derived from it) cannot perform I/O — the compiler enforces this at compile time, not at runtime. There is no hidden global process object. The check keyword handles fallible operations. If world.out.write(…) can fail, check surfaces that failure along the call stack. The raises annotation on main signals that the function can propagate errors — making error paths visible in signatures rather than buried in runtime exceptions. Getting Started Installation requires one command: Copy CodeCopiedUse a different Browser curl -fsSL https://zerolang.ai/install.sh | bash export PATH=”$HOME/.zero/bin:$PATH” zero –version The installer downloads the latest binary from the GitHub release and places it in $HOME/.zero/bin/zero. Packages are defined with a zero.json manifest and source files under src/, initialized with zero new cli <name>. A VS Code extension for .0 file syntax highlighting ships in the repository under extensions/vscode/. Marktechpost’s Visual Explainer Vercel Labs 01 / 09  ·  Overview Zero The Programming Languagefor Agents An experimental systems language that gives AI agents structured diagnostics, typed repair metadata, and machine-readable docs — alongside sub-10 KiB native binaries. Systems Language Agent-Native v0.1.1 Apache-2.0 Experimental Released May 15, 2026 Author Chris Tate · Vercel Labs Repo vercel-labs/zero File Extension .0 Context 02 / 09  ·  Why Zero Exists The Agent Repair Loop Problem Most programming languages produce compiler output written for human readers — unstructured text that AI agents must parse to determine what failed and how to fix it. This creates a fragile loop. Agent writes code — compiler emits an error as unstructured text Agent parses text — error format can change between compiler versions No repair hint — there’s no built-in concept of a “repair action” Human steps in — the loop requires manual intervention to resolve errors Zero was designed from day zero so agents can read the code, interpret the diagnostics, and repair the program — without human translation. Core Feature 03 / 09  ·  JSON Diagnostics Structured Compiler Output Running zero check ——json emits machine-readable diagnostics instead of plain text. Every error includes a stable code, a human message, a line number, and a typed repair ID. $ zero check –json { “ok”: false, “diagnostics”: [{ “code”: “NAM003”, “message”: “unknown identifier”, “line”: 3, “repair”: { “id”: “declare-missing-symbol” } }] } code — stable identifier agents can match reliably (NAM003) message — human-readable description of the error repair — typed repair ID agents can act on without text parsing Core Feature 04 / 09

Vercel Labs Introduces Zero, a Systems Programming Language Designed So AI Agents Can Read, Repair, and Ship Native Programs Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

Zyphra Releases ZAYA1-8B-Diffusion-Preview: The First MoE Diffusion Model Converted From an Autoregressive LLM With Up to 7.7x Speedup

Zyphra, the San Francisco-based AI lab behind the ZAYA1 model family, released ZAYA1-8B-Diffusion-Preview — a preview of its early work in diffusion-language models. The release demonstrates that an existing autoregressive language model can be converted into a discrete diffusion model with no systematic loss of evaluation performance, while delivering substantial inference speedups on AMD hardware. https://www.zyphra.com/post/zaya1-8b-diffusion-preview The Problem With Autoregressive Decoding To understand why this matters, it helps to first understand how most language models generate text today. Standard large language models are autoregressive: they decode one token at a time in sequence. For each new token, the attention mechanism has to look back over all previously generated tokens and load their stored representations — called the KV-cache — from GPU memory. Crucially, because every user in a batch has a different history of tokens, each user’s KV-cache must be loaded separately and cannot be shared across requests. This creates a bottleneck. When the GPU spends more time moving data from memory than performing actual computation, the system becomes memory-bandwidth bound rather than compute-bound. This limits how efficiently modern GPU hardware — which has been scaling compute FLOPs faster than memory bandwidth — can be used during inference. Diffusion offers an alternative. Instead of generating one token at a time, a diffusion model generates multiple drafts of N tokens simultaneously and iterates this drafting process multiple times. Because all N tokens in the block share the same KV-cache, the operation shifts from memory-bandwidth bound to compute-bound, which means the GPU can be utilized more efficiently. In ZAYA1-8B-Diffusion-Preview specifically, the model performs a single-step transformation from mask to token for each token in the block — meaning it directly predicts the unmasked token in one step rather than iteratively denoising. Converting Autoregression to Diffusion Without Training From Scratch Training a diffusion language model from scratch is technically difficult, and there are few established recipes for doing so. Zyphra team offers two reasons for preferring conversion over training from scratch: first, it is simply hard, with few known recipes; second, there is no advantage to training in diffusion-mode because training is already compute-bound — the memory-bandwidth bottleneck that diffusion solves only appears at inference time. This means all the benefits of diffusion are inference-time benefits, and an existing pretraining stack can be reused as-is. Building on the TiDAR recipe, Zyphra took the ZAYA1-8B-base checkpoint and performed an additional 600 billion tokens of diffusion-conversion mid-training at a 32k context length, followed by 500 billion tokens of native context extension to 128k, and then a diffusion supervised fine-tuning (SFT) phase. ZAYA1-8B-Diffusion-Preview is the first MoE diffusion model converted from an autoregressive LLM, and the first diffusion-language model to be trained on AMD GPUs. Zyphra reports minimal evaluation degradation compared to the base autoregressive checkpoint, with gains on some benchmarks such as LCB-v6. They attribute this partly to improved mid-training datasets and partly to the greater expressivity of diffusion-style within-block non-causal inference compared to causal autoregression. How the Diffusion Sampler Works During inference, ZAYA1-8B-Diffusion-Preview generates a draft of 16 tokens simultaneously. A fraction of these tokens are accepted based on a sampling criterion borrowed from speculative decoding. The key advantage here is that the same model acts as both speculator and verifier within a single forward pass, which removes the overhead associated with running two separate models as in traditional methods like EAGLE or dFlash. In heavily memory-bandwidth-bound regimes, almost all accepted tokens represent free speedup over autoregressive decoding — the GPU is already loaded and the extra tokens cost very little additional compute. Zyphra team reports two samplers with different speed-quality trade-offs: Lossless diffusion sampler: Uses the standard speculative decoding acceptance criterion of min(1, p(x)/q(x)), where p is the autoregressive model’s logit distribution and q is the diffusion model’s distribution. Upon rejection, the next token is sampled from the residual distribution of p(x)-q(x). This sampler achieves a 4.6x speedup with no systematic evaluation degradation. Logit-mixing sampler: First mixes the logits from the diffusion speculator and the autoregressive model, then uses the averaged distribution for verification. This improves acceptance rates because the verification logits are closer to the diffusion logits, but has some impact on quality. This sampler achieves a 7.7x speedup. The trade-off between speed and quality can be chosen at runtime. One important caveat on these numbers: because ZAYA1-8B-Diffusion-Preview is a base mid-train checkpoint that has not yet undergone RL training, Zyphra uses pass@ evaluations rather than standard accuracy benchmarks to better represent the model’s ultimate potential after RL training. Readers comparing these figures to other models’ reported benchmarks should keep this in mind. Zyphra team also notes that the speedups observed from diffusion are higher than those from alternative methods such as multi-token prediction (MTP) and various speculative decoding strategies such as EAGLE3. Since TiDAR-style diffusion models utilize a single forward pass only, acceptance rates comparable to dFlash still yield substantial speedups. https://www.zyphra.com/post/zaya1-8b-diffusion-preview Architecture Details ZAYA1-8B-Diffusion-Preview is a single-step speculative diffusion model that uses order constrained generation which means the diffusion model is only capable of generating tokens in a contiguous subsequence starting from the prefix. This constraint increases training stability dramatically compared to unconstrained mask diffusion objectives or set block decoding, and was a primary reason Zyphra built on the TiDAR recipe. The model uses ZAYA1-8B’s existing CCA attention variant from Zyphra. CCA dramatically reduces prefill FLOPs in attention, which is directly beneficial for diffusion because diffusion converts decoding into a prefill-like operation. This means CCA lets the model diffuse more tokens in parallel before hitting compute limits. More specifically, the architecture uses CCGQA with a 4:1 ratio between query heads and key heads. One design choice behind this was deliberately avoiding MLA (Multi-Head Latent Attention), whose high arithmetic intensity was seen as a mismatch compared to CCGQA. Since block diffusion accesses the same cache, arithmetic intensity scales with block size and with the number of blocks per forward pass. On AMD MI300x hardware in bf16, the system supports roughly three block-sized proposals per single forward pass; on MI355x, this

Zyphra Releases ZAYA1-8B-Diffusion-Preview: The First MoE Diffusion Model Converted From an Autoregressive LLM With Up to 7.7x Speedup Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

Musk v. Altman week 3: Elon Musk and Sam Altman traded blows over each other’s credibility. Now the jury will pick a side.

In the final week of the Musk v. Altman trial, lawyers traded blows over Elon Musk’s and OpenAI CEO Sam Altman’s credibility. Altman was grilled on his alleged history of lying and self-dealing involving companies that do business with OpenAI. But he fired back, painting Musk as a power-seeker who wanted to control the development of artificial general intelligence (AGI)—powerful AI that can compete with humans on most cognitive tasks.  As evidence of their commitment to AI safety, OpenAI brought out a golden trophy of a donkey’s ass that was gifted to an employee after he was called a “jackass” for standing up to Musk’s plans to race toward AGI.  Lawyers for both sides also presented their closing arguments, floating unflattering mugshot-style photos of Musk and Altman next to each other on a giant screen. Musk’s lawyer Steven Molo argued that Altman and OpenAI president Greg Brockman broke their promise to use money Musk donated to maintain OpenAI as a nonprofit that develops AI for the benefit of humanity. Instead, they created a for-profit subsidiary that made them extraordinarily wealthy. OpenAI’s lawyer Sarah Eddy argued that Altman and Brockman never promised to keep OpenAI a nonprofit. She added that even though it’s been restructured, OpenAI remains a nonprofit dedicated to developing AI safely. She claimed that Musk sued too late—and that his real motive is to sabotage a competitor to his own AI company, xAI, which he launched in 2023.  Musk is asking the court to unwind the 2025 restructuring that converted OpenAI’s for-profit subsidiary into a public benefit corporation and to remove Altman and Brockman from their roles. He is also seeking as much as $134 billion in damages from OpenAI and Microsoft, to be awarded to OpenAI’s nonprofit.  The jury will begin deliberating on Monday and deliver an advisory verdict as soon as next week. The jury verdict is not binding on the judge, who will decide the case. If the judge rules in Musk’s favor, it could upend OpenAI’s race toward an IPO at a valuation approaching $1 trillion. Meanwhile, xAI is expected to go public as a part of Musk’s rocket company SpaceX as early as June, at a target valuation of $1.75 trillion. Musk the power-seeker, Altman the liar. In the first week of the trial, Musk said he was suing to save OpenAI’s mission to build AI safely for the benefit of humanity. This week, Altman denied Musk was a paladin of AI safety and painted him as a power-seeker who wanted to control OpenAI.  Altman told the jury that in 2017, when Musk and other cofounders were discussing creating a for-profit arm, they asked Musk what would happen to his control over such an entity if he died. “Maybe the control of OpenAI should pass to my children,” Musk said, according to Altman. Musk’s lawyer shot back, grilling Altman on his alleged history of lying. He pointed out that OpenAI’s former executives Ilya Sutskever and Mira Murati, and former board members Helen Toner and Tasha McCauley, all testified that Altman had lied to them. In 2023, Altman was briefly fired as CEO over the alleged behavior. Molo also pressed Altman about his personal investments in startups that do business with OpenAI. Altman testified that he tried to steer OpenAI to buying power from the nuclear energy company Helion Energy, a third of which he owns. (Last Friday, the US House oversight committee launched an investigation into Altman’s potential conflicts of interest. Attorneys general from more than a half-dozen states called for the Securities and Exchange Commission to review them.) During his closing statement, Molo put Altman’s credibility on the stand again. “Imagine that you’re on a hike, and you come upon one of those wooden bridges that you see on a trail, and it’s over a gorge,” he said. “A woman standing by the entry to the bridge says, ‘Don’t worry—the bridge is built on Sam Altman’s version of the truth.’ Would you walk across that bridge?” Altman, who sat behind his lawyers, looked up uneasily every time his name was mentioned.  During her closing argument, Eddy fired back. Musk “never cared about the nonprofit structure,” she said. “What he cared about was winning.”  Musk, though, was absent. Despite the judge’s order that he remain available, he flew to China with President Trump. Did Altman promise to keep OpenAI a nonprofit? During her closing argument, Eddy argued that no testimony or evidence showed any conditions on Musk’s donations, or any promises made by Altman and Brockman to keep the company a nonprofit. “No commitments or promises were made. No restrictions were placed on Mr. Musk’s donations,” she said. Eddy added that it was evident Musk wasn’t truly committed to keeping OpenAI a nonprofit. She noted that in 2017, he tried to create a for-profit subsidiary and fought a bitter battle with Altman and Brockman to have control over it. “I was not opposed to there being a small for-profit that provides funding to the nonprofit,” Musk told the jury earlier in the trial, “as long as the tail didn’t wag the dog.”  Eddy then argued that Musk sued too late, filing in 2024 after the statutes of limitations on his claims ran out. In 2019, OpenAI created a for-profit subsidiary, under which employees and investors received a capped return on their investment.  But Musk testified that he discovered OpenAI had abandoned its nonprofit mission only in 2022, when Microsoft was preparing to invest $10 billion in OpenAI—a deal that closed in 2023. “I was disturbed to see OpenAI with a $20B valuation,” he texted Altman after reading the news. “This is a bait and switch.” Musk told the jury that the $20 billion valuation made him realize “the for-profit is the tail wagging the dog.”  “The 2023 deal was different,” Molo hammered home during his closing argument. Is OpenAI still a nonprofit committed to its mission? A central question raised in the last week of trial was whether OpenAI remains a nonprofit

Musk v. Altman week 3: Elon Musk and Sam Altman traded blows over each other’s credibility. Now the jury will pick a side. Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

NVIDIA Introduces SANA-WM: A 2.6B-Parameter Open-Source World Model That Generates Minute-Scale 720p Video on a Single GPU

World models (systems that synthesize realistic video sequences from an initial image and a set of actions) are becoming central to embodied AI, simulation, and robotics research. The core challenge is scaling these systems to generate minute-long, high-resolution video without requiring prohibitively large clusters for both training and inference. Most competitive open-source baselines either require multi-GPU inference or sacrifice resolution to stay within compute budgets. NVIDIA’s SANA-WM directly targets these bottlenecks. Built on the SANA-Video codebase and available through the NVlabs/Sana GitHub repository, it is a 2.6B-parameter Diffusion Transformer (DiT) trained natively for one-minute generation at 720p with metric-scale 6-DoF camera control. It supports three single-GPU inference variants: a bidirectional generator for high-quality offline synthesis, a chunk-causal autoregressive generator for sequential rollout, and a few-step distilled autoregressive generator for faster deployment. The distilled variant denoises a 60-second 720p clip in 34 seconds on a single RTX 5090 with NVFP4 quantization. https://arxiv.org/pdf/2605.15178 The Architecture: Four Core Design Decisions 1. Hybrid Linear Attention with Gated DeltaNet (GDN) Standard softmax attention has memory and compute complexity that grows quadratically with sequence length — a serious problem when generating 961 latent frames for a 60-second video at 720p. SANA-Video, the predecessor, used cumulative ReLU-based linear attention, which maintains a constant-size recurrent state. However, this has no decay mechanism: all past frames accumulate with equal weight, causing drift over minute-scale sequences. SANA-WM replaces most attention blocks with frame-wise Gated DeltaNet (GDN). Unlike token-wise GDN used in language models, SANA-WM’s frame-wise variant processes one entire latent frame per recurrent step. The GDN update rule incorporates a decay gate γ (which down-weights stale past frames) and a delta-rule correction (which updates only the residual between the target value and the current state prediction), keeping the recurrent state at a constant D×D size regardless of video length. To stabilize training, the research team introduces an algebraic key-scaling approach: keys are scaled by 1/√(D·S), where D is the head dimension and S is the number of spatial tokens per frame. This ensures the spectral norm of the transition matrix remains bounded and eliminates the NaN divergence events observed with standard L2 key normalization (1/√D) or no scaling at all, both of which triggered NaN events at steps 16 and 1, respectively. The final backbone interleaves 15 frame-wise GDN blocks with 5 softmax attention blocks (at layers 3, 7, 11, 15, and 19) across 20 total transformer blocks. The softmax blocks provide exact long-range recall where GDN’s recurrence alone is insufficient. 2. Dual-Branch Camera Control Camera-controlled world modeling requires the model to faithfully follow a continuous 6-DoF trajectory, not just align with a text description of motion. SANA-WM uses two complementary branches that operate at different temporal rates: Coarse branch (UCPE attention): Operates at the latent-frame rate. For each latent token, it computes a ray-local camera basis from the camera-to-world pose and intrinsics, then applies a Unified Camera Positional Encoding (UCPE) to the geometric channels of each attention head. This captures global trajectory structure across the full sequence. Fine branch (Plücker mixing): Addresses a compression mismatch. Each latent token summarizes eight raw frames, each with its own distinct camera pose. The fine branch computes pixel-wise Plücker raymaps (a 6D representation: ray direction d and moment o×d) from all eight raw frames within one VAE temporal stride, packs them into a 48-channel tensor, and injects this embedding after each self-attention output via a zero-initialized projection. This restores intra-stride camera motion that the coarse branch cannot see at latent-frame resolution. Ablations on OmniWorld show that neither branch alone matches the dual approach: UCPE-only achieves a Camera Motion Consistency (CamMC) of 0.2453, while UCPE + Plücker mixing reaches 0.2047. 3. Two-Stage Generation Pipeline Stage-1 SANA-WM outputs, while spatiotemporally consistent, can contain structural artifacts over long sequences. A second-stage refiner, initialized from the 17B LTX-2 model with rank-384 LoRA adapters fine-tuned on paired synthetic and real video data, corrects these artifacts. It uses truncated-σ flow matching: stage-1 latents are perturbed with a large starting noise (σ_start = 0.9), and the refiner learns to map this noisy input toward the high-fidelity target. Only three Euler denoising steps are needed at inference. The refiner reduces long-horizon visual drift (ΔIQ) from 3.79 to 1.17 on the Simple-Trajectory split, and from 3.09 to 0.31 on the Hard-Trajectory split. 4. Robust Data Annotation Pipeline Training camera-controlled video generation requires metric-scale 6-DoF pose annotations, the information not available in standard video datasets. The research team modified VIPE (a camera-pose annotation engine) by replacing its depth backend with Pi3X (for long-sequence-consistent depth) fused with MoGe-2 (for accurate per-frame metric scale). They also extended the bundle adjustment stage to treat focal lengths and principal points as per-frame variables rather than shared global intrinsics, enabling more robust annotation on internet video with varying focal lengths. The resulting pipeline processes seven training corpus entries drawn from multiple open-source sources: SpatialVID-HQ (real, 10s clips), DL3DV real clips (10s), DL3DV GS Refined synthetic clips (60s, rendered via 3D Gaussian Splatting), OmniWorld (synthetic, 60s), Sekai Game (synthetic, 60s), Sekai Walking-HQ (real, 60s), and MiraData (real, 60s). This yields a total of 212,975 clips with metric-scale pose annotations. The LTX2-VAE used for compression is 2.0× smaller than ST-DC-AE and 8.0× smaller than Wan2.1-VAE, which directly improves training and inference efficiency. For DL3DV, which contains static 3D scene captures rather than native one-minute videos, the research team fit one FCGS 3D Gaussian Splatting reconstruction per scene, designed diverse one-minute camera paths, rendered long videos with known intrinsics and extrinsics, and then refined the rendered outputs with DiFix3D to reduce splatting artifacts. Training Strategy and Infrastructure SANA-WM’s compute involves two phases on 64 H100 GPUs. First, before DiT training, the team adapts the LTX2 VAE to the SANA-Video SFT training data in approximately 50K steps, taking roughly 3.5 days. The main DiT training then follows a four-stage progressive schedule lasting approximately 15 days: Stage 1 (~2.75 days): Adapt the pre-trained SANA-Video model to the frame-wise GDN architecture on short (5s) video clips. This replaces cumulative linear attention with the

NVIDIA Introduces SANA-WM: A 2.6B-Parameter Open-Source World Model That Generates Minute-Scale 720p Video on a Single GPU Leggi l'articolo »

AI, Committee, Notizie, Uncategorized

Supertone Releases Supertonic v3: On-Device Text-to-Speech Model with 31-Language Support, Fewer Reading Failures, and Expression Tags

Supertone released Supertonic 3, the third generation of its on-device, ONNX-based text-to-speech system. Supertonic 3 ships with 31-language support, improved reading accuracy, fewer repeat and skip failures, and v2-compatible public ONNX assets. It is Lightning Fast, On-Device, Multilingual and Accurate TTS. What Changed from v2 to v3 Compared with Supertonic 2, Supertonic 3 reduces repeat and skip failures, improves speaker similarity across the shared-language set, and expands language coverage from 5 to 31 languages. Version 2 supported English, Korean, Spanish, Portuguese, and French. Version 3 adds Japanese, Arabic, Bulgarian, Czech, Danish, German, Greek, Estonian, Finnish, Croatian, Hungarian, Indonesian, Italian, Lithuanian, Latvian, Dutch, Polish, Romanian, Russian, Slovak, Slovenian, Swedish, Turkish, Ukrainian, and Vietnamese — 31 total ISO language codes. There is also a special na fallback for text whose language is unknown or outside the supported set. The model grows modestly to accommodate the added languages. At about 99M parameters across the public ONNX assets, Supertonic 3 is much smaller than 0.7B to 2B class open TTS systems. The smaller model size is a practical advantage for download size, startup time, and on-device inference. The update also brings the total disk footprint of the public ONNX assets to 404 MB. Additionally, Supertone recently launched the Voice Builder, allowing developers to create custom, edge-native TTS models from their own voice recordings. Expressive Tags One new capability in v3 that wasn’t present in v2 is expressive tag support. Supertonic 3 supports simple expression tags such as <laugh>, <breath>, and <sigh>. These let you embed prosodic cues directly into input text without a separate preprocessing step or a separate model for expressiveness. For engineers building voice interfaces or accessibility tools, this means you can specify breathing pauses or laughter inline in your text payload. Architecture and Runtime The underlying architecture carries over from prior versions: a speech autoencoder that encodes waveforms into continuous latent representations, a flow-matching based text-to-latent module that maps text to audio features, and a duration predictor that controls natural timing. Flow matching is a generative modeling technique that learns a vector field to transform a simple distribution into a target distribution — it samples faster than diffusion models at low step counts, which is why Supertonic can produce usable output in just 2 inference steps. To further refine output, v3 integrates Length-Aware Rotary Position Embedding (LARoPE) for superior text-speech alignment and utilizes a Self-Purifying Flow Matching technique during training to remain robust against noisy data labels. On runtime efficiency, Supertonic 3 runs fast on CPU, even compared with larger baselines measured on A100 GPU, and uses substantially less memory. It does not require a GPU, which makes local, browser, and edge deployment much easier. Reading Accuracy Across measured languages, Supertonic 3 stays within a competitive WER/CER range against much larger open TTS models such as VoxCPM2, while preserving a lightweight on-device deployment path. WER (Word Error Rate) and CER (Character Error Rate) are standard TTS readability metrics: you synthesize a passage, run ASR over the output, and compare the transcription to the original text. CER is used for languages without clear word boundaries; the others use WER. The system’s efficiency is best demonstrated on extreme edge hardware; it achieves an average RTF of 0.3x on an Onyx Boox Go 6 (an E-ink e-reader) in airplane mode. Furthermore, the ecosystem has expanded to include Flutter (with macOS support), .NET 9, and Go, while the web implementation leverages onnxruntime-web for pure client-side execution. Text Normalization A differentiating property carried forward from v2 is built-in text normalization. Supertonic handles complex surface forms — financial expressions like $5.2M, phone numbers with area codes and extensions like (212) 555-0142 ext. 402, time and date formats like 4:45 PM on Wed, Apr 3, 2024, and technical units like 2.3h and 30kph — without any preprocessing pipeline or phonetic annotations. The financial expression “$5.2M” must read as “five point two million dollars,” and “$450K” as “four hundred fifty thousand dollars.” All four competing systems failed this. The technical unit “2.3h” must read as “two point three hours” and “30kph” as “thirty kilometers per hour.” All four competitors also failed this category. The competing systems evaluated include ElevenLabs Flash v2.5, OpenAI TTS-1, Gemini 2.5 Flash TTS, and Microsoft. https://github.com/supertone-inc/supertonic Getting Started The Python SDK install is pip install supertonic. On first run, the SDK downloads the model assets from Hugging Face automatically. A minimal example: Copy CodeCopiedUse a different Browser from supertonic import TTS tts = TTS(auto_download=True) style = tts.get_voice_style(voice_name=”M1″) text = “A gentle breeze moved through the open window while everyone listened to the story.” wav, duration = tts.synthesize(text, voice_style=style, lang=”en”) tts.save_audio(wav, “output.wav”) print(f”Generated {duration:.2f}s of audio”) Marktechpost’s Visual Explainer Supertonic 3 — Developer Guide 1 / 7 Overview Supertonic 3: On-Device TTS,Now in 31 Languages Supertonic 3 is a lightweight, open-weight text-to-speech system by Supertone Inc. It runs entirely via ONNX Runtime on your device — no cloud, no API call, no data leaving your machine. v3 expands from 5 to 31 languages, adds expressive tags, reduces reading failures, and stays compatible with the v2 ONNX interface. 31 Languages ~99M Parameters 404 MB ONNX Assets MIT Code License What’s New in v3 Four Core Improvements Over Supertonic 2 Version 3 is a focused upgrade — same inference contract, meaningfully better output. 31 languages — Expanded from the 5-language v2 release (en, ko, es, pt, fr). Now includes Japanese, Arabic, German, Hindi, Russian, Turkish, Vietnamese, and 20 more ISO codes, plus a special na fallback for unknown languages. More stable reading — Fewer repeat and skip failures, especially on short and long utterances. This was a known limitation in v2 that v3 directly addresses. Expression tags — Supports <laugh>, <breath>, and <sigh> inline in text, without any separate preprocessing or external model. Higher speaker similarity — Improved similarity across the shared-language set compared with Supertonic 2. Voices are more consistent across languages. Installation Get Running in Under a Minute Install the Python SDK via pip. On first run, model assets are downloaded automatically from Hugging Face

Supertone Releases Supertonic v3: On-Device Text-to-Speech Model with 31-Language Support, Fewer Reading Failures, and Expression Tags Leggi l'articolo »

We use cookies to improve your experience and performance on our website. You can learn more at Politica sulla privacy and manage your privacy settings by clicking Settings.

Privacy Preferences

You can choose your cookie settings by turning on/off each type of cookie as you wish, except for essential cookies.

Allow All
Manage Consent Preferences
  • Always Active

Save
it_IT