YouZum

Committee

AI, Committee, 新闻, Uncategorized

A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation

In this tutorial, we work with MSEB, the Massive Sound Embedding Benchmark from Google Research, and approach it from the perspective of what a leaderboard number actually means: the evaluator surface. We install the package and map its three layers, then write two deliberately different encoders against the framework’s own abstract base class: one that measures loudness over time and one that measures timbre, and encode a small synthetic corpus we generate in the notebook so nothing has to be downloaded. We drive the classification, clustering, retrieval, and segmentation evaluators over those embeddings, call the metric functions directly to see what each one rewards, and finish by assembling the TaskMetadata a real submission carries. The result is a comparison in which the two encoders trade places depending on which evaluator is asked, which is the argument for a multi-task benchmark made in numbers rather than in prose. Copy CodeCopiedUse a different Browser import os import sys import json import math import traceback import subprocess import numpy as np RESULTS = {} BENCH = {} def banner(title): print(“n” + “=” * 78) print(title) print(“=” * 78) def section(name): def wrap(fn): def run(*a, **kw): banner(name) try: out = fn(*a, **kw) RESULTS[name] = out if isinstance(out, str) else “ok” return out except Exception as e: RESULTS[name] = f”SKIPPED / FAILED -> {type(e).__name__}: {e}” print(f”n[!] {name} did not complete: {type(e).__name__}: {e}”) traceback.print_exc(limit=3) return None return run return wrap banner(“0. Install MSEB and map the three layers we will use”) subprocess.run([sys.executable, “-m”, “pip”, “install”, “-q”, “mseb==0.1.0″], check=True) import mseb from mseb import types, encoder as encoder_lib, evaluator as evaluator_lib, metrics from mseb.evaluators import ( classification_evaluator, clustering_evaluator, retrieval_evaluator, segmentation_evaluator, ) print(f” mseb {mseb.__version__} | Python {sys.version.split()[0]} | numpy {np.__version__}”) print(“n MSEB is three layers, and a benchmark run walks down them:”) print(” types -> Sound, SoundEmbedding, Score, TaskMetadata: the shapes every task speaks”) print(” encoder -> MultiModalEncoder: the contract YOUR model implements”) print(” evaluators -> classification, clustering, retrieval, reranking, transcription, segmentation, …”) print(“n evaluator entry points we will drive:”) for module, cls in [(classification_evaluator, “ClassificationEvaluator”), (clustering_evaluator, “ClusteringEvaluator”), (retrieval_evaluator, “RetrievalEvaluator”), (segmentation_evaluator, “SegmentationEvaluator”)]: print(f” {module.__name__.split(‘.’)[-1]:28s} {cls}”) print(“n Everything below runs on CPU with no dataset download: we synthesise the audio.”) We install mseb and import the three layers that a benchmark run walks down. The types module holds the shapes every task speaks, Sound, SoundEmbedding, Score and TaskMetadata; the encoder module holds MultiModalEncoder, the contract our own model implements; and the evaluators package holds one module per task family. We import only the four evaluators this notebook drives, because the classification, clustering, retrieval, and segmentation modules depend on nothing heavier than NumPy and scikit-learn. In contrast, the reranking and transcription evaluators pull in Whisper and the task runner pulls in TensorFlow and apache-beam. Everything below therefore runs on a free CPU runtime with no dataset download and no accelerator. Copy CodeCopiedUse a different Browser SR = 16000 @section(“1. The type contract: Sound, SoundEmbedding, Score”) def type_contract(): t = np.arange(SR) / SR waveform = (0.5 * np.sin(2 * np.pi * 440 * t)).astype(np.float32) sound = types.Sound( waveform=waveform, context=types.SoundContextParams(id=”demo_000″, sample_rate=SR, length=len(waveform), language=”en_us”, text=”a 440 Hz tone”), ) print(f” Sound id={sound.context.id!r} {sound.waveform.shape} @ {sound.context.sample_rate} Hz” f” -> {sound.size_bytes:,} bytes”) embedding = types.SoundEmbedding( embedding=np.zeros((1, 16), dtype=np.float32), # (N, D): one utterance-level vector timestamps=np.array([[0.0, 1.0]], dtype=np.float32), # (M, 2): [start, end] in seconds context=sound.context, encoding_stats=types.EncodingStats(input_size_bytes=sound.size_bytes, embedding_size_bytes=16 * 4), ) print(f” SoundEmbedding embedding{embedding.embedding.shape} timestamps{embedding.timestamps.shape}” f” -> {embedding.size_bytes} bytes”) print(f” compression_ratio = {embedding.encoding_stats.compression_ratio:.5f}” f” ({1 / embedding.encoding_stats.compression_ratio:,.0f}x smaller than the audio)”) print(” N embeddings and M timestamps: M == N is frame-aligned, M == 1 is utterance-level.”) print(” `embedding` may also hold N strings instead of vectors – step 8 uses exactly that.”) score = types.Score(metric=”Accuracy”, description=”Overall classification accuracy”, value=0.875, min=0.0, max=1.0) print(f”n Score {score.metric}={score.value} in [{score.min}, {score.max}] :: {score.description}”) for bad, why in [(dict(metric=””, description=”d”, value=0.5, min=0.0, max=1.0), “empty metric name”), (dict(metric=”m”, description=”d”, value=0.5, min=1.0, max=0.0), “min > max”)]: try: types.Score(**bad) except Exception as e: print(f” rejected at construction ({why}): {type(e).__name__}: {e}”) return f”Sound {sound.size_bytes:,} B -> embedding {embedding.size_bytes} B” type_contract() We start with the type contract, because every other layer is expressed in it. A Sound carries a waveform, along with SoundContextParams, the identifier, sample rate, length, language, and optional transcript, which follow the audio through the whole pipeline. A SoundEmbedding carries an array of N embeddings and an array of M timestamp pairs, and the relation between N and M is the benchmark’s vocabulary: M equal to N means one vector per frame, while M equal to one means a single utterance-level vector, which is what our encoders produce. EncodingStats records the input and embedding sizes and exposes compression_ratio, here a thousandfold reduction from audio to vector. A Score is a metric name, a value and its bounds, and it validates itself at construction, rejecting an empty metric name or a minimum above its maximum, so a malformed number cannot reach a leaderboard. The embedding field also accepts N strings instead of N vectors, which is the door that step 8 walks through. Copy CodeCopiedUse a different Browser class EnergyEnvelopeEncoder(encoder_lib.MultiModalEncoder): “””Baseline: average energy in `n_bins` equal time slices. Loud/quiet, nothing about timbre.””” def __init__(self, n_bins: int = 16): super().__init__() self.n_bins = n_bins def _setup(self): self._ready = True # a real encoder loads weights here def _check_input_types(self, batch): for item in batch: if not isinstance(item, types.Sound): raise ValueError(f”{type(self).__name__} takes types.Sound, got {type(item).__name__}”) def _encode(self, batch) -> list[types.SoundEmbedding]: out = [] for sound in batch: slices = np.array_split(sound.waveform.astype(np.float32), self.n_bins) vec = np.array([[float(np.sqrt(np.mean(s ** 2) + 1e-12)) for s in slices]], dtype=np.float32) vec /= np.linalg.norm(vec) + 1e-9 out.append(types.SoundEmbedding( embedding=vec, timestamps=np.array([[0.0, sound.context.length / sound.context.sample_rate]], dtype=np.float32), context=sound.context)) return out class SpectralProfileEncoder(encoder_lib.MultiModalEncoder): “””Contender: mean log-magnitude spectrum pooled into `n_bands` bands. Describes timbre.””” def __init__(self, n_bands: int = 16, frame: int = 512): super().__init__() self.n_bands, self.frame = n_bands, frame def _setup(self): self._window = np.hanning(self.frame).astype(np.float32) def _check_input_types(self, batch): for item in batch: if not isinstance(item, types.Sound): raise ValueError(f”{type(self).__name__} takes types.Sound, got {type(item).__name__}”) def _encode(self, batch) -> list[types.SoundEmbedding]: out = [] for sound in batch: w = sound.waveform.astype(np.float32) n_frames =

A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation Read Post »

AI, Committee, 新闻, Uncategorized

End-to-End Multimodal Data Augmentation and Adversarial Robustness Benchmark with AugLy for Images, Text, Audio, and PyTorch

In this tutorial, we build a comprehensive multimodal augmentation and robustness workflow with AugLy for images, text, and audio. We start by addressing modern dependency compatibility issues and generating deterministic synthetic datasets so the experiments remain self-contained and reproducible. We then explore AugLy’s functional and class-based APIs, metadata, and intensity tracking, probabilistic composition, bounding-box-aware transformations, and custom transforms. We extend the workflow into practical robustness experiments by benchmarking perceptual-hash copy detection under image distortions and evaluating text classifiers against adversarial perturbations, Unicode obfuscation, sanitization, and adversarial training. We also integrate audio augmentation, build a queryable metadata warehouse, and connect AugLy transformations directly to PyTorch datasets and DataLoaders, giving us an end-to-end view of augmentation as both a data-generation mechanism and a measurable robustness tool. Copy CodeCopiedUse a different Browser import subprocess, sys, importlib def _sh(cmd): print(f”$ {cmd}”) subprocess.run(cmd, shell=True, check=False, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL) def _need(mod): try: importlib.import_module(mod) return False except ImportError: return True if _need(“augly”): _sh(“apt-get -qq install -y libmagic1 > /dev/null 2>&1″) _sh(f'”{sys.executable}” -m pip install -q –no-deps augly’) _sh(f'”{sys.executable}” -m pip install -q “iopath>=0.1.8” “python-magic>=0.4.22″ ‘ f'”regex>=2021.4.4” “nlpaug==1.1.3″‘) import numpy as np from PIL import Image, ImageDraw, ImageFont, ImageFilter for _name, _builtin in ((“float”, float), (“int”, int), (“bool”, bool)): if not hasattr(np, _name): setattr(np, _name, _builtin) def _size(font, text): left, top, right, bottom = font.getbbox(text) return (right, bottom) if not hasattr(ImageFont.FreeTypeFont, “getsize”): ImageFont.FreeTypeFont.getsize = lambda self, t, *a, **k: _size(self, t) if not hasattr(ImageFont.FreeTypeFont, “getsize_multiline”): def _getsize_multiline(self, text, direction=None, spacing=4, features=None, language=None, stroke_width=0): lines = text.split(“n”) w = max((_size(self, ln)[0] for ln in lines), default=0) h = sum(_size(self, ln)[1] for ln in lines) + spacing * (len(lines) – 1) return (w, h) ImageFont.FreeTypeFont.getsize_multiline = _getsize_multiline import os, io, json, math, random, string, textwrap, unicodedata, warnings from dataclasses import dataclass from typing import Any, Dict, List, Optional, Tuple import matplotlib.pyplot as plt import pandas as pd import augly.image as imaugs import augly.text as textaugs import augly.utils as augutils from augly.image.transforms import BaseTransform as ImageBaseTransform warnings.filterwarnings(“ignore”) pd.set_option(“display.width”, 160) SEED = 1234 random.seed(SEED) np.random.seed(SEED) print(“n” + “=” * 78) print(“AugLy ready. assets at:”, augutils.ASSETS_BASE_DIR) print(“image augs :”, len([f for f in dir(imaugs) if f[0].islower()])) print(“text augs :”, len([f for f in dir(textaugs) if f[0].islower()])) print(“=” * 78 + “n”) def make_image(idx: int, w: int = 320, h: int = 240) -> Tuple[Image.Image, Tuple[int, int, int, int]]: “””Procedurally generated ‘photo’ + a ground-truth bbox in pascal_voc format.””” rng = random.Random(SEED + idx) img = Image.new(“RGB”, (w, h), tuple(rng.randint(20, 90) for _ in range(3))) d = ImageDraw.Draw(img) for _ in range(70): x0, y0 = rng.randint(0, w), rng.randint(0, h) d.line([x0, y0, x0 + rng.randint(-60, 60), y0 + rng.randint(-60, 60)], fill=tuple(rng.randint(60, 160) for _ in range(3)), width=rng.randint(1, 3)) ow, oh = rng.randint(70, 130), rng.randint(60, 110) ox, oy = rng.randint(10, w – ow – 10), rng.randint(10, h – oh – 10) box = (ox, oy, ox + ow, oy + oh) colour = tuple(rng.randint(150, 255) for _ in range(3)) if idx % 3 == 0: d.ellipse(box, fill=colour, outline=(255, 255, 255), width=3) elif idx % 3 == 1: d.rectangle(box, fill=colour, outline=(255, 255, 255), width=3) else: d.polygon([(ox + ow // 2, oy), (ox + ow, oy + oh), (ox, oy + oh)], fill=colour, outline=(255, 255, 255)) return img, box N_IMAGES = 24 IMAGES, BOXES = zip(*[make_image(i) for i in range(N_IMAGES)]) IMAGES, BOXES = list(IMAGES), list(BOXES) DEMO_IMG, DEMO_BOX = IMAGES[0], BOXES[0] def make_text_dataset(n_per_class: int = 260): “””Tiny sentiment corpus built from templates -> learnable but not trivial.””” rng = random.Random(SEED) pos_adj = [“excellent”, “delightful”, “superb”, “charming”, “brilliant”, “flawless”, “wonderful”, “outstanding”, “impressive”, “lovely”] neg_adj = [“terrible”, “awful”, “dreadful”, “disappointing”, “clumsy”, “broken”, “miserable”, “useless”, “painful”, “sloppy”] subj = [“the movie”, “this restaurant”, “the hotel room”, “their support team”, “the new phone”, “the sequel”, “this laptop”, “the delivery service”] tail_p = [“and I would recommend it to anyone”, “worth every rupee”, “I left completely satisfied”, “easily the best of the year”, “it exceeded all my expectations”] tail_n = [“and I want a refund”, “a total waste of money”, “I left extremely frustrated”, “easily the worst of the year”, “it failed every expectation”] rows = [] for _ in range(n_per_class): rows.append((f”{rng.choice(subj)} was {rng.choice(pos_adj)} {rng.choice(tail_p)}”, 1)) rows.append((f”{rng.choice(subj)} was {rng.choice(neg_adj)} {rng.choice(tail_n)}”, 0)) rng.shuffle(rows) return [r[0] for r in rows], [r[1] for r in rows] TEXTS, LABELS = make_text_dataset() DEMO_TEXT = “The quick brown fox jumps over the lazy dog near the river bank” def make_audio(seconds: float = 2.0, sr: int = 16000) -> Tuple[np.ndarray, int]: “””A chirp + harmonics + a little noise = something you can actually hear change.””” t = np.linspace(0, seconds, int(sr * seconds), endpoint=False) f = np.linspace(220, 880, t.size) sig = 0.5 * np.sin(2 * np.pi * f * t) + 0.2 * np.sin(2 * np.pi * 2 * f * t) sig += 0.02 * np.random.RandomState(SEED).randn(t.size) env = np.minimum(1.0, np.minimum(t * 8, (seconds – t) * 8)) return (sig * env).astype(np.float32), sr AUDIO, SR = make_audio() def show_grid(pairs, cols=4, title=””, figsize_scale=2.9): “””pairs: list of (caption, PIL.Image).””” rows = math.ceil(len(pairs) / cols) fig, axes = plt.subplots(rows, cols, figsize=(cols * figsize_scale, rows * figsize_scale)) axes = np.atleast_1d(axes).ravel() for ax, (cap, im) in zip(axes, pairs): ax.imshow(im) ax.set_title(cap, fontsize=8) ax.axis(“off”) for ax in axes[len(pairs):]: ax.axis(“off”) if title: fig.suptitle(title, fontsize=13, y=1.0) plt.tight_layout() plt.show() def as_str(out) -> str: “””AugLy text augs return str for str input in some transforms, list in others.””” return out[0] if isinstance(out, list) else out print(“n### §2 IMAGE AUGMENTATION + METADATA ” + “#” * 38) functional_result = imaugs.pixelization(DEMO_IMG, ratio=0.25) class_result = imaugs.Pixelization(ratio=0.25, p=1.0)(DEMO_IMG) print(“functional == class:”, np.array_equal(np.array(functional_result), np.array(class_result))) IMAGE_ZOO = { “blur”: lambda im, m: imaugs.blur(im, radius=3.0, metadata=m), “brightness”: lambda im, m: imaugs.brightness(im, factor=1.7, metadata=m), “color_jitter”: lambda im, m: imaugs.color_jitter(im, brightness_factor=1.3, contrast_factor=1.4, saturation_factor=1.6, metadata=m), “crop”: lambda im, m: imaugs.crop(im, x1=.15, y1=.15, x2=.85, y2=.85, metadata=m), “encoding_quality”: lambda im, m: imaugs.encoding_quality(im, quality=8, metadata=m), “grayscale”: lambda im, m: imaugs.grayscale(im, metadata=m), “hflip”: lambda im, m: imaugs.hflip(im, metadata=m), “meme_format”: lambda im, m: imaugs.meme_format(im, text=”TOP TEXT”, caption_height=90, metadata=m), “opacity”: lambda im, m: imaugs.opacity(im, level=0.45, metadata=m), “overlay_emoji”: lambda im, m: imaugs.overlay_emoji(im, opacity=0.9, emoji_size=0.35, metadata=m), “overlay_screenshot”: lambda im, m: imaugs.overlay_onto_screenshot(im, metadata=m), “overlay_stripes”: lambda im, m: imaugs.overlay_stripes(im, line_width=0.4, line_opacity=0.7, metadata=m),

End-to-End Multimodal Data Augmentation and Adversarial Robustness Benchmark with AugLy for Images, Text, Audio, and PyTorch Read Post »

AI, Committee, 新闻, Uncategorized

Exa Launches Agent Ultra: A Subagent Swarm Deep Research API Built for Exhaustive List Building

Exa has released Agent Ultra, the highest effort level of its Exa Agent API. It is built for research that must run to exhaustion: large list building, entity enrichment, and questions that need thousands of sources. Exa team reports that Ultra beats Opus 5.5, GPT-6 Astra, and Perplexity Agent, each at maximum effort, on 4 research benchmarks. Is it deployable? Yes, as a hosted API. Agent Ultra is live today on the Exa API by setting effort: “ultra”. It is not open weights and cannot be self-hosted. What is Exa Agent Ultra? Exa Agent splits a task into subtasks and assigns subagents to research several domains at once. It routes frontier models to steps that need them and faster models where those are enough. Ultra is the mode that spends the most compute. According to the Agent Ultra docs, it runs longer than any other effort to return the most complete results. Ultra runs typically finish complex tasks in about 30 minutes. Very hard tasks can take up to 3 hours. Benchmark Results All figures below are from Exa’s launch post. Competitors ran at their maximum effort setting. Benchmark (metric) Agent Ultra Opus 5.5 GPT-6 Astra Perplexity Agent WANDR (soft recall) 81.4% 72.3% 26.0% 40.1% DeepSearchQA (F1) 93.9% 77.6% 85.3% 89.7% WideSearch (row-level F1) 58.9% 51.6% 54.7% 56.0% Company Find-All (avg. passing entities per task) 2,451 146 113 98 Exa pairs each result with a cost claim: WANDR: +12.6% over Opus 5.5, at half its cost per task. DeepSearchQA: +4.7% over Perplexity, at 46% lower cost per task than GPT-6 Astra. WideSearch: +5.2% over Perplexity, at the lowest cost per task of the 4 systems. Company Find-All: +1579% over Opus 5.5, at the lowest cost per entity found. These gains are relative, not percentage points. On WANDR, the absolute gap to Opus 5.5 is 9.1 points. Understanding These Numbers WANDR is Perplexity’s benchmark of 500 wide and deep data-collection tasks, with an open harness. Exa’s grader shares the upstream evaluation logic. It swaps in Exa as the contents tool, changes transport logic, and uses gpt-6-luna as the judge. Where a vendor had published a result on this harness, Exa reports that figure. Otherwise, Exa ran the benchmark itself. DeepSearchQA is Google DeepMind’s 900-prompt multi-step search benchmark. WideSearch tests broad information gathering. Exa evaluated up to 200 tasks each for WANDR and DeepSearchQA, and 100 each for WideSearch and Company Find-All. Graded task counts vary by provider. All results are vendor-reported and not yet independently reproduced. Where Agent Ultra Fits Exa lists 3 target user groups: Model providers: assemble training data, such as every paper and repo implementing a given technique. Verify criteria like ‘released weights, not just an API’. Financial services: build diligence market maps, run KYC research across filings and court records, and monitor portfolio signals. Go-to-market teams: build account lists and enrich rows with judgment fields, each backed by a cited URL. Ultra can also expand an existing list. Pass the rows you already have, and they are excluded from new results. API, Pricing, and Controls Ultra uses the standard Agent run endpoint. The request supports outputSchema, input.data, and streaming. Copy CodeCopiedUse a different Browser from exa_py import Exa exa = Exa() run = exa.agent.runs.create( query=”Find all companies building browser automation tools in the United States.”, effort=”ultra”, ) run = exa.agent.runs.poll_until_finished(run.id, timeout_ms=3 * 60 * 60 * 1000) print(run.stop_reason) Pricing: metered at standard Agent usage rates, up to a default $20 per run. Runs that finish early cost less. Budget: maxCostDollars accepts $1 to $100. maxDurationSeconds accepts 300 to 10,800 seconds. Stopping: a stop call ends a run early, keeps its results, and bills usage up to that point. Timeouts: SDK polling helpers time out after 1 hour by default, so set a longer timeout or stream events. OpenAI compatibility: on /responses, set reasoning.effort: “ultra” with streaming or background mode. You can test it in the Exa API Playground. Comparison Key Takeaways Agent Ultra is Exa Agent’s highest effort mode, live now via API. It orchestrates parallel subagents and mixes frontier and faster models. Exa reports top scores on WANDR, DeepSearchQA, WideSearch, and Company Find-All. Runs cost up to $20 by default, adjustable from $1 to $100. Typical runs take about 30 minutes, with a 3 hour ceiling. Check out the Technical Details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Exa Launches Agent Ultra: A Subagent Swarm Deep Research API Built for Exhaustive List Building appeared first on MarkTechPost.

Exa Launches Agent Ultra: A Subagent Swarm Deep Research API Built for Exhaustive List Building Read Post »

AI, Committee, 新闻, Uncategorized

Supersonic Labs Releases Julia 1: A 144.3M-Parameter Open Decision Model That Runs on a CPU

Supersonic Labs, a small AI lab from Brazil, has released Julia 1. It is a compact decision model, not a chatbot. You pass it context, a question, and 2 to 20 candidate answers. It picks one and returns a probability for every option. The model has 144.3M parameters and runs on a plain CPU. Is it deployable? Yes. The weights are on Hugging Face under Apache 2.0 and run locally with Python 3.11+ on CPU or a BF16-capable GPU. An ONNX build also runs in the browser via WebGPU. A hosted API is announced but not open yet. What Julia 1 Does Julia 1 handles three decision types through one API: choice: pick one label from 2 to 20 described options (classification, routing). score: return the expected index on an ordered rubric, such as low, medium, high. noul: return the probability that a yes-or-no statement is true. Results come back in the caller’s option order with full softmax probabilities. Caller IDs such as billing are returned unchanged. The model does not generate text. Architecture and Training Budget Julia 1 starts from JHU CLSP’s mmBERT-small, a 140M-parameter multilingual ModernBERT encoder trained on 1,800+ languages. Supersonic Labs kept the encoder and tokenizer, added a decision head, and trained on decision-format examples. The lab states Julia 1 is not a fine-tuned Qwen model. The runtime supports 8,192 combined tokens, but published benchmarks used a 1,024-token limit. Total cloud GPU spend for training and experiments was about R$540 (US$104.08). The FP32 weights occupy 550.5 MiB. The private training pipeline is not released. Julia 2, with the lab’s own foundation architecture, is in development. Benchmark Results The September 24, 2026 evaluation ran on H200 BF16 with strict encoding. The comparison baseline is TypeSafe’s Jev, using reference values from the Jev benchmark protocol, not a new Jev run. Typed Decisions: 73.15% (1,463/2,000) vs 72.70% reference. AG News, 4 labels: 94/100 vs 91% reference. DAIR Emotion, 6 labels: 86/100 vs 48% reference. Banking77, 72 labels: 64/100 vs 87% reference. This is the clear failure. MASSIVE, 18 scenarios: 71.50% macro accuracy across 52 locales; 86.25% pt-PT, 86.75% en-US. The classification pilots use only 100 examples each. A September 25 CPU run reproduced most numbers: 72.55% on Typed Decisions and 60/100 on Banking77 with 3 abstentions. On-Device Latency The lab published per-device measurements. On an Apple M4, one decision per call took a 33.15 ms median. On a Samsung SM-X510 tablet via ONNX Runtime, the median was 203 ms with 393.1 MB peak RSS. On an Intel Core i5-1235U, AG News decisions took a 107.83 ms median. Banking77 took 3,713.54 ms because it narrows 72 labels first. On X, @supersonicai claims Julia 1 classifies 5x faster than Jev on an i5 laptop. Treat that carefully. The Jev pilot measured Jev as a hosted service called from France, so latencies are not like-for-like. Introducing Julia-1:Our first classification model that runs on almost anything. Learn more https://t.co/YJCdEeIBSo pic.twitter.com/3cizsbG9ZB — Supersonic Labs (@supersonicai) September 26, 2026 Interactive Explainer Julia 1 vs Closest Competitors Feature Julia 1 TypeSafe Jev GLiNER2.5 Multi Developer Supersonic Labs TypeSafe AI Fastino Access Open weights Hosted API, early access Open weights License Apache 2.0 Proprietary Apache 2.0 Parameters 144.3M Not disclosed 287M Base encoder mmBERT-small Not disclosed mDeBERTa-v3-base Decision types Choice, score, yes/no Typed structured decisions Classification, NER, relations, records Options per call 2 to 20 (Router for more) Up to 255 Label list per schema Runs locally on CPU Yes No Yes Input price per 1M tokens $0.025 (planned API) $0.042 Free (self-hosted) AG News pilot 94% 91% 70% DAIR Emotion pilot 86% 48% 44% Banking77 pilot 64% 87% 61% Sources: Julia 1 model card, TypeSafe launch post, GLiNER2.5 Multi card, Jev benchmark pilot. Julia 1 pilots ran separately from the Jev and GLiNER runs. Limitations Julia 1 compares the answers you supply. It cannot be counted on for missing facts, algebra, or multi-step calculation. The Router can drop the correct label during narrowing. It is not a drop-in Transformers pipeline, and no Hugging Face inference provider serves it. Supersonic Labs advises evaluating on your own questions and keeping humans in the loop for consequential decisions. Key Takeaways Julia 1 is a 144.3M-parameter, Apache 2.0 decision model that runs on CPU. One API covers choice, ordered score, and yes-or-no decisions over 2 to 20 options. It beat Jev references on 3 of 4 pilots but trailed badly on 72-label Banking77. Median latency hit 33.15 ms per decision on an Apple M4. Training cost about US$104 in cloud GPUs; a $0.025/MTok API is planned. Check out the Model Weights, ONNX/WebGPU build, and Technical details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Supersonic Labs Releases Julia 1: A 144.3M-Parameter Open Decision Model That Runs on a CPU appeared first on MarkTechPost.

Supersonic Labs Releases Julia 1: A 144.3M-Parameter Open Decision Model That Runs on a CPU Read Post »

AI, Committee, 新闻, Uncategorized

Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English

Sarvam AI has released Saaras V4, the newest generation of its speech recognition model. It covers all 22 scheduled Indian languages plus English, now including global English accents. Sarvam reports state-of-the-art accuracy across all 22 languages. Is it deployable? Yes, through Sarvam’s API today, using model=”saaras:v4″. Weights are not public, and Sarvam’s SageMaker self-hosting docs currently cover Saaras v3 only. What is Inside Saaras V4 Saaras V4 is an encoder-decoder system. An audio encoder converts the waveform into embeddings that carry phonetic and acoustic detail. A temporal-downsampling adapter then shortens that sequence and projects it into the language model’s embedding space. This keeps long recordings inside the decoder’s context budget. The decoder is Sarvam-3B, a 3B-parameter hybrid state-space language model trained from scratch in-house. It reads the audio features alongside a text prompt. It then emits the transcript autoregressively, feeding each token back as input for the next. Benchmark Results English: Sarvam evaluated 7 English datasets. Six come from Hugging Face’s Open ASR Leaderboard: AMI, GigaSpeech, LibriSpeech clean, LibriSpeech other, SPGISpeech and VoxPopuli. The seventh is AI4Bharat’s Indian-accented Svarah. Scoring follows the leaderboard’s normalization code. Saaras V4 posts the lowest average WER among the models Sarvam benchmarked. Indic: On Vistaar, Sarvam reports results across 10 Indian languages using both WER and LLM-WER. LLM-WER adds a semantic check. It separates real meaning errors from harmless spelling or formatting variants common in Indic scripts. Noisy audio: On Kathbath Noisy, measured with LLM-WER, Sarvam says Saaras V4’s error rate is under half that of Deepgram Nova-3 and GPT-4o Transcribe. The set includes compressed, clipped and background-heavy recordings. Language ID: On verified IndicVoices utterances, language identification error is 2.9% across the top 10 Indian languages. It is 5.22% across all 22. It is important to note that all numbers above are vendor-reported. Independent reproduction has not been published yet. 5 Output Modes From 1 Model The same audio can return 5 representations, selected through the mode parameter: transcribe (default): native script with numbers and dates normalized. verbatim: every word as spoken, fillers and spoken numbers kept. codemix: native script, with English words left in English. translit: the full utterance in Latin script. translate: an English translation with numbers normalized. Sarvam’s argument is simple. Handling these inside the model removes post-processing steps that can compound errors. Keyterm Prompting Keyterm prompting is new in V4 and works only with saaras:v4. You pass a JSON list under keyterms, with up to 50 terms of 64 characters each. Keyterms bias recognition; they do not guarantee output. Use codemix mode when a brand such as PhonePe must stay in Latin script. On IndicContextEval (paper, Interspeech 2026), Saaras V4 reports 16.03% WER in the L5 keyword-prompting setting. Sarvam says that is the lowest score on the benchmark. Streaming, Long Audio and Pricing Streaming: WebSocket with partial results and time to first token below 150 ms. REST: synchronous transcription for clips up to 30 seconds. Batch: asynchronous jobs up to 2 hours per file, with optional speaker diarization. SDKs: Python 3.9+ and Node.js 18+, plus LiveKit Agents, Pipecat and Vercel AI SDK integrations. Price: Sarvam lists speech-to-text at ₹30 per hour for real-time, streaming and batch, and ₹45 per hour with diarization. Saaras v3 stays the default model. V4 uses the same request shape, so switching is a 1-line change. Saaras V4 vs Closest Competitors These are the 3 systems Sarvam benchmarked against. Figures come from each vendor’s public docs and pricing pages, checked on September 26, 2026. Feature Sarvam Saaras V4 Deepgram Nova-3 ElevenLabs Scribe v2 OpenAI GPT-4o Transcribe Indian scheduled languages (of 22) 22 11 14 Not listed per language Total languages 23 (22 Indian + English) 45+ 90+ Multilingual Keyterm biasing Up to 50 terms Yes, paid add-on Up to 1,000 (batch), 50 (realtime), paid add-on Free-text prompt Built-in output modes 5 (transcribe, verbatim, codemix, translit, translate) Transcript plus Smart Formatting Verbatim or no_verbatim Transcript Real-time streaming WebSocket, under 150 ms TTFT (vendor claim) Yes (WebSocket) Scribe v2 Realtime, about 150 ms File streaming; live via Realtime API Speaker diarization Batch API Yes Up to 32 speakers Separate gpt-4o-transcribe-diarize model List price ₹30/hour $0.0052/min (multilingual, pre-recorded) $0.22/hour (batch) ~$0.006/min Self-hosting Not for V4 yet (v3 on SageMaker) Yes Cloud API Cloud API Key Takeaways Saaras V4 covers all 22 Indian languages plus global English in 1 model. A 3B hybrid state-space decoder, trained from scratch, sits behind an audio encoder. Keyterm prompting accepts up to 50 terms and scored 16.03% WER on IndicContextEval L5. 5 output modes and sub-150 ms streaming TTFT come from the same model. API-only today at ₹30 per hour; self-hosting docs still cover v3. Check out the Technical Details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English appeared first on MarkTechPost.

Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English Read Post »

AI, Committee, 新闻, Uncategorized

The Download: the Pentagon’s AI-powered lie detector and young organ limits

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. The Pentagon wants $30 million to build an AI-powered lie detector The US government wants to spend $30.3 million over the next five years on an improved lie detector, according to a Department of Defense budget request. The program, called “Polygraph+” or “Polygraph Next,” will focus on scoring algorithms that use AI and machine learning, as well as a technique called “standoff sensing,” which can take physiological readings from a person without attaching a device to them. The project aims to improve the accuracy and reliability of polygraph assessments. But it could just be the latest in a long line of failed attempts to use technology to detect lies. Here’s how the program could work—and why experts say AI-enabled lie detection is “the worst of both worlds.” —Amit Katwala Young organs may not be a fountain of youth for recipients Around this time last year, a hot mic caught Vladimir Putin and Xi Jinping discussing the possibility of living forever. “With the developments of biotechnology, human organs can be continuously transplanted, and people can live younger and younger, and even achieve immortality,” Putin reportedly said. He seemed to be referring to the “replacement” theory of longevity, supported by experiments that involved physically stitching young mice to old ones. Unfortunately for Putin, new research on transplanted hearts pours a little cold water on this idea. But it also offers useful insights for organ transplantation. Find out what the research reveals about the limits of rejuvenation. —Jessica Hamzelou This story is from The Checkup, our weekly biotech newsletter. Sign up to receive it in your inbox every Thursday. Roundtables: the deadly failures of the virtual border wall The US has spent billions building a “virtual wall” of surveillance towers along its southern border, promising to detect and apprehend border crossers and save lives. But an MIT Technology Review investigation found more than a thousand people died within the advertised range of the towers without getting caught. On Monday September 28, our editor-in-chief Mat Honan, senior AI reporter James O’Donnell and senior reporter for features and investigations Eileen Guo will join a subscriber-only conversation about the investigation. They’ll examine the failures of border surveillance technology and uncover the stories of the people who die in the borderlands. Register now to attend on Monday, September 28 at 7:00pm BST / 2:00pm EDT / 11:00am PDT. Want to join the conversation? Subscribe to MIT Technology Review for exclusive access to all our Roundtables. Read the full MIT Technology Review investigation here. The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 ChatGPT helped the Tumbler Ridge shooter focus on attack tacticsA new investigation found it also advised on evading safeguards. (Mother Jones)+ British Columbia is suing OpenAI over its failure to alert police. (BBC)+ And wants OpenAI to pay for a replacement school. (Ars Technica)+ Do AI chatbots cause delusions or amplify them? (MIT Technology Review) 2 Ukraine just used drones to drop military robots behind Russian linesThe “world-first” assault sent the devices into enemy territory. (Ars Technica)+ The robots can attack, scout and clear routes for troops. (Business Insider)+ Europe has a drone-filled vision for future wars. (MIT Technology Review) 3 Google is about to launch AI chips into spaceThe satellite will run simple AI queries from orbit. (Reuters $)+ The project is part of plans to put AI data centers in space. (Gizmodo)+ Here’s how we could put data centers in space. (MIT Technology Review) 4 A new DNA test can diagnose brain tumors in hoursThe technology analyzes a tumor’s DNA to identify its type. (BBC) 5 Tesla’s Semi is finally ready to hit the roadThe long-delayed electric truck is reaching customers this week.(Verge)+ It could be a big deal for electric trucking. (MIT Technology Review) 6 Meta has gained an early lead over OpenAI in the AI device marketThe Muse Charm is expected to ship before OpenAI’s hardware. (CNBC)+ Meta promises to put privacy at the center of the products. (Axios) 7 A mathematician has discovered a rare eight-faced shapeHe argues that the strange three-holed object can exist in 3D. (New Scientist $) 8 AI agents are flooding researchers with collaboration requestsSome bots are asking for data, money, and research partnerships. (Nature) 9 Venus may have swallowed its own moonThe planet’s slow rotation could have dragged the moon to its doom. (Wired $) 10 Mark Zuckerberg has finally explained his fashion glow-upHis makeover is intertwined with Meta’s wearables push. (NYT $) Quote of the day “They’ve got to inflict an enormous amount of pain and suffering on you so that they can save you. And so I think AI’s kind of like that.”  —Jensen Huang compares AI’s need for fossil fuels to surgery in an interview on The Ezra Klein Show. One more thing This scientist rewarmed and studied pieces of his friend’s cryopreserved brain L. Stephen Coles’s brain sits in a vat at a storage facility in Arizona. It has been held there at a temperature of around −146°C for more than a decade, largely undisturbed. Before he died in 2014, Coles had the brain frozen with an ambitious goal in mind: reanimation. His friend, cryobiologist Greg Fahy, believes it could be revived one day. But other experts are less optimistic.  Still, Fahy’s research could lead to new ways to study the brain. And using cryopreservation for organ transplantation is becoming a viable reality.  Read the full story to find out what the future holds for the technology. —Jessica Hamzelou We can still have nice things A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.) + A French village recently hosted its annual pig-squealing championship.+ Things get deliciously unattractive at the Iowa State Fair’s annual ugliest cake contest.+ Greater Victoria is home to more than 1,000 Little Free Libraries, each

The Download: the Pentagon’s AI-powered lie detector and young organ limits Read Post »

AI, Committee, 新闻, Uncategorized

Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation

Perplexity Research published a new post-training study. It trains a model inside Perplexity Computer on real user sessions, including failed ones. The method pairs rejection sampling fine-tuning with hint-guided self-distillation. In a live A/B test, tool-call failures fell from 2.24% to 1.77% between 2 trained checkpoints. Perplexity team reports this as a statistically significant 21.2% relative reduction. Is it deployable? Not directly. Perplexity has not released the post-trained weights or training code. The model runs only as a model option inside Perplexity Computer. The base model, GLM 5.2, is openly available on Hugging Face. Why Outcome-Only Filtering Falls Short Standard rejection sampling fine-tuning (RFT) judges each session and imitates only the successful ones. A successful outcome does not mean every step was correct. An agent can recover from a bad tool call and still deliver the right answer. Imitating that full trajectory can reinforce the error. Discarding failed sessions also throws away clear evidence of avoidable mistakes. Imitate, Correct, or Keep as Context Perplexity team separates 2 decisions: which sessions hold behavior worth imitating, and which turns hold mistakes worth correcting. Each assistant turn gets 1 of 3 treatments: Imitate: non-error turns in successful sessions receive cross-entropy (CE) loss. Correct: error turns with a validated hint receive Kullback-Leibler (KL) divergence loss, in any session. Keep as context: remaining turns stay in the input but receive no loss. Successful sessions can supply both imitation and correction targets. Unsuccessful sessions supply only correction targets. How a Hint Becomes a Training Signal A hint is a short corrective instruction grounded in information the model already had. In one example, a search call set recency_filter to ‘year.’ The schema allowed only ‘day,’ ‘week,’ or ‘month.’ The hint names the failed call, includes the validation error, and suggests an allowed value or omitting the optional field. The corrective part uses On-Policy Self-Distillation (OPSD). The trainer runs the same GLM 5.2 checkpoint twice on the recorded turn. The teacher pass sees the hint; the student pass does not. Both use teacher forcing, so no replacement answer is generated. The teacher’s next-token probabilities are detached and act as a soft target through forward KL. The combined loss is (CE + λ × KL), divided by the number of imitated tokens. Setting λ to 0 recovers standard SFT. The CE term matters. Correction-only training can let teacher and student agree by ignoring context. Tracing Complaints to the Real Mistake The pipeline draws from training-eligible Computer sessions served by GLM 5.2. Sessions with personally identifiable information and users who opted out are excluded. An LLM judge keeps tasks rated 4 or 5 on a 5-point difficulty scale. Two LLM judges must both approve the final delivery for a session to count as successful. For user feedback, threes LLM judges locate the responsible turn, and at least 2 must agree. This is important because the last assistant turn before a complaint is the root cause only about half the time. Each hint is also checked against information available before the mistake. That check reduces hindsight bias. One example: a user asked for their ‘w3’ on Paychex. The model assumed a W-2 typo and searched for the wrong form. The hint targets that earlier interpretation, not just the final answer. Interactive Explainer What the Evaluations Show Hints work before training: On 985 held-out tool-error turns, the unchanged base model avoided the original failure in 93.7% of cases with hints, up from 75.1%. The share taking the corrected action rose from 60.6% to 82.3%. On user-feedback turns, fixed or on-track rates rose from 40.0% to 75.0% for explicit evidence. For inferred intent, they rose from 32.5% to 80.0%. Offline tool errors fell: Recorded tool-error rates were 2.79% for stock GLM 5.2 and 1.35% for RFT only. The RFT plus OPSD checkpoint reached 0.87%. Perplexity notes these checkpoints used different training data, so this is not a matched ablation. Task-level benchmark results on suites like BrowseComp and SpreadsheetBench were mixed. Live results are narrower: Each A/B test used about 100,000 users per condition. An early checkpoint versus stock GLM 5.2 showed 2.82% versus 2.94% failures, which was not significant. The later checkpoint comparison produced the significant 21.2% drop, without hints at inference. Strong dissatisfaction moved from 2.58% to 2.54%, which was also not significant. Perplexity did not compare the later checkpoint directly against stock GLM 5.2 online. Key Takeaways Perplexity learns from failed sessions, not just successful ones. Validated hints turn avoidable mistakes into KL correction targets. 1 model acts as teacher (with hint) and student (without). Live tool-call failures fell from 2.24% to 1.77%. User dissatisfaction showed no significant change. Check out the Technical Details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation appeared first on MarkTechPost.

Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation Read Post »

AI, Committee, 新闻, Uncategorized

Aikido Security Releases Altar-1: An Open-Weight Security Model Pruned From GLM-5.3 to 328 GB

Aikido Security has released Altar-1, its first open-weight security model. It is a compressed version of Z.AI’s GLM-5.3, built to run inside infrastructure the customer controls. Altar-1 powers Aikido Machine, the company’s autonomous pentesting appliance for on-prem and air-gapped networks. Is it deployable? Yes, the weights are public on Hugging Face and run with vLLM on a single node of 4x NVIDIA H200 GPUs. The Problem: Security Context Cannot Leave the Network Closed frontier models run on someone else’s infrastructure. Using them sends source code, architecture docs, and unremediated findings outside the network. Aikido points to banks under data-residency mandates and OT operators with no internet route. Open-weight models solve the residency problem but create a deployment gap. Mixture-of-experts (MoE) models must store every expert, even when a workload uses only a few of them. Security agents also build long-running context. That KV cache competes with model weights for the same GPU memory. How Altar-1 Was Built GLM-5.3 is a 753B parameter MoE model. Each token routes to 8 of 256 experts per layer, which is about 40B active parameters. Aikido applied 2 compression steps: Step 1-Quantization: Altar-1 starts from the cyankiwi GLM-5.3-AWQ-INT4 checkpoint. AWQ stores routed expert weights in 4 bits, with 16-bit activations (W4A16). Attention, the shared expert, dense layers, and the head stay in BF16. Step 2- Expert pruning: Aikido used Cerebras REAP (Router-weighted Expert Activation Pruning). REAP scores each expert by router weight and output magnitude, not just by how often it is selected. Altar-1 keeps 168 of 256 routed experts per layer and removes 88 (34.4%). No retraining is involved. Calibration used traces from Aikido’s pentesting harness, plus coding, tool calling, reasoning, and multilingual Wikipedia text. Aikido states no customer data was used. Each expert is scored by its largest share of any single domain’s routed work. That protects the specialist experts for code, rare languages, and structured output. Routing is unchanged. The router still picks 8 experts per token, now from 168, with about 40B active parameters. Checkpoint Stored weights GLM-5.3, BF16 1,506.7 GB GLM-5.3, AWQ INT4 488.2 GB Altar-1, pruned W4A16 328.0 GB Altar-1 is 78.2% smaller than BF16 and 32.8% smaller than the AWQ parent. On fidelity, Altar-1 has a KL divergence of 0.506 nats against full BF16 on a sealed 25-prompt panel. An EXL3 build of the same cut scores 0.511. The details are in the public fidelity study. Benchmark Results Aikido team tested Altar-1 on its internal CVE benchmark. The benchmark covers 32 known vulnerabilities across 30 repositories, with 3 runs per case. Model Avg recall per run Found at least once GLM-5.3, BF16 65.6% 25 of 32 GLM-5.3, AWQ INT4 61.5% 23 of 32 Altar-1 60.4% 23 of 32 Compared with the AWQ checkpoint, pruning cost about 1 point of recall and no coverage. Compared with the parent, Altar-1 keeps 23 of 25 covered vulnerabilities (92%) at 5.2 points lower recall. The benchmark’s scope is narrow. It measures targeted CVE rediscovery inside a pipeline that uses other models for the surrounding stages. It does not measure blind discovery, exploit validation, or fix proposals. Aikido also reports that Altar-1 found a valid critical-severity vulnerability during a client’s production pentest. That is a single result reported by the vendor. Deployment and License The model card requires Hopper GPUs (H100 or H200). Aikido says 328 GB across 4x H200 leaves room for a 128k-context KV cache at production batch sizes. Copy CodeCopiedUse a different Browser vllm serve AikidoSec/altar-1 –tensor-parallel-size 4 –trust-remote-code –max-model-len 131072 vLLM selects the Marlin MoE kernel automatically. A 4x H100 80 GB node has only 320 GB of memory, which is less than the 328 GB of weights. Altar-1 inherits the GLM-5.3 License. The license permits commercial use, modification, and redistribution. Model-as-a-Service operators with more than $10B in revenue over 12 months must first pass a Z.AI security review. Altar-1 is open-weight, not OSI-approved open source. Altar-1 also powers Aikido Attack, AI Code Analysis, and Deep Review. Next, Aikido plans to try lower-bit formats like EXL3 so it can keep more experts. It also plans to fine-tune models for security workflows. Key Takeaways Altar-1 compresses GLM-5.3 from 1,506.7 GB to 328 GB, a 78.2% cut. REAP pruning keeps 168 of 256 experts per layer, with 8 active per token. Recall drops from 65.6% to 60.4%, keeping 92% of the parent’s CVE coverage. It runs on 1 node of 4x H200 with vLLM, including air-gapped setups. The GLM-5.3 License allows commercial use, with a review clause above $10B revenue. Check out the Model Weights and Technical Details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Aikido Security Releases Altar-1: An Open-Weight Security Model Pruned From GLM-5.3 to 328 GB appeared first on MarkTechPost.

Aikido Security Releases Altar-1: An Open-Weight Security Model Pruned From GLM-5.3 to 328 GB Read Post »

AI, Committee, 新闻, Uncategorized

Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding

Liquid AI has announced LFM2.5-VL-3B-DSpark, an experimental speculative-decoding draft model for its LFM2.5-VL-3B vision-language model. The drafter adds about 280M parameters and speeds up decoding without changing the model’s output. Liquid AI team reports up to 3.13x faster decoding on Apple silicon and up to 2.66x on an NVIDIA H100. Is it deployable? Yes, Weights are live on Hugging Face in Safetensors and GGUF, with day-one support in SGLang, MLX-VLM, and llama.cpp. Liquid AI team labels the release experimental, and it ships under the LFM Open License v1.0, which allows free commercial use only for companies under $10M in annual revenue. What Speculative Decoding Changes for a VLM A standard model generates one token per forward pass. Speculative decoding adds a small drafter that proposes several tokens ahead. The large target model then checks the whole block in one pass and keeps the tokens it agrees with. DSpark follows the recipe from Liquid AI’s text-model DSpark drafters, described in the DSpark paper. The drafter reads the target model’s hidden states from several layers and predicts the next k tokens. The key design point: modality does not matter to the drafter. By the time tokens reach the hidden layers, text and image patches are both just tensors. So Liquid AI team reuses the exact same inference algorithm for its vision-language model. Drafter Architecture and Training The drafter is a simplified attention-only model. Ablations picked 4 layers and a block size of 9. Liquid AI recommends a block size of 8 or 9 at inference, depending on hardware. Apple silicon runs use 8. Component Parameters Decoder stack (4 layers) 193.0M Hidden-state projection 21.0M Markov head 65.5M Norms + confidence head 6.4k Total 279.5M The embedding and LM head are tied to the target, so the drafter does not carry them. Liquid AI says this raises the deployed parameter count by 8.9%. Training used supervised fine-tuning data covering common vision-language tasks for 10 epochs. All ablations and training ran exclusively on AMD hardware. Benchmark Results Evaluation follows the MMSpec benchmark across 6 task types: General VQA, Text VQA, Image Captioning, Chart VQA, Complex Reasoning, and Multi-turn Conversation. All runs used batch size 1, temperature 0, and 16-bit weights for the vision encoder and backbone. Data was collected on Pipette, Liquid AI’s public device-benchmarking infrastructure. Stack Decode speedup End-to-end speedup Accepted tokens per pass MLX-VLM, M5 Max MacBook Pro (block 8) 2.30x to 3.13x 1.56x to 2.62x 3.24 to 4.34 llama.cpp, M3 Ultra (block 8) 1.57x to 2.14x 1.30x to 1.77x 3.31 to 4.50 SGLang, 1x H100 80GB (block 9) 2.04x to 2.66x 1.64x to 2.27x 3.46 to 4.57 The ‘up to’ decode and end-to-end figures often come from different tasks. On the M5 Max, 3.13x decode is from COCO captioning, while 2.62x end-to-end is from MMMU-Pro. Acceptance landed in a similar range on both Apple stacks. Liquid AI reads this as acceptance depending on the drafter and workload, not the runtime. At higher concurrency, DSpark kept a throughput advantage at every measured level on a single H100 in SGLang. The gap narrows as concurrency rises. Output Quality and Temperature Under greedy decoding, the target verifies every proposed token, so output is identical to the base model. At non-zero temperatures with matched sampling, speculative decoding preserves the target’s output distribution, as proven by Leviathan et al. Temperature does affect speed. Higher temperatures spread probability across more candidate tokens, so drafter and target disagree more often. In Liquid AI’s tests, this lowered acceptance and throughput. Why End-to-End Gains Are Smaller on Edge Speculative decoding only accelerates decoding. Image encoding and prefill run at the same speed. A VLM must encode the image, then process hundreds of visual tokens alongside the prompt. On edge devices with less compute than data center GPUs, prefill takes a larger share of latency. Liquid AI frames this as Amdahl’s law: total speedup is bounded by the part left unaccelerated. This explains cases like TextVQA on the M5 Max, where 2.69x faster decoding yields 1.56x end to end. How to Run It SGLang requires v0.5.19 or newer. Launch LiquidAI/LFM2.5-VL-3B with –speculative-algorithm DSPARK and point –speculative-draft-model-path at the drafter. On Apple silicon, MLX-VLM v0.7.2 or newer accepts the drafter through –draft-model. DSpark in MLX-VLM currently supports greedy sampling only, so set temperature to 0. For llama.cpp, pair the GGUF drafter with the LFM2.5-VL-3B-GGUF target. Integration work is public in the llama.cpp, SGLang, and MLX-VLM pull requests. Acceleration of quantized models is outside the scope of this release. Key Takeaways A 279.5M drafter adds 8.9% parameters to LFM2.5-VL-3B. Decoding runs up to 3.13x faster on M5 Max, 2.66x on H100. Output is identical under greedy decoding; distribution preserved when sampling. Prefill and vision encoding cap end-to-end gains, especially on edge. Tested at 16-bit only; quantized acceleration is not covered yet. Check out the Technical Details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding appeared first on MarkTechPost.

Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding Read Post »

We use cookies to improve your experience and performance on our website. You can learn more at 隱私權政策 and manage your privacy settings by clicking Settings.

Privacy Preferences

You can choose your cookie settings by turning on/off each type of cookie as you wish, except for essential cookies.

Allow All
Manage Consent Preferences
  • Always Active

Save
zh_CN