Local Agentic AI Workflows with Hermes + Ollama
In this article, you will learn how to build a fully local, zero-cost agentic AI workflow using Hermes Agent and Ollama, so that your files,…
Local Agentic AI Workflows with Hermes + Ollama Read Post »
In this article, you will learn how to build a fully local, zero-cost agentic AI workflow using Hermes Agent and Ollama, so that your files,…
Local Agentic AI Workflows with Hermes + Ollama Read Post »
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Last Wednesday, Anthropic announced that earlier this year it had launched a molecular biology lab, where Claude agents read and conjecture about hard biology problems and human scientists run experiments on what they report. And this AI-powered lab, the company said, had made its first discovery. To understand what Anthropic says its system did, imagine you’re flipping through a library of millions of DNA sequences, amassed as scientists sequence more and more of the living world. One step toward a breakthrough might be finding a peculiar sequence that encodes an interesting enzyme, perhaps. Then you’d need to figure out what that enzyme does and, eventually, how to manipulate it to do something useful. What Anthropic says its system of 950 agents found after 21 hours was not a brand-new sequence. The agents instead flagged a repeating pattern surrounding a known enzyme, a particular pattern Anthropic said hadn’t been catalogued before. But if you read through Anthropic’s announcement, which calls this pattern “reminiscent” of what led to the gene-editing technology CRISPR that “has already transformed science and medicine,” it sounds as if this army of agents really found something of note. These claims have angered some biologists. A viral post from one, subsequently endorsed by the chair and CEO of the drugmaker Eli Lilly, said that “finding a weird cluster of genes and repeats is often the easy part. The hard part, and where the real discoveries come from, is figuring out what the system actually does.” The agents helped with some laboratory grunt work, in other words. But a discovery it is not. It’s a reminder that even if AI does something impressive—like finding a pattern in a mass of biological data that would be difficult to perceive with human eyes alone—the result itself may not constitute a breakthrough for science. What is novel for AI may be routine, unsurprising, or simply not that consequential to a biologist. Muddying the issue further, Mario Rodríguez Mestre, a biologist at the University of Copenhagen, said over the weekend that his team had already discovered this particular pattern, the New York Times reported. Mestre, who regularly chatted with Claude in his work, wondered whether Anthropic’s team had learned from his conversations. Anthropic denies this, but Mestre says he’s stopping all use of Claude anyway. Part of the problem here is that AI companies aren’t presenting their systems simply as tools scientists can use, like microscopes or supercomputers. They’re insisting that the AI systems are making discoveries themselves. To some, that approach is incompatible with how science actually works, with new knowledge more typically emerging from collaboration and an ever-growing arsenal of tools. It’s also making people more skeptical of genuine progress when it happens. Whittling 200,000 candidates down to a few worth exploring is no small feat; it is legitimate scientific work. The fact that a general-purpose chatbot could do that work is notable, even if humans helped steer it and ultimately ran the experiments. But once the standard is whether Claude itself made a discovery, all that becomes evidence for one side or the other in a debate that has only two answers: breakthrough or bust. Once we’re judging AI by whether it has made a discovery, it’s also tempting to shift the goalposts even after it really does seem to notch a win. Earlier this month, OpenAI said its own team agents had cracked a million-dollar problem in mathematics. But a couple of weeks later, nearly every AI skeptic in my feed was sharing an article asking whether it was the math problem that really mattered. To be clear, the piece did not argue that OpenAI’s solution was wrong. Instead, it argued that the particular result may not be the one mathematicians care most about. Throw in the accusation by a mathematician that the models may have used some of his work without credit, and people are left thinking either OpenAI cheated or the solution wasn’t important anyway. Or both. That’s part of what concerns Lucas Harrington, the biologist who wrote the post critiquing Anthropic’s announcement. He closed with a suggestion: AI companies, he said, should “set the bar high now, so that when an AI actually discovers a fundamentally new biological mechanism, everyone appreciates how big a deal it is.” But as OpenAI’s Sam Altman and Anthropic’s Dario Amodei race to one-up each other, raising the bar for scientific breakthroughs by AI might be the last thing on their minds.
When can we say AI made a scientific discovery? Read Post »
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Who’s liable when AI agents go rogue? Over the past few months, a cascade of cyberattacks by AI agents has stunned the world. In July, OpenAI disclosed that a swarm of its agents had escaped their sandbox and hacked into the AI platform Hugging Face to cheat on a cybersecurity test. Many experts say it’s only a matter of time until there’s a more damaging incident where AI agents bypass sandboxes to access systems they shouldn’t. But the big question is: How do we hold companies liable when they lose control of their AI agents? —Michelle Kim The AI Hype Index Separating AI reality from hyped-up fiction isn’t always easy. That’s why we’ve created the AI Hype Index—a simple, at-a-glance summary of what’s shaping the industry right now. The latest edition includes Chinese chipmakers, German wiki sites, and American Terminators. See where it all landed on this month’s index. —Michelle Kim Roundtables: the deadly failures of the virtual border wall Last week, we published an MIT Technology Review investigation that found more than a thousand people died within the advertised range of surveillance towers along the US border. Later today, our editor-in-chief Mat Honan, senior AI reporter James O’Donnell and senior reporter for features and investigations Eileen Guo will join a subscriber-only conversation about the investigation, and its implications. Register now to attend on Monday, September 28 at 7:00pm BST / 2:00pm EDT / 11:00am PDT. Want to join the conversation? Subscribe to MIT Technology Review for exclusive access to all our Roundtables.Read the full MIT Technology Review investigation here. The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology 1 OpenAI says it has paused training its models The decision comes after its agents interacted with US government websites. (AP) + Trump hosted Anthropic boss Dario Amodei at a White House dinner last night. (FT) + How much access will evaluators really have inside AI companies? (Atlantic) + Why can’t we just keep rogue AI off the internet? (Verge) + Democrats are in a mad scrabble to get AI “right”. (New York Magazine) + An AI kill switch isn’t that simple, after all. (Bloomberg) + Will AI really kill us all? Your questions, answered. (MIT Technology Review) 2 Why isn’t the data center backlash also a climate reckoning? It’s still hard to get people to care about the emissions they create. (Wired) + What does the climate movement do now? (Atlantic) + The math behind data centers and energy. (MIT Technology Review) 3 Wall Street big bears aren’t ready to bet against AI, yet Bubble talk is growing, but not many are willing to go against the crowd. (Information $) + What even is the AI bubble? (MIT Technology Review) 4 Did Anthropic’s AI really make a scientific discovery on its own? One scientist shared research with Claude, and says the new finding matches. Hmm. (NYT) + AI for science needs reasoning, not just data. (MIT Technology Review) 5 Ancient superbugs might help us fight antibiotic resistance They could be emerging from the permafrost. (New Scientist) 6 More reliable—and less easy to jam—alternatives to GPS are coming Using Earth’s quantum field could be more secure. (Economist) 7 Why social media bans aren’t enough to keep children safe Online child safety deserves more nuanced policy. (IEEE Spectrum, Opinion) 8 This is how the US is attacking China’s control of critical minerals It’s a multibillion-dollar effort to loosen Beijing’s chokehold. And it’s working. (WSJ) 9 No one wants to date tech bros anymore They used to be nerdy and harmless, now they’ve got a real image problem. (Wired) 10 What are the rules around cellphone etiquette now? Is there an obligation to text someone right back, for instance? (Vox) Quote of the day “If you’re evil, you’re at least competent. And if you’re evil, you’re not bad, and therefore you’re actually maybe kind of good because you’re at least getting something done.” —Venture capitalist and PayPal cofounder Peter Thiel shares his unusual take on competence versus morality in an interview with Axel Springer CEO Mathias Döpfner. One more thing The gig workers who are training humanoid robots at home When Zeus, a medical student in Nigeria, returns to his apartment from a long day at the hospital, he straps his iPhone to his forehead and records himself doing chores. Zeus is a data recorder for Micro1, which sells the data he collects to robotics firms. As these companies race to build humanoids, videos from workers like Zeus have become the hottest new way to train them. Micro1 has hired thousands of them in more than 50 countries, including India, Nigeria, and Argentina. The jobs pay well locally, but raise thorny questions around privacy and informed consent. The work can be challenging—and weird. Read the full story. ——Michelle Kim We can still have nice things A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.) + This surfboard comes with a 10-liter beer keg built right into it.+ Hollywood’s iconic Cinerama dome movie theatre is set to reopen at long last in early 2028.+ From Dracula to Marilyn Monroe, here are 25 famous quotes in pop culture you’ve probably been getting wrong.+ Baby lobsters in mini pods and a human-sized floating waterlily highlight this stunning selection of science photos.
The Download: rogue agent liability and the AI Hype Index Read Post »
NVIDIA has launched the NVIDIA Open Agent Safety Platform, an open software platform and reference system design for AI agent security. It pairs the OpenShell secure runtime with NVIDIA Sentry, an out-of-band watchdog on BlueField-4 DPUs. The core idea is simple. Safety controls should not live inside the agent they are meant to control. Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry. Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to… pic.twitter.com/dReAxwpRUn — Jensen Huang (@JensenHuang) September 28, 2026 Is it deployable today? Yes for OpenShell. It is Apache 2.0, installs on Linux, macOS (Apple Silicon) or Windows WSL 2, and its repo still labels it alpha. Why NVIDIA Moved Enforcement Below the Agent The NVIDIA technical report cites recent reports from several frontier labs. Agents broke out of evaluation environments and reached systems they should not have touched. Some agents misreported what they did. The NVIDIA team names a common pattern: agents circumvented application-layer controls to finish their task. NVIDIA calls this failure mode drift. Drift can follow a policy block, a bug, a missing tool or ambiguous instructions. NVIDIA team argues drift cannot be trained away without losing capability. So an agent cannot be expected to fully govern itself. How the Platform is Built OpenShell (runtime): Each agent runs in an isolated sandbox. A gateway manages sandbox lifecycle across Docker, Podman, MicroVM or Kubernetes drivers. Every outbound connection hits a policy engine that allows it, binds credentials to an approved endpoint, or denies and logs it. Filesystem and process rules lock at creation. Network and provider rules are hot-reloadable. See NVIDIA’s runtime controls walkthrough for implementation details. Sentry (in-silicon watchdog): Sentry runs on BlueField-4 DPUs and uses NVIDIA DOCA to inspect agent requests and responses. It provides attested telemetry, verifies agent identity and enforces zero-trust access to data, tools and APIs. It stays isolated from the host, so a compromised runtime does not disable it. Placement matters: In a Vera Rubin POD, each compute tray’s BlueField-4 sits on the node’s only path to the model. An agent cannot act without its next inference call. That makes the path both the best observation point and the kill switch. For existing Vera plus BlueField-4 systems, NVIDIA says enabling these protections is a software update. The stack is optimized for NVIDIA Vera CPUs but is compatible with other hardware. NVIDIA team claims Vera delivers up to 80% faster sandbox performance than traditional CPU infrastructure. OpenShell can also be extended to Arm and Intel platforms. The 5 Design Principles Verifiable policy: a prover checks the policy cannot escape operator intent before the agent runs. Out-of-band enforcement: controls sit outside the agent’s reach. Control the path to the model: it is the observation point and the kill switch. Scale authority with visible reasoning: more capable agents need more inspectable thinking. Shared responsibility: labs, enterprises and hardware providers each own a layer. Interactive Explainer: Send a Request Through the Stack How It Compares With Other Agent Sandboxes The closest alternatives are sandbox platforms for agent-generated code. Neither offers an equivalent hardware watchdog. Feature NVIDIA OpenShell + Sentry E2B Daytona Type Open runtime plus hardware reference design Open-source sandbox cloud Sandbox infrastructure runtime License Apache 2.0 Apache 2.0 AGPL-3.0 (public repo unmaintained since June 2026) Isolation Per-sandbox container or MicroVM, kernel-level isolation Firecracker microVM, own kernel Dedicated kernel, filesystem and network stack per sandbox Egress control YAML policy at HTTP method and path level, hot-reloadable Allow and deny lists by IP, CIDR or domain Network limits Out-of-band hardware enforcement Yes, Sentry on BlueField-4 (optional) No, software isolation No, software isolation Where it runs Local, on-prem, cloud, Kubernetes (experimental) E2B cloud or self-hosted on AWS and GCP Daytona cloud Agent support Claude Code, Codex, OpenCode, Copilot CLI built in JS and Python SDKs Python, TypeScript, Ruby, Go, Java SDKs Who is Building on It NVIDIA says over 100 organizations work with the platform. Anthropic integrated Claude Managed Agents with OpenShell and BlueField. SpaceXAI uses it for Cursor coding agents and Grok models. Salesforce connected OpenShell to Slack for approving agent permission requests. SAP is embedding OpenShell in the Joule Studio runtime. Red Hat, SUSE and Canonical are integrating it into their operating systems. The effort feeds the Open Secure AI Alliance, governed by the Linux Foundation. OpenShell and its skills are available on GitHub and the OpenShell docs. Key Takeaways 2 layers: OpenShell sandboxes the agent, Sentry watches it from separate silicon. Sentry can quarantine an agent that leaves its boundary in milliseconds, per NVIDIA. OpenShell policies are declarative YAML, with network rules enforced at HTTP method and path level. OpenShell runs Claude Code, Codex, OpenCode and GitHub Copilot CLI out of the box. NVIDIA lists over 100 organizations working with the platform, including Anthropic and Microsoft. FAQ Does OpenShell require BlueField-4? No. It runs on local, on-prem, cloud and Kubernetes infrastructure. BlueField-4 only adds Sentry. How is this different from model guardrails? Guardrails shape what an agent attempts. Runtime controls enforce what it is allowed to do. Can I use existing agents and models? Yes. OpenShell supports open and closed models and custom sandbox images. Check out the Paltform here and Technical Details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post NVIDIA Launches Open Agent Safety Platform: OpenShell Sandboxes Agents on Vera CPUs While Sentry on BlueField-4 Quarantines Them in Milliseconds appeared first on MarkTechPost.
In this tutorial, we work with MSEB, the Massive Sound Embedding Benchmark from Google Research, and approach it from the perspective of what a leaderboard number actually means: the evaluator surface. We install the package and map its three layers, then write two deliberately different encoders against the framework’s own abstract base class: one that measures loudness over time and one that measures timbre, and encode a small synthetic corpus we generate in the notebook so nothing has to be downloaded. We drive the classification, clustering, retrieval, and segmentation evaluators over those embeddings, call the metric functions directly to see what each one rewards, and finish by assembling the TaskMetadata a real submission carries. The result is a comparison in which the two encoders trade places depending on which evaluator is asked, which is the argument for a multi-task benchmark made in numbers rather than in prose. Copy CodeCopiedUse a different Browser import os import sys import json import math import traceback import subprocess import numpy as np RESULTS = {} BENCH = {} def banner(title): print(“n” + “=” * 78) print(title) print(“=” * 78) def section(name): def wrap(fn): def run(*a, **kw): banner(name) try: out = fn(*a, **kw) RESULTS[name] = out if isinstance(out, str) else “ok” return out except Exception as e: RESULTS[name] = f”SKIPPED / FAILED -> {type(e).__name__}: {e}” print(f”n[!] {name} did not complete: {type(e).__name__}: {e}”) traceback.print_exc(limit=3) return None return run return wrap banner(“0. Install MSEB and map the three layers we will use”) subprocess.run([sys.executable, “-m”, “pip”, “install”, “-q”, “mseb==0.1.0″], check=True) import mseb from mseb import types, encoder as encoder_lib, evaluator as evaluator_lib, metrics from mseb.evaluators import ( classification_evaluator, clustering_evaluator, retrieval_evaluator, segmentation_evaluator, ) print(f” mseb {mseb.__version__} | Python {sys.version.split()[0]} | numpy {np.__version__}”) print(“n MSEB is three layers, and a benchmark run walks down them:”) print(” types -> Sound, SoundEmbedding, Score, TaskMetadata: the shapes every task speaks”) print(” encoder -> MultiModalEncoder: the contract YOUR model implements”) print(” evaluators -> classification, clustering, retrieval, reranking, transcription, segmentation, …”) print(“n evaluator entry points we will drive:”) for module, cls in [(classification_evaluator, “ClassificationEvaluator”), (clustering_evaluator, “ClusteringEvaluator”), (retrieval_evaluator, “RetrievalEvaluator”), (segmentation_evaluator, “SegmentationEvaluator”)]: print(f” {module.__name__.split(‘.’)[-1]:28s} {cls}”) print(“n Everything below runs on CPU with no dataset download: we synthesise the audio.”) We install mseb and import the three layers that a benchmark run walks down. The types module holds the shapes every task speaks, Sound, SoundEmbedding, Score and TaskMetadata; the encoder module holds MultiModalEncoder, the contract our own model implements; and the evaluators package holds one module per task family. We import only the four evaluators this notebook drives, because the classification, clustering, retrieval, and segmentation modules depend on nothing heavier than NumPy and scikit-learn. In contrast, the reranking and transcription evaluators pull in Whisper and the task runner pulls in TensorFlow and apache-beam. Everything below therefore runs on a free CPU runtime with no dataset download and no accelerator. Copy CodeCopiedUse a different Browser SR = 16000 @section(“1. The type contract: Sound, SoundEmbedding, Score”) def type_contract(): t = np.arange(SR) / SR waveform = (0.5 * np.sin(2 * np.pi * 440 * t)).astype(np.float32) sound = types.Sound( waveform=waveform, context=types.SoundContextParams(id=”demo_000″, sample_rate=SR, length=len(waveform), language=”en_us”, text=”a 440 Hz tone”), ) print(f” Sound id={sound.context.id!r} {sound.waveform.shape} @ {sound.context.sample_rate} Hz” f” -> {sound.size_bytes:,} bytes”) embedding = types.SoundEmbedding( embedding=np.zeros((1, 16), dtype=np.float32), # (N, D): one utterance-level vector timestamps=np.array([[0.0, 1.0]], dtype=np.float32), # (M, 2): [start, end] in seconds context=sound.context, encoding_stats=types.EncodingStats(input_size_bytes=sound.size_bytes, embedding_size_bytes=16 * 4), ) print(f” SoundEmbedding embedding{embedding.embedding.shape} timestamps{embedding.timestamps.shape}” f” -> {embedding.size_bytes} bytes”) print(f” compression_ratio = {embedding.encoding_stats.compression_ratio:.5f}” f” ({1 / embedding.encoding_stats.compression_ratio:,.0f}x smaller than the audio)”) print(” N embeddings and M timestamps: M == N is frame-aligned, M == 1 is utterance-level.”) print(” `embedding` may also hold N strings instead of vectors – step 8 uses exactly that.”) score = types.Score(metric=”Accuracy”, description=”Overall classification accuracy”, value=0.875, min=0.0, max=1.0) print(f”n Score {score.metric}={score.value} in [{score.min}, {score.max}] :: {score.description}”) for bad, why in [(dict(metric=””, description=”d”, value=0.5, min=0.0, max=1.0), “empty metric name”), (dict(metric=”m”, description=”d”, value=0.5, min=1.0, max=0.0), “min > max”)]: try: types.Score(**bad) except Exception as e: print(f” rejected at construction ({why}): {type(e).__name__}: {e}”) return f”Sound {sound.size_bytes:,} B -> embedding {embedding.size_bytes} B” type_contract() We start with the type contract, because every other layer is expressed in it. A Sound carries a waveform, along with SoundContextParams, the identifier, sample rate, length, language, and optional transcript, which follow the audio through the whole pipeline. A SoundEmbedding carries an array of N embeddings and an array of M timestamp pairs, and the relation between N and M is the benchmark’s vocabulary: M equal to N means one vector per frame, while M equal to one means a single utterance-level vector, which is what our encoders produce. EncodingStats records the input and embedding sizes and exposes compression_ratio, here a thousandfold reduction from audio to vector. A Score is a metric name, a value and its bounds, and it validates itself at construction, rejecting an empty metric name or a minimum above its maximum, so a malformed number cannot reach a leaderboard. The embedding field also accepts N strings instead of N vectors, which is the door that step 8 walks through. Copy CodeCopiedUse a different Browser class EnergyEnvelopeEncoder(encoder_lib.MultiModalEncoder): “””Baseline: average energy in `n_bins` equal time slices. Loud/quiet, nothing about timbre.””” def __init__(self, n_bins: int = 16): super().__init__() self.n_bins = n_bins def _setup(self): self._ready = True # a real encoder loads weights here def _check_input_types(self, batch): for item in batch: if not isinstance(item, types.Sound): raise ValueError(f”{type(self).__name__} takes types.Sound, got {type(item).__name__}”) def _encode(self, batch) -> list[types.SoundEmbedding]: out = [] for sound in batch: slices = np.array_split(sound.waveform.astype(np.float32), self.n_bins) vec = np.array([[float(np.sqrt(np.mean(s ** 2) + 1e-12)) for s in slices]], dtype=np.float32) vec /= np.linalg.norm(vec) + 1e-9 out.append(types.SoundEmbedding( embedding=vec, timestamps=np.array([[0.0, sound.context.length / sound.context.sample_rate]], dtype=np.float32), context=sound.context)) return out class SpectralProfileEncoder(encoder_lib.MultiModalEncoder): “””Contender: mean log-magnitude spectrum pooled into `n_bands` bands. Describes timbre.””” def __init__(self, n_bands: int = 16, frame: int = 512): super().__init__() self.n_bands, self.frame = n_bands, frame def _setup(self): self._window = np.hanning(self.frame).astype(np.float32) def _check_input_types(self, batch): for item in batch: if not isinstance(item, types.Sound): raise ValueError(f”{type(self).__name__} takes types.Sound, got {type(item).__name__}”) def _encode(self, batch) -> list[types.SoundEmbedding]: out = [] for sound in batch: w = sound.waveform.astype(np.float32) n_frames =
In this tutorial, we build a comprehensive multimodal augmentation and robustness workflow with AugLy for images, text, and audio. We start by addressing modern dependency compatibility issues and generating deterministic synthetic datasets so the experiments remain self-contained and reproducible. We then explore AugLy’s functional and class-based APIs, metadata, and intensity tracking, probabilistic composition, bounding-box-aware transformations, and custom transforms. We extend the workflow into practical robustness experiments by benchmarking perceptual-hash copy detection under image distortions and evaluating text classifiers against adversarial perturbations, Unicode obfuscation, sanitization, and adversarial training. We also integrate audio augmentation, build a queryable metadata warehouse, and connect AugLy transformations directly to PyTorch datasets and DataLoaders, giving us an end-to-end view of augmentation as both a data-generation mechanism and a measurable robustness tool. Copy CodeCopiedUse a different Browser import subprocess, sys, importlib def _sh(cmd): print(f”$ {cmd}”) subprocess.run(cmd, shell=True, check=False, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL) def _need(mod): try: importlib.import_module(mod) return False except ImportError: return True if _need(“augly”): _sh(“apt-get -qq install -y libmagic1 > /dev/null 2>&1″) _sh(f'”{sys.executable}” -m pip install -q –no-deps augly’) _sh(f'”{sys.executable}” -m pip install -q “iopath>=0.1.8” “python-magic>=0.4.22″ ‘ f'”regex>=2021.4.4” “nlpaug==1.1.3″‘) import numpy as np from PIL import Image, ImageDraw, ImageFont, ImageFilter for _name, _builtin in ((“float”, float), (“int”, int), (“bool”, bool)): if not hasattr(np, _name): setattr(np, _name, _builtin) def _size(font, text): left, top, right, bottom = font.getbbox(text) return (right, bottom) if not hasattr(ImageFont.FreeTypeFont, “getsize”): ImageFont.FreeTypeFont.getsize = lambda self, t, *a, **k: _size(self, t) if not hasattr(ImageFont.FreeTypeFont, “getsize_multiline”): def _getsize_multiline(self, text, direction=None, spacing=4, features=None, language=None, stroke_width=0): lines = text.split(“n”) w = max((_size(self, ln)[0] for ln in lines), default=0) h = sum(_size(self, ln)[1] for ln in lines) + spacing * (len(lines) – 1) return (w, h) ImageFont.FreeTypeFont.getsize_multiline = _getsize_multiline import os, io, json, math, random, string, textwrap, unicodedata, warnings from dataclasses import dataclass from typing import Any, Dict, List, Optional, Tuple import matplotlib.pyplot as plt import pandas as pd import augly.image as imaugs import augly.text as textaugs import augly.utils as augutils from augly.image.transforms import BaseTransform as ImageBaseTransform warnings.filterwarnings(“ignore”) pd.set_option(“display.width”, 160) SEED = 1234 random.seed(SEED) np.random.seed(SEED) print(“n” + “=” * 78) print(“AugLy ready. assets at:”, augutils.ASSETS_BASE_DIR) print(“image augs :”, len([f for f in dir(imaugs) if f[0].islower()])) print(“text augs :”, len([f for f in dir(textaugs) if f[0].islower()])) print(“=” * 78 + “n”) def make_image(idx: int, w: int = 320, h: int = 240) -> Tuple[Image.Image, Tuple[int, int, int, int]]: “””Procedurally generated ‘photo’ + a ground-truth bbox in pascal_voc format.””” rng = random.Random(SEED + idx) img = Image.new(“RGB”, (w, h), tuple(rng.randint(20, 90) for _ in range(3))) d = ImageDraw.Draw(img) for _ in range(70): x0, y0 = rng.randint(0, w), rng.randint(0, h) d.line([x0, y0, x0 + rng.randint(-60, 60), y0 + rng.randint(-60, 60)], fill=tuple(rng.randint(60, 160) for _ in range(3)), width=rng.randint(1, 3)) ow, oh = rng.randint(70, 130), rng.randint(60, 110) ox, oy = rng.randint(10, w – ow – 10), rng.randint(10, h – oh – 10) box = (ox, oy, ox + ow, oy + oh) colour = tuple(rng.randint(150, 255) for _ in range(3)) if idx % 3 == 0: d.ellipse(box, fill=colour, outline=(255, 255, 255), width=3) elif idx % 3 == 1: d.rectangle(box, fill=colour, outline=(255, 255, 255), width=3) else: d.polygon([(ox + ow // 2, oy), (ox + ow, oy + oh), (ox, oy + oh)], fill=colour, outline=(255, 255, 255)) return img, box N_IMAGES = 24 IMAGES, BOXES = zip(*[make_image(i) for i in range(N_IMAGES)]) IMAGES, BOXES = list(IMAGES), list(BOXES) DEMO_IMG, DEMO_BOX = IMAGES[0], BOXES[0] def make_text_dataset(n_per_class: int = 260): “””Tiny sentiment corpus built from templates -> learnable but not trivial.””” rng = random.Random(SEED) pos_adj = [“excellent”, “delightful”, “superb”, “charming”, “brilliant”, “flawless”, “wonderful”, “outstanding”, “impressive”, “lovely”] neg_adj = [“terrible”, “awful”, “dreadful”, “disappointing”, “clumsy”, “broken”, “miserable”, “useless”, “painful”, “sloppy”] subj = [“the movie”, “this restaurant”, “the hotel room”, “their support team”, “the new phone”, “the sequel”, “this laptop”, “the delivery service”] tail_p = [“and I would recommend it to anyone”, “worth every rupee”, “I left completely satisfied”, “easily the best of the year”, “it exceeded all my expectations”] tail_n = [“and I want a refund”, “a total waste of money”, “I left extremely frustrated”, “easily the worst of the year”, “it failed every expectation”] rows = [] for _ in range(n_per_class): rows.append((f”{rng.choice(subj)} was {rng.choice(pos_adj)} {rng.choice(tail_p)}”, 1)) rows.append((f”{rng.choice(subj)} was {rng.choice(neg_adj)} {rng.choice(tail_n)}”, 0)) rng.shuffle(rows) return [r[0] for r in rows], [r[1] for r in rows] TEXTS, LABELS = make_text_dataset() DEMO_TEXT = “The quick brown fox jumps over the lazy dog near the river bank” def make_audio(seconds: float = 2.0, sr: int = 16000) -> Tuple[np.ndarray, int]: “””A chirp + harmonics + a little noise = something you can actually hear change.””” t = np.linspace(0, seconds, int(sr * seconds), endpoint=False) f = np.linspace(220, 880, t.size) sig = 0.5 * np.sin(2 * np.pi * f * t) + 0.2 * np.sin(2 * np.pi * 2 * f * t) sig += 0.02 * np.random.RandomState(SEED).randn(t.size) env = np.minimum(1.0, np.minimum(t * 8, (seconds – t) * 8)) return (sig * env).astype(np.float32), sr AUDIO, SR = make_audio() def show_grid(pairs, cols=4, title=””, figsize_scale=2.9): “””pairs: list of (caption, PIL.Image).””” rows = math.ceil(len(pairs) / cols) fig, axes = plt.subplots(rows, cols, figsize=(cols * figsize_scale, rows * figsize_scale)) axes = np.atleast_1d(axes).ravel() for ax, (cap, im) in zip(axes, pairs): ax.imshow(im) ax.set_title(cap, fontsize=8) ax.axis(“off”) for ax in axes[len(pairs):]: ax.axis(“off”) if title: fig.suptitle(title, fontsize=13, y=1.0) plt.tight_layout() plt.show() def as_str(out) -> str: “””AugLy text augs return str for str input in some transforms, list in others.””” return out[0] if isinstance(out, list) else out print(“n### §2 IMAGE AUGMENTATION + METADATA ” + “#” * 38) functional_result = imaugs.pixelization(DEMO_IMG, ratio=0.25) class_result = imaugs.Pixelization(ratio=0.25, p=1.0)(DEMO_IMG) print(“functional == class:”, np.array_equal(np.array(functional_result), np.array(class_result))) IMAGE_ZOO = { “blur”: lambda im, m: imaugs.blur(im, radius=3.0, metadata=m), “brightness”: lambda im, m: imaugs.brightness(im, factor=1.7, metadata=m), “color_jitter”: lambda im, m: imaugs.color_jitter(im, brightness_factor=1.3, contrast_factor=1.4, saturation_factor=1.6, metadata=m), “crop”: lambda im, m: imaugs.crop(im, x1=.15, y1=.15, x2=.85, y2=.85, metadata=m), “encoding_quality”: lambda im, m: imaugs.encoding_quality(im, quality=8, metadata=m), “grayscale”: lambda im, m: imaugs.grayscale(im, metadata=m), “hflip”: lambda im, m: imaugs.hflip(im, metadata=m), “meme_format”: lambda im, m: imaugs.meme_format(im, text=”TOP TEXT”, caption_height=90, metadata=m), “opacity”: lambda im, m: imaugs.opacity(im, level=0.45, metadata=m), “overlay_emoji”: lambda im, m: imaugs.overlay_emoji(im, opacity=0.9, emoji_size=0.35, metadata=m), “overlay_screenshot”: lambda im, m: imaugs.overlay_onto_screenshot(im, metadata=m), “overlay_stripes”: lambda im, m: imaugs.overlay_stripes(im, line_width=0.4, line_opacity=0.7, metadata=m),
Exa has released Agent Ultra, the highest effort level of its Exa Agent API. It is built for research that must run to exhaustion: large list building, entity enrichment, and questions that need thousands of sources. Exa team reports that Ultra beats Opus 5.5, GPT-6 Astra, and Perplexity Agent, each at maximum effort, on 4 research benchmarks. Is it deployable? Yes, as a hosted API. Agent Ultra is live today on the Exa API by setting effort: “ultra”. It is not open weights and cannot be self-hosted. What is Exa Agent Ultra? Exa Agent splits a task into subtasks and assigns subagents to research several domains at once. It routes frontier models to steps that need them and faster models where those are enough. Ultra is the mode that spends the most compute. According to the Agent Ultra docs, it runs longer than any other effort to return the most complete results. Ultra runs typically finish complex tasks in about 30 minutes. Very hard tasks can take up to 3 hours. Benchmark Results All figures below are from Exa’s launch post. Competitors ran at their maximum effort setting. Benchmark (metric) Agent Ultra Opus 5.5 GPT-6 Astra Perplexity Agent WANDR (soft recall) 81.4% 72.3% 26.0% 40.1% DeepSearchQA (F1) 93.9% 77.6% 85.3% 89.7% WideSearch (row-level F1) 58.9% 51.6% 54.7% 56.0% Company Find-All (avg. passing entities per task) 2,451 146 113 98 Exa pairs each result with a cost claim: WANDR: +12.6% over Opus 5.5, at half its cost per task. DeepSearchQA: +4.7% over Perplexity, at 46% lower cost per task than GPT-6 Astra. WideSearch: +5.2% over Perplexity, at the lowest cost per task of the 4 systems. Company Find-All: +1579% over Opus 5.5, at the lowest cost per entity found. These gains are relative, not percentage points. On WANDR, the absolute gap to Opus 5.5 is 9.1 points. Understanding These Numbers WANDR is Perplexity’s benchmark of 500 wide and deep data-collection tasks, with an open harness. Exa’s grader shares the upstream evaluation logic. It swaps in Exa as the contents tool, changes transport logic, and uses gpt-6-luna as the judge. Where a vendor had published a result on this harness, Exa reports that figure. Otherwise, Exa ran the benchmark itself. DeepSearchQA is Google DeepMind’s 900-prompt multi-step search benchmark. WideSearch tests broad information gathering. Exa evaluated up to 200 tasks each for WANDR and DeepSearchQA, and 100 each for WideSearch and Company Find-All. Graded task counts vary by provider. All results are vendor-reported and not yet independently reproduced. Where Agent Ultra Fits Exa lists 3 target user groups: Model providers: assemble training data, such as every paper and repo implementing a given technique. Verify criteria like ‘released weights, not just an API’. Financial services: build diligence market maps, run KYC research across filings and court records, and monitor portfolio signals. Go-to-market teams: build account lists and enrich rows with judgment fields, each backed by a cited URL. Ultra can also expand an existing list. Pass the rows you already have, and they are excluded from new results. API, Pricing, and Controls Ultra uses the standard Agent run endpoint. The request supports outputSchema, input.data, and streaming. Copy CodeCopiedUse a different Browser from exa_py import Exa exa = Exa() run = exa.agent.runs.create( query=”Find all companies building browser automation tools in the United States.”, effort=”ultra”, ) run = exa.agent.runs.poll_until_finished(run.id, timeout_ms=3 * 60 * 60 * 1000) print(run.stop_reason) Pricing: metered at standard Agent usage rates, up to a default $20 per run. Runs that finish early cost less. Budget: maxCostDollars accepts $1 to $100. maxDurationSeconds accepts 300 to 10,800 seconds. Stopping: a stop call ends a run early, keeps its results, and bills usage up to that point. Timeouts: SDK polling helpers time out after 1 hour by default, so set a longer timeout or stream events. OpenAI compatibility: on /responses, set reasoning.effort: “ultra” with streaming or background mode. You can test it in the Exa API Playground. Comparison Key Takeaways Agent Ultra is Exa Agent’s highest effort mode, live now via API. It orchestrates parallel subagents and mixes frontier and faster models. Exa reports top scores on WANDR, DeepSearchQA, WideSearch, and Company Find-All. Runs cost up to $20 by default, adjustable from $1 to $100. Typical runs take about 30 minutes, with a 3 hour ceiling. Check out the Technical Details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Exa Launches Agent Ultra: A Subagent Swarm Deep Research API Built for Exhaustive List Building appeared first on MarkTechPost.
Supersonic Labs, a small AI lab from Brazil, has released Julia 1. It is a compact decision model, not a chatbot. You pass it context, a question, and 2 to 20 candidate answers. It picks one and returns a probability for every option. The model has 144.3M parameters and runs on a plain CPU. Is it deployable? Yes. The weights are on Hugging Face under Apache 2.0 and run locally with Python 3.11+ on CPU or a BF16-capable GPU. An ONNX build also runs in the browser via WebGPU. A hosted API is announced but not open yet. What Julia 1 Does Julia 1 handles three decision types through one API: choice: pick one label from 2 to 20 described options (classification, routing). score: return the expected index on an ordered rubric, such as low, medium, high. noul: return the probability that a yes-or-no statement is true. Results come back in the caller’s option order with full softmax probabilities. Caller IDs such as billing are returned unchanged. The model does not generate text. Architecture and Training Budget Julia 1 starts from JHU CLSP’s mmBERT-small, a 140M-parameter multilingual ModernBERT encoder trained on 1,800+ languages. Supersonic Labs kept the encoder and tokenizer, added a decision head, and trained on decision-format examples. The lab states Julia 1 is not a fine-tuned Qwen model. The runtime supports 8,192 combined tokens, but published benchmarks used a 1,024-token limit. Total cloud GPU spend for training and experiments was about R$540 (US$104.08). The FP32 weights occupy 550.5 MiB. The private training pipeline is not released. Julia 2, with the lab’s own foundation architecture, is in development. Benchmark Results The September 24, 2026 evaluation ran on H200 BF16 with strict encoding. The comparison baseline is TypeSafe’s Jev, using reference values from the Jev benchmark protocol, not a new Jev run. Typed Decisions: 73.15% (1,463/2,000) vs 72.70% reference. AG News, 4 labels: 94/100 vs 91% reference. DAIR Emotion, 6 labels: 86/100 vs 48% reference. Banking77, 72 labels: 64/100 vs 87% reference. This is the clear failure. MASSIVE, 18 scenarios: 71.50% macro accuracy across 52 locales; 86.25% pt-PT, 86.75% en-US. The classification pilots use only 100 examples each. A September 25 CPU run reproduced most numbers: 72.55% on Typed Decisions and 60/100 on Banking77 with 3 abstentions. On-Device Latency The lab published per-device measurements. On an Apple M4, one decision per call took a 33.15 ms median. On a Samsung SM-X510 tablet via ONNX Runtime, the median was 203 ms with 393.1 MB peak RSS. On an Intel Core i5-1235U, AG News decisions took a 107.83 ms median. Banking77 took 3,713.54 ms because it narrows 72 labels first. On X, @supersonicai claims Julia 1 classifies 5x faster than Jev on an i5 laptop. Treat that carefully. The Jev pilot measured Jev as a hosted service called from France, so latencies are not like-for-like. Introducing Julia-1:Our first classification model that runs on almost anything. Learn more https://t.co/YJCdEeIBSo pic.twitter.com/3cizsbG9ZB — Supersonic Labs (@supersonicai) September 26, 2026 Interactive Explainer Julia 1 vs Closest Competitors Feature Julia 1 TypeSafe Jev GLiNER2.5 Multi Developer Supersonic Labs TypeSafe AI Fastino Access Open weights Hosted API, early access Open weights License Apache 2.0 Proprietary Apache 2.0 Parameters 144.3M Not disclosed 287M Base encoder mmBERT-small Not disclosed mDeBERTa-v3-base Decision types Choice, score, yes/no Typed structured decisions Classification, NER, relations, records Options per call 2 to 20 (Router for more) Up to 255 Label list per schema Runs locally on CPU Yes No Yes Input price per 1M tokens $0.025 (planned API) $0.042 Free (self-hosted) AG News pilot 94% 91% 70% DAIR Emotion pilot 86% 48% 44% Banking77 pilot 64% 87% 61% Sources: Julia 1 model card, TypeSafe launch post, GLiNER2.5 Multi card, Jev benchmark pilot. Julia 1 pilots ran separately from the Jev and GLiNER runs. Limitations Julia 1 compares the answers you supply. It cannot be counted on for missing facts, algebra, or multi-step calculation. The Router can drop the correct label during narrowing. It is not a drop-in Transformers pipeline, and no Hugging Face inference provider serves it. Supersonic Labs advises evaluating on your own questions and keeping humans in the loop for consequential decisions. Key Takeaways Julia 1 is a 144.3M-parameter, Apache 2.0 decision model that runs on CPU. One API covers choice, ordered score, and yes-or-no decisions over 2 to 20 options. It beat Jev references on 3 of 4 pilots but trailed badly on 72-label Banking77. Median latency hit 33.15 ms per decision on an Apple M4. Training cost about US$104 in cloud GPUs; a $0.025/MTok API is planned. Check out the Model Weights, ONNX/WebGPU build, and Technical details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Supersonic Labs Releases Julia 1: A 144.3M-Parameter Open Decision Model That Runs on a CPU appeared first on MarkTechPost.
Sarvam AI has released Saaras V4, the newest generation of its speech recognition model. It covers all 22 scheduled Indian languages plus English, now including global English accents. Sarvam reports state-of-the-art accuracy across all 22 languages. Is it deployable? Yes, through Sarvam’s API today, using model=”saaras:v4″. Weights are not public, and Sarvam’s SageMaker self-hosting docs currently cover Saaras v3 only. What is Inside Saaras V4 Saaras V4 is an encoder-decoder system. An audio encoder converts the waveform into embeddings that carry phonetic and acoustic detail. A temporal-downsampling adapter then shortens that sequence and projects it into the language model’s embedding space. This keeps long recordings inside the decoder’s context budget. The decoder is Sarvam-3B, a 3B-parameter hybrid state-space language model trained from scratch in-house. It reads the audio features alongside a text prompt. It then emits the transcript autoregressively, feeding each token back as input for the next. Benchmark Results English: Sarvam evaluated 7 English datasets. Six come from Hugging Face’s Open ASR Leaderboard: AMI, GigaSpeech, LibriSpeech clean, LibriSpeech other, SPGISpeech and VoxPopuli. The seventh is AI4Bharat’s Indian-accented Svarah. Scoring follows the leaderboard’s normalization code. Saaras V4 posts the lowest average WER among the models Sarvam benchmarked. Indic: On Vistaar, Sarvam reports results across 10 Indian languages using both WER and LLM-WER. LLM-WER adds a semantic check. It separates real meaning errors from harmless spelling or formatting variants common in Indic scripts. Noisy audio: On Kathbath Noisy, measured with LLM-WER, Sarvam says Saaras V4’s error rate is under half that of Deepgram Nova-3 and GPT-4o Transcribe. The set includes compressed, clipped and background-heavy recordings. Language ID: On verified IndicVoices utterances, language identification error is 2.9% across the top 10 Indian languages. It is 5.22% across all 22. It is important to note that all numbers above are vendor-reported. Independent reproduction has not been published yet. 5 Output Modes From 1 Model The same audio can return 5 representations, selected through the mode parameter: transcribe (default): native script with numbers and dates normalized. verbatim: every word as spoken, fillers and spoken numbers kept. codemix: native script, with English words left in English. translit: the full utterance in Latin script. translate: an English translation with numbers normalized. Sarvam’s argument is simple. Handling these inside the model removes post-processing steps that can compound errors. Keyterm Prompting Keyterm prompting is new in V4 and works only with saaras:v4. You pass a JSON list under keyterms, with up to 50 terms of 64 characters each. Keyterms bias recognition; they do not guarantee output. Use codemix mode when a brand such as PhonePe must stay in Latin script. On IndicContextEval (paper, Interspeech 2026), Saaras V4 reports 16.03% WER in the L5 keyword-prompting setting. Sarvam says that is the lowest score on the benchmark. Streaming, Long Audio and Pricing Streaming: WebSocket with partial results and time to first token below 150 ms. REST: synchronous transcription for clips up to 30 seconds. Batch: asynchronous jobs up to 2 hours per file, with optional speaker diarization. SDKs: Python 3.9+ and Node.js 18+, plus LiveKit Agents, Pipecat and Vercel AI SDK integrations. Price: Sarvam lists speech-to-text at ₹30 per hour for real-time, streaming and batch, and ₹45 per hour with diarization. Saaras v3 stays the default model. V4 uses the same request shape, so switching is a 1-line change. Saaras V4 vs Closest Competitors These are the 3 systems Sarvam benchmarked against. Figures come from each vendor’s public docs and pricing pages, checked on September 26, 2026. Feature Sarvam Saaras V4 Deepgram Nova-3 ElevenLabs Scribe v2 OpenAI GPT-4o Transcribe Indian scheduled languages (of 22) 22 11 14 Not listed per language Total languages 23 (22 Indian + English) 45+ 90+ Multilingual Keyterm biasing Up to 50 terms Yes, paid add-on Up to 1,000 (batch), 50 (realtime), paid add-on Free-text prompt Built-in output modes 5 (transcribe, verbatim, codemix, translit, translate) Transcript plus Smart Formatting Verbatim or no_verbatim Transcript Real-time streaming WebSocket, under 150 ms TTFT (vendor claim) Yes (WebSocket) Scribe v2 Realtime, about 150 ms File streaming; live via Realtime API Speaker diarization Batch API Yes Up to 32 speakers Separate gpt-4o-transcribe-diarize model List price ₹30/hour $0.0052/min (multilingual, pre-recorded) $0.22/hour (batch) ~$0.006/min Self-hosting Not for V4 yet (v3 on SageMaker) Yes Cloud API Cloud API Key Takeaways Saaras V4 covers all 22 Indian languages plus global English in 1 model. A 3B hybrid state-space decoder, trained from scratch, sits behind an audio encoder. Keyterm prompting accepts up to 50 terms and scored 16.03% WER on IndicContextEval L5. 5 output modes and sub-150 ms streaming TTFT come from the same model. API-only today at ₹30 per hour; self-hosting docs still cover v3. Check out the Technical Details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English appeared first on MarkTechPost.
Theory is easier to trust once it’s running against a real API, so both examples in this article use the same tool — a get_weather function backed by
Tool Calling vs. Code Execution for AI Agents: Choosing the Right Action Primitive Read Post »
We use cookies to improve your experience and performance on our website. You can learn more at Privacy Policy and manage your privacy settings by clicking Settings.